Run LangChain's Open Deep Research multi-agent pipeline locally
LangChain's open-source deep research agent, built on LangGraph, runs research sub-agents in parallel, compresses what they find and writes a final report. It works with many model providers, search tools and MCP servers.
The problem
Commercial deep-research tools are closed and tied to one vendor. This open reference implementation lets teams run, inspect and adjust a multi-agent research pipeline with their own models and search.
Steps
- Clone langchain-ai/open_deep_research and create an environment with uv venv and uv sync.
- Copy .env.example to .env and add your model and search API keys. Tavily is the default search.
- In configuration.py, choose models for the summarization, research, compression and final-report steps.
- Start LangGraph Studio with the command below and send a research question.
- Optionally, connect MCP servers or switch to native web search for Anthropic and OpenAI models.
From the official docs
uv sync
cp .env.example .env
uvx --refresh --from "langgraph-cli[inmem]" --with-editable . --python 3.11 langgraph dev --allow-blocking
Results
- On Deep Research Bench, the default GPT-4.1 configuration scored 0.4309 and the GPT-5 configuration 0.4943.
- Earlier plan-and-execute and supervisor-researcher versions are kept in the repo for comparison.
As reported by the source (LangChain on GitHub); AgentGid did not measure these figures.
This suits a developer or small engineering team already comfortable with Python, LangGraph Studio and API keys, since the setup means choosing models for four separate pipeline steps. LangGraph is free (OSS), but you pay for model and search usage across parallel sub-agents, so costs can climb. The benchmark scores are from the project's own configurations, so test on your own questions. AI SDK (Free (OSS)) is a growing alternative if you prefer a different stack.
The agent used here
Similar use cases
How NBIM rolled out Claude to investment and compliance staff
NBIM, which manages Norway's sovereign wealth fund, deployed Claude across investment research, ESG analysis, compliance and data operations. Anthropic reports weekly time savings…
According to Anthropic, employees save 20% of their week on the analytical and operational work they now do with Claude.
How The AA answered routine trading questions in Teams with Databricks Genie
UK roadside assistance provider The AA used the Genie Conversation API to answer trading teams' routine data questions inside Microsoft Teams. Databricks reports faster answers to…
Databricks reports a 70% efficiency gain in answering routine queries.
How Oxford PharmaGenesis ran a 500-paper literature review in a week with Elicit
Health-science communications consultancy Oxford PharmaGenesis used Elicit for screening, data extraction and reporting in literature reviews, with experts checking the results. E…
Elicit reports answering 40 research questions across 500 papers in under a week.
Guides
AI Agent vs Chatbot vs Agentic AI: What Is the Difference?
Plain-language definitions of AI agents, chatbots, agentic AI, LLMs, workflows and skills, how they relate, and how to tell which one you are looking at, with real examples.
How to Build an AI Agent: No-Code Builders, Frameworks and SDKs
A practical guide to building your first AI agent: choosing between no-code builders, open-source frameworks and model vendors' SDKs, the core agent loop, tools, guardrails and testing, with live usage data for each option.
What Is an AI Agent? A Practical Explanation With Real Examples
What makes software an AI agent, how agents differ from chatbots and plain language models, the main types in use today and how to judge one.