Multi-Agent Orchestration with LangChain and LangGraph: Collaborate or Compete?
Kuaray Deep Dive 03: when AI agents should collaborate, when they should compete, and when one is enough. Evidence, LangGraph patterns and cost math.
One agent, a team, or a tournament?
Kuaray Deep Dive #03 — Agent orchestration with LangChain and LangGraph: when agents should collaborate, when they should compete, and when a single agent is the right answer. The patterns, the evidence, and the math on what each one costs.
The problem: "multi-agent" is not an architecture
On September 23, 2026, Anthropic reported that roughly 950 Claude agents, running in parallel for 21 hours and consuming 210 million tokens, combed through more than 200,000 reverse transcriptases and surfaced a previously uncharacterized, CRISPR-like enzyme system [1]. Outside scientists called it intriguing, early and not yet understood [2]. As a demonstration of orchestration, though, it's hard to ignore.
Now the other side. Google Research and MIT ran 180 agent configurations across four benchmarks and three model families. On a parallelizable finance task, a coordinated team beat a single agent by 80.8%. On a sequential planning task, every multi-agent variant got worse, by 39% to 70% [3].
Both results are true. The lesson isn't "multi-agent good" or "multi-agent bad". It's that how agents are wired together, and whether they should be cooperating at all, decides the outcome. This Deep Dive is about those decisions, using the stack most teams actually ship with: LangChain for the agent loop and LangGraph for orchestration.
Part 1 — The vocabulary: collaborate or compete
Every multi-agent system answers two questions: who decides what happens next, and do the agents share one goal, or race toward it.
Collaboration patterns (one goal, split the work). LangChain's own documentation names them [6][7]:
- Subagents (orchestrator–worker). A main agent calls specialists as tools. Workers are stateless and see only the task they're given, which gives strong context isolation. Cost: one extra model call per hop, because results flow back through the orchestrator.
- Handoffs. The active agent changes based on state: a tool call flips a variable and the next turn is handled by another agent. Natural for support flows. Cost: sequential by design, and state management is on you.
- Router. A classification step sends the input to one or more specialists in parallel, then synthesizes. Good for distinct knowledge domains.
- Skills. Not really multi-agent: one agent loads specialized prompts on demand. Often the right answer when you thought you needed a team.
Competition patterns (same goal, independent attempts, someone picks):
- Best-of-N + judge. N agents attempt the task independently; a judge (a model, a test suite, or both) picks the winner.
- Debate. Agents argue over several rounds and converge. Popularized by Du et al. [16], and more fragile than it looks (Part 5).
- Generator vs. critic. One agent produces, a second agent with a clean context attacks the result. The cheapest form of competition, and the one that works best in production today [14].
Part 2 — LangChain vs. LangGraph in 2026: who does what
The two names still cause confusion, so here's the current split.
LangChain 1.x is the agent layer. Since 1.0 (October 2025), the center of gravity is create_agent: a model, tools, a system prompt, and middleware for things like summarization, human approval and dynamic tool selection [19]. The latest release as of this writing is langchain==1.4.3 (Sept 28, 2026) [13].
LangGraph 1.x is the runtime underneath. It's a graph of nodes and edges with typed state, checkpointing (every step persisted, resumable after a crash), interrupt for human-in-the-loop, Command for jumps between nodes, and Send for dynamic fan-out [9][10][11]. Current release: langgraph==1.2.12 (Sept 21, 2026), with 65M+ monthly downloads according to LangChain [11][13].
Deep Agents is the ready-made harness on top: a planning agent with a filesystem and a task() tool that spins up subagents in an isolated context ("context quarantine"), so the parent only sees final results, not dozens of intermediate tool calls [12].
One change matters if you followed 2025 tutorials: langgraph-supervisor is no longer actively maintained. The recommended replacement is simpler: wrap each worker agent as a tool and give them to a normal create_agent [8].
from langchain.agents import create_agent
from langchain.tools import tool
research_agent = create_agent(model, tools=[web_search], system_prompt="...")
math_agent = create_agent(model, tools=[calculator], system_prompt="...")
@tool("research_expert")
def call_research_agent(query: str) -> str:
"""Delegate research questions. Returns a short, sourced answer."""
result = research_agent.invoke({"messages": [{"role": "user", "content": query}]})
return result["messages"][-1].content
# ...same for call_math_agent
orchestrator = create_agent(
model,
tools=[call_research_agent, call_math_agent],
system_prompt="Split the task. Delegate with precise, self-contained instructions.",
)
Source: adapted from LangChain's supervisor migration guide [8]. The worker sees only query: that's the context isolation.
Handoffs use Command to move control to another agent node in the parent graph [9]:
from langchain.tools import tool, ToolRuntime
from langchain.messages import ToolMessage
from langgraph.types import Command
@tool
def transfer_to_billing(runtime: ToolRuntime) -> Command:
"""Hand the conversation to the billing agent."""
return Command(
goto="billing_agent",
graph=Command.PARENT,
update={
"active_agent": "billing_agent",
"messages": [ToolMessage("Transferred to billing.",
tool_call_id=runtime.tool_call_id)],
},
)
Simplified. The docs' main warning: always pair the tool call with its ToolMessage, and decide explicitly which messages travel with the handoff, or the history gets malformed and bloated [9].
And competition is a fan-out with Send plus a judge node:
import operator
from typing import Annotated
from typing_extensions import TypedDict
from langgraph.graph import StateGraph, START, END
from langgraph.types import Send
class State(TypedDict):
task: str
candidates: Annotated[list[dict], operator.add] # parallel writes are merged
winner: dict
def fan_out(state: State):
return [Send("attempt", {"task": state["task"], "seed": i}) for i in range(3)]
def attempt(s: dict):
answer = solver.invoke(s["task"]) # fresh context per attempt
return {"candidates": [{"seed": s["seed"], "answer": answer}]}
def judge(state: State):
# shuffled, anonymized candidates + tests; never the solvers' reasoning
return {"winner": pick_best(state["task"], state["candidates"])}
g = StateGraph(State)
g.add_node("attempt", attempt)
g.add_node("judge", judge)
g.add_conditional_edges(START, fan_out, ["attempt"])
g.add_edge("attempt", "judge")
g.add_edge("judge", END)
app = g.compile(checkpointer=checkpointer)
Illustrative: solver, pick_best and checkpointer are yours to define. With a checkpointer, a crashed attempt is retried from the last step instead of re-running all three [10][11].
Part 3 — The evidence: when more agents help, and when they hurt
Google Research + MIT, "Towards a Science of Scaling Agent Systems" [3] is the most controlled study so far: 180 configurations, five architectures (single agent, independent, centralized, decentralized, hybrid), GPT, Gemini and Claude.
| Finding | Number |
|---|---|
| Parallelizable task (Finance-Agent), centralized team vs. single agent | +80.8% |
| Sequential task (PlanCraft), every multi-agent variant | −39% to −70% |
| Error amplification, independent agents (no cross-checking) | 17.2× |
| Error amplification, centralized (orchestrator verifies) | 4.4× |
| Single-agent accuracy above which coordination stops paying | ~45% |
| Architecture correctly predicted on unseen configurations | 87% |
Two more findings from the same paper: tool-heavy tasks (16 tools, software engineering) pay a coordination tax, and once the single agent is already decent (≈45% on the task), adding agents gives diminishing or negative returns [3].
Anthropic's research system [4] points the same way from production. A Claude Opus lead with Sonnet subagents beat a single Opus agent by 90.2% on their internal research eval, a breadth-first, highly parallel task. The bill: agents use about 4× the tokens of a chat, and multi-agent systems about 15×. Token usage alone explained 80% of the performance variance.
Why multi-agent systems fail [5]. The MAST study annotated 1,642 traces from seven popular frameworks (ChatDev, MetaGPT, Magentic-One, OpenManus and others) and found 14 failure modes in three groups: system design, inter-agent misalignment, and task verification. Failure rates ranged from 41% to 86.7%. Adding a high-level verification step to ChatDev alone raised success by 15.6%: same model, better organization.
The pattern across all three: multi-agent wins on breadth (independent subproblems that can run in parallel) and loses on depth (one long chain of dependent decisions). And unchecked independence multiplies errors; a coordinator that verifies contains them.
Part 4 — Collaboration done right
1. Start with one agent. LangChain's own guidance: most agentic tasks are best handled by a single agent with well-designed tools; graduate to multi-agent only when you hit clear limits [6]. Skills often solve "too many specializations" without a second agent.
2. Context engineering is the whole game. "At the center of multi-agent design is context engineering: deciding what information each agent sees" [7]. In LangChain's own comparison on a multi-domain query, subagents processed about 9K tokens versus ~15K for skills, roughly 40% fewer, purely from context isolation. On repeated requests, stateful patterns (handoffs, skills) saved 40–50% of model calls [6].
| Pattern | Best for | Watch out for |
|---|---|---|
| Subagents | Parallel, independent subtasks; distinct domains | +1 model call per hop; vague delegation |
| Handoffs | Multi-stage conversations (support, sales) | No parallelism; message pairing; context bloat |
| Router | Separate knowledge bases, multi-source queries | Re-routing overhead when history matters |
| Skills | One agent, many specializations | Context accumulates over the session |
3. Parallelize reads, single-thread writes. Cognition, which in 2025 published "Don't Build Multi-Agents", updated its position in April 2026: multi-agent systems work best today "when writes stay single-threaded and the additional agents contribute intelligence rather than actions" [14]. Subagents research, explore and review; one agent commits.
4. Teach the orchestrator to delegate. Anthropic's first lesson: vague instructions like "research the semiconductor shortage" led subagents to duplicate each other's work. Each delegation needs an objective, an output format, tool guidance and clear boundaries, and effort should scale with query complexity [4].
5. Put verification in the graph, not in the prompt. MAST's cheapest fix was a verification step [5]; Google's 17.2× vs. 4.4× is the same lesson at scale [3]. In LangGraph that's a node with its own edge back to the worker, not a line in a system prompt.
Part 5 — Competition done right
Debate is weaker than it looks. Choi, Zhu and Li tested multi-agent debate on seven benchmarks and found that majority voting alone accounts for most of the gains usually attributed to debate. They prove that unguided debate forms a martingale over the agents' beliefs: on its own, it doesn't improve expected correctness [15]. If you want the benefit of several opinions, take independent answers and vote. Only add debate rounds with an explicit bias toward correction (for example, evidence that must be cited).
Best-of-N is only as good as the judge. Here's the math. Say one attempt succeeds 60% of the time and attempts are independent. With 3 attempts:
- all three wrong: 0.4³ = 6.4%
- all three right: 0.6³ = 21.6%
- mixed: 72%, and here the judge has to pick a right one
Success = 21.6% + 72% × judge accuracy.
| Judge accuracy (when there's a choice) | Best-of-3 | Best-of-5 |
|---|---|---|
| Random pick (≈53–57%) | 60% | 60% |
| 70% | 72.0% | 71.6% |
| 80% | 79.2% | 80.7% |
| 90% | 86.4% | 89.9% |
| Perfect | 93.6% | 99.0% |
Two things jump out. A judge that picks at random gives you nothing for 3× the cost. And going from 3 to 5 attempts barely matters until the judge is above ~80%. Invest in the judge before you invest in N. The best judges aren't models at all: tests, type checkers, schema validators, reproducible metrics. When a model has to judge, show it anonymized, shuffled candidates and the acceptance criteria, never the solvers' reasoning.
And a caveat the table hides: it assumes independent attempts. Same model, same prompt and same context make errors correlated, and correlated attempts fail together. Vary the seed, the prompt, or the model.
Generator vs. critic is the cheapest tournament. Cognition's review agent, running with a clean context (no shared history with the coding agent), catches about 2 bugs per PR, and 58% of them are severe [14]. It's competition in its most useful form: the critic is rewarded for finding what the generator missed, and writes stay with one agent.
Part 6 — Quantifying: what each topology costs
A reproducible scenario. Prices: Claude Sonnet 5.5, $2 input / $10 output per million tokens [18]. A single agent run on a medium task: 200K input + 20K output (accumulated across turns), no caching to keep it simple.
| Topology | How we estimate it | Cost per task |
|---|---|---|
| Single agent + tools | 200K in + 20K out | $0.60 |
| Generator + clean-context critic | + critic: 60K in + 4K out | $0.76 |
| Best-of-3 + model judge | 3 × $0.60 + judge: 40K in + 2K out | $1.90 |
| Orchestrator + workers | ≈ 15× / 4× of a single agent [4] | ≈ $2.25 |
When does it pay? Extra cost ÷ success gain = the minimum value of one successful task.
- Best-of-3 with an 80% judge: 60% → 79.2% (+19.2 pp) for +$1.30. Pays if a success is worth more than $6.77.
- Orchestrator: +$1.65 per task. If a success is worth $50 (half an hour of an engineer), it needs to raise success by 3.3 percentage points. On breadth tasks, Anthropic saw far more [4]. On sequential tasks, Google saw it go negative [3].
- Critic: +$0.16 per task. Almost always worth it where errors are expensive (code, contracts, money).
Being honest about the math. Token volumes are an assumption: measure your own (LangSmith or your gateway shows it). The 3.75× orchestrator factor comes from Anthropic's averages, not your workload. Prompt caching changes everything for the better, especially for parallel attempts that share a long prefix (see Deep Dive 01 on LLM gateways and routing). And the best-of-N table assumes independent attempts.
Part 7 — A production checklist for LangGraph
- Checkpointer from day one. A multi-agent run that dies in step 14 of 20 should resume, not restart. Use a durable checkpointer (Postgres) in production [10].
- Budgets per agent, not per run. Recursion limits, max tool calls and a token ceiling per worker. Anthropic's orchestrator scales effort to query complexity by rule, not by vibes [4].
- One writer. Parallel
Sendbranches for reading, exploring and attempting; a single node that commits side effects [14]. - Verification as a node. Tests, validators or a critic with clean context, with an edge back for retries [3][5].
- Human approval on irreversible actions via
interrupt, ideally on the writer node only. We covered gate design in Human in the loop approval for AI agents. - Trace everything. Per-agent tokens, latency and outcomes. Without it you can't tell whether the team is better than the single agent, and that's the only question that matters.
- Cross-vendor agents? Use a protocol. A2A reached 1.0 and 150+ supporting organizations in April 2026, with native support in Azure AI Foundry and Bedrock AgentCore [17]. Inside one codebase, a LangGraph graph is simpler. For the governance side of agent-to-agent protocols, see the agent protocol governance gap.
The caveats
- Benchmarks aren't your workload. Google's 87% prediction accuracy is on their task mix [3]; Anthropic's +90.2% is on an internal research eval [4]. Run your own single-agent baseline first. We catalogued the failure modes that show up in production in why AI agents fail in production.
- The 950-agent discovery is early. The function of the enzyme system is still unknown, the preprint isn't peer-reviewed, and the lab work was done by humans [1][2]. It shows what parallel search can do, not that swarms replace scientists.
- Fast-moving APIs. LangChain's multi-agent docs changed substantially in 2026 (
langgraph-supervisorno longer maintained, subagents as tools recommended) [8]. Pin versions. - More agents, more attack surface. Every agent that reads untrusted content can pass a prompt injection to the next. Isolate contexts, and don't let workers write to shared memory without review.
Conclusion
"Multi-agent" isn't an upgrade. It's a set of trade-offs. Teams of agents win on breadth: independent subproblems, parallel search, separate domains. They lose on depth: long chains of dependent decisions, where every handoff is a chance to drop context. Competition works when a good judge exists, and the cheapest version, a critic with a clean context, is the one that pays off most often.
In LangChain and LangGraph terms: start with create_agent and good tools. Add subagents as tools when context gets crowded. Use Send for parallel attempts, a verification node to contain errors, and keep a single writer.
At Kuaray Tech, we design agent topologies the same way we design distributed systems: measure the baseline, isolate state, verify every write. That's the core of our AI and agent engineering work, running on the same tracing and infrastructure stack as the rest of your platform through our DevOps practice. If you're weighing a framework against custom code, we compared the options in n8n vs custom AI agent development.
Talk to Kuaray about your agent architecture — we benchmark your single-agent baseline, pick the topology the evidence supports, and build it in LangGraph with verification and budgets from day one.
References
- Anthropic — Claude discovers a novel enzyme system (Sept 23, 2026). https://www.anthropic.com/news/claude-discovers-novel-enzyme-system
- Smithsonian Magazine — Anthropic Says Its A.I. Discovered a New Enzyme System That Resembles CRISPR (Sept 29, 2026). https://www.smithsonianmag.com/smart-news/anthropic-says-its-ai-discovered-a-new-enzyme-system-that-resembles-the-revolutionary-gene-editing-tool-crispr-180989578/
- Kim, Y., Liu, X. et al. — Towards a Science of Scaling Agent Systems. arXiv:2512.08296; Google Research blog (Jan 28, 2026). https://research.google/blog/towards-a-science-of-scaling-agent-systems-when-and-why-agent-systems-work/
- Anthropic Engineering — How we built our multi-agent research system (Jun 13, 2025; updated Jan 2026). https://www.anthropic.com/engineering/multi-agent-research-system
- Cemri, M. et al. — Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657 (NeurIPS 2025). https://arxiv.org/abs/2503.13657
- LangChain Blog — Choosing the right multi-agent architecture (Jan 14, 2026). https://www.langchain.com/blog/choosing-the-right-multi-agent-architecture
- LangChain Docs — Multi-agent. https://docs.langchain.com/oss/python/langchain/multi-agent
- LangChain Docs — Migrate from langgraph-supervisor. https://docs.langchain.com/oss/python/migrate/langgraph-supervisor
- LangChain Docs — Handoffs. https://docs.langchain.com/oss/javascript/langchain/multi-agent/handoffs
- LangChain Docs — LangGraph fault tolerance. https://docs.langchain.com/oss/python/langgraph/fault-tolerance
- LangChain Blog — 3 years of graph engineering with LangGraph (Jul 22, 2026). https://www.langchain.com/blog/3-years-of-graph-engineering-with-langgraph
- LangChain Docs — Deep Agents: subagents. https://docs.langchain.com/oss/python/deepagents/subagents
- PyPI — langgraph 1.2.12 (Sept 21, 2026), https://pypi.org/project/langgraph/ · Releasebot — LangChain release notes, September 2026, https://releasebot.io/updates/langchain-ai
- Cognition — Multi-Agents: What's Actually Working (Apr 22, 2026). https://cognition.com/blog/multi-agents-working
- Choi, H. K., Zhu, X., Li, Y. — Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models? arXiv:2508.17536 (NeurIPS 2025). https://arxiv.org/abs/2508.17536
- Du, Y. et al. — Improving Factuality and Reasoning in Language Models through Multiagent Debate. arXiv:2305.14325 (ICML 2024). https://arxiv.org/abs/2305.14325
- Linux Foundation — A2A protocol surpasses 150 organizations… (Apr 9, 2026). https://www.linuxfoundation.org/press/a2a-protocol-surpasses-150-organizations-lands-in-major-cloud-platforms-and-sees-enterprise-production-use-in-first-year
- Anthropic — Claude pricing (accessed Oct 2, 2026). https://platform.claude.com/docs/en/about-claude/pricing
- Taskade — What Is LangChain? History, release dates & roadmap (2026). https://www.taskade.com/blog/langchain-history
Versions, prices and benchmarks checked at the time of writing. The cost and best-of-N math is ours, using Anthropic list prices and the assumptions stated above.