LLM orchestration is the layer that decides which model call happens next, with what context, and what to do when it fails. Multi-agent orchestration is one shape of that layer: more than one independent LLM, each with its own context and its own decisions, whose outputs have to be reconciled into a single result. That reconciliation is where production systems break. Not the prompts, not the model — the coordination: shared state two agents both write to, a retry at hop three that re-fires a side effect from hop two, a supervisor that can’t tell which worker actually finished.
This guide covers those mechanics — topologies, state, error propagation, cost, observability — and the question most vendor pitches skip: whether you need a second agent at all. If you’re still working out what the surrounding layer does, start with what an AI orchestration engine is. This piece assumes one agent already in production and someone asking you to add another.
Key takeaways
- Most systems described as “multi-agent” are one agent with several tools — usually the right architecture, because one decision-maker means no coordination problem to solve.
- Cost is not linear. In Anthropic’s own data, agents use about 4× the tokens of a chat interaction and multi-agent systems about 15× (Anthropic Engineering).
- Choose the topology by the failure you can tolerate: supervisor (single point of failure), handoff (routing loops, inherited context you pay for), pipeline (one bad stage poisons everything downstream).
- Every action with a side effect needs an idempotency key — retries happen at every hop, not only at the top.
- Good fit: parallelizable work, context beyond one window, distinct tool and policy domains. Bad fit: sequential, write-heavy, tightly coupled tasks.
- Evaluate the system’s end task, not each agent’s output — an agent can be locally correct and still produce a wrong answer.
Deciding whether you need a multi-agent system at all? Skip the mechanics and jump to when multi-agent is the wrong answer. Building one? Read in order.
What’s actually different about coordinating several LLM agents, versus one agent with tools?
The difference is the number of decision-makers, not the number of capabilities. An agent with ten tools still has one control loop: the model picks a tool, your code runs it, the result returns into the same context, and the loop continues until the task is done. It reads as one trace, and when it fails, one component decided wrong.
Add a second agent and you have a distributed system. Two models hold two views of the task, each keeps its own context, and each can be locally right while the combined result is wrong. You are no longer debugging a prompt but a protocol: who owns the task, what crosses the boundary, who decides it’s finished, and what happens when one side goes quiet.
Which is why the major published playbooks start from the same place — don’t, unless you must. Microsoft’s Azure pattern catalogue is blunt: “Decision-making and flow-control overhead often exceed the benefits of breaking the task into multiple agents” (Microsoft Azure Architecture Center). Anthropic’s advice is to “find the simplest solution possible, and only increasing complexity when needed” (Anthropic Engineering).
In practice, most systems marketed as multi-agent are one well-built agent with a typed tool registry — that’s the architecture that ships. So the useful question isn’t “single agent vs multi-agent” in the abstract; it’s whether your task contains work that genuinely runs in parallel over independent context. If it doesn’t, a second agent adds cost and failure surface and buys nothing. (Still deciding whether an agent beats a rules engine or a chatbot? We covered that separately.)
What topology should coordinate your agents — supervisor, handoff, or pipeline?
Three topologies cover almost every real multi-agent system. What separates them isn’t their shape on a diagram but who decides the next step — and each choice buys one specific failure mode. Pick the one whose failure you can detect and survive.
Supervisor / router: one orchestrator keeps control
One LLM stays in charge: it breaks the task down, delegates sub-tasks to workers, and synthesizes their outputs. Anthropic describes this orchestrator-workers pattern as “a central LLM dynamically break[ing] down tasks, delegate[ing] them to worker LLMs, and synthesizing their results” (Anthropic Engineering). It has the best observability story, because one component knows the whole plan — which is why supervisor implementations (LangGraph’s among them) keep shared state in one place instead of letting workers hold it.
Its failures are structural: a vague delegation prompt is inherited by every worker, the synthesis step is the quality ceiling, and you pay coordination cost — tokens plus a round trip — on every delegation.

Handoff / peer: the task changes owner
One agent transfers the live task to another, which then owns it. OpenAI’s cookbook defines a handoff as “an agent (or routine) handing off an active conversation to another agent, much like when you get transferred to someone else on a phone call” (OpenAI Cookbook). In its reference implementation the full message history travels with the task — convenient, until that history is longer than the new agent’s useful context.
So handoff fails in two directions: pass too little and the next agent re-asks what the user already answered; pass everything and you pay for context it doesn’t need. And because agents route themselves, the path is unpredictable — Microsoft names “infinite handoff loops” and “unpredictable routing paths” as this pattern’s characteristic failures. Cap handoffs per task; log every transfer.
Pipeline / sequential: no shared control loop at all
Each stage’s output is the next stage’s input, in an order fixed in code. It’s the most predictable topology and the easiest to test, because control flow isn’t a model decision. Microsoft’s summary of the weakness is exact: “Failures in early stages propagate. No parallelism.” The real danger is quiet failure — stage two doesn’t crash on a plausible-but-wrong input, it elaborates on it. Validate every stage’s output against a schema before it becomes the next stage’s input.
| Topology | Who decides the next step | Dominant failure mode |
|---|---|---|
| Supervisor / router | The orchestrator LLM, per delegation | Single point of failure; coordination cost per hop; synthesis caps quality |
| Handoff / peer | Whichever agent currently holds the task | Routing loops; an inherited transcript the new agent pays for but didn’t need |
| Pipeline / sequential | Nobody — the order is fixed in code | An early stage’s wrong-but-plausible output poisons everything after it |
How do agents share state and pass context without corrupting each other’s work?
State is the real design decision behind every multi-agent orchestration pattern, and there’s no free option. Isolating context per agent is what makes the architecture worth having — Anthropic lists “information that exceeds single context windows” among the conditions where multi-agent systems earn their cost (Anthropic Engineering). But the less agents share, the more they misunderstand the job: thin delegation instructions meant “subagents misinterpreted the task or performed the exact same searches as other agents” — two agents, billing you twice for the same work.
So the question at every boundary isn’t “how much context can I pass?” but “what does the next agent need in order to be right?” Microsoft’s guidance is to choose deliberately between full raw context, a compacted summary, and a fresh instruction set with no history — and to compact between hops, because context grows with every agent’s reasoning and tool results.
Concurrent writes turn this from a quality problem into a data problem. LangChain puts the asymmetry plainly: “read actions are inherently more parallelizable than write actions,” and “conflicting write actions typically produce far worse outcomes than conflicting read actions” (LangChain). Parallelizing writes means solving two problems at once — communicating context between agents, then merging their outputs coherently. Microsoft names the antipattern directly: “sharing mutable state between concurrent agents, which can result in transactionally inconsistent data because of assuming synchronous updates across agent boundaries.” Three rules hold up in production:
- One writer per key. If two agents can write the same record, route the write through a single agent or service — not through both with a hope of ordering.
- Persist shared state outside the context window, so a long-running task resumes after an interruption instead of replaying from the top.
- Version what you write. Compare-and-set turns a silent overwrite into a detectable conflict — the difference between a bug in a trace and a bug in a customer’s data.
What happens when one agent fails — and how do you stop it becoming a runaway loop?
Two failure classes matter, and they need different controls. The first is propagation: agent A’s bad output becomes agent B’s input, and neither knows anything is wrong. Nondeterminism makes it expensive to chase — Anthropic notes agents “are non-deterministic between runs, even with identical prompts,” and “one step failing can cause agents to explore entirely different trajectories, leading to unpredictable outcomes” (Anthropic Engineering). You can’t reproduce the incident by re-running it.
The control is validation at the boundary, not better prompts. Microsoft’s reliability guidance is to validate agent output before passing it on, because “low-confidence, malformed, or off-topic responses can cascade through a pipeline,” and to surface errors rather than hide them. In practice: a typed contract at every hop, a timeout, a retry policy, and a circuit breaker on any agent that depends on a shared model endpoint. When two agents contradict each other and neither is clearly wrong, the resolution is usually a human checkpoint rather than a tie-break heuristic — coordination moves the reviewer, it doesn’t remove them.
The second class is runaway: a loop that keeps spending. A supervisor that can re-delegate, or a handoff graph that can route back, has no natural stopping point, and a stuck run costs tokens per minute. Ceilings must be explicit — maximum steps, maximum tool calls, maximum wall-clock time, and a token budget per run that halts rather than degrades. Anthropic scales the allowance to the task inside the prompt itself: “simple fact-finding requires just 1 agent with 3-10 tool calls… complex research might use more than 10 subagents with clearly divided responsibilities”. Without a rule like that, a trivial request can spawn a research project.
Why idempotency keys belong on every hop, not just the top
Agent A asks agent B to post a summary to a channel. B posts it, then its acknowledgement is lost — timeout, dropped connection, redeploy. A’s retry policy does exactly what you told it to and re-issues the request. Now the channel has two posts; had the action been a payment or an outbound email, you’d have an incident.
The fix is standard distributed-systems practice that agent frameworks won’t do for you: derive an idempotency key from the intent, not the attempt — a hash of (task id, action, target) — and have the executing tool treat a repeat key as a no-op that returns the original result. Two refinements: separate read retries (free) from write retries (keyed only), and keep retry policy, ceilings and key generation in the orchestration layer rather than inside each agent, so a new agent can’t quietly ship without them.

How much does running multiple agents actually cost, and what latency can it afford?
Adding a coordinating agent isn’t a 2× cost decision. The clearest published figures come from Anthropic’s research system, stated as their own measurement rather than a universal law: “in our data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats” (Anthropic Engineering). Your ratio will differ; the mechanism generalizes. Each agent pays separately for its instructions, accumulated context, reasoning and tool results — and the supervisor pays again to synthesize.
Topology shapes the bill. Sequential and handoff patterns invoke agents one at a time and accumulate cost per step; concurrent fan-out raises throughput but spikes consumption when several agents call models at once; planner-style orchestrations are hardest to forecast, because the planner iterates until it’s satisfied (Microsoft).
Latency runs the opposite way. Every sequential hop adds a full model round trip, so a four-stage pipeline stacks four of them additively — Microsoft lists “overlooking latency impacts of multiple-hop communication” among the common pitfalls. A supervisor fanning out parallel workers trades money for wall-clock time, but only if the sub-tasks are genuinely independent. If they depend on each other, you pay the multi-agent tax and still wait.
Three levers to build in before launch rather than after the first invoice: size the model per agent (classification, extraction and formatting rarely need your most capable model), meter tokens per agent and per run, and compact context between agents. Then answer one question honestly before adding an agent: is this task actually parallelizable, or am I paying the multi-agent tax for a sequential problem?
If that cost curve is why you’re reading this, our AI orchestration work usually starts by instrumenting what you already have — the per-agent numbers tend to change the design decision.
How do you observe and evaluate a multi-agent system once it’s live?
Per-agent tracing isn’t optional, because the failure you’ll be asked to explain is a cascade, and a cascade is only legible after the fact. Instrument every agent operation and every handoff: input, output, tool calls, tokens, cost and latency, tagged with a run id that ties the hops together. Microsoft’s framing is that this is ordinary distributed-systems troubleshooting.
Evaluation is where teams most often fool themselves. Score the system’s end task, not each agent’s local output: an agent can give a defensible answer to the sub-task it was handed and still contribute to a wrong final result. Exact-match assertions don’t work on nondeterministic output, so rubric or model-as-judge scoring is the practical route.
There’s a sharper reason to insist on an end-task eval. In Anthropic’s benchmark, “token usage by itself explains 80% of the variance, with the number of tool calls and the model choice as the two other explanatory factors” (Anthropic Engineering). Read that as a warning, not a recipe: if spend correlates that strongly with score, any eval that doesn’t measure the outcome will reward spend. You can make the dashboard improve by making the system more expensive.
When is multi-agent the wrong answer?
The symptom set is recognizable: you added a second agent, and now you have double the cost, double the failure surface, more latency, and no clear answer to “which agent broke this?” That combination almost always means the work was never parallel to begin with.
Anthropic’s poor-fit criteria are worth self-diagnosing against: “most coding tasks involve fewer truly parallelizable tasks than research,” “LLM agents are not yet great at coordinating and delegating to other agents in real time,” and domains that “require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today” (Anthropic Engineering). Add write-heavy work, for the concurrency reason above.
So the decision rule is short. Multiple coordinating agents can be worth it when all three hold: the work splits into genuinely independent sub-tasks, the information exceeds one context window, and the sub-tasks have distinct tool or policy domains — a read-only research agent and an agent permitted to send money are different security boundaries, not just different prompts. If any one is missing, the better build is usually one agent with a well-designed tool layer: typed tool registry, schema-validated calls, explicit ceilings, and a human checkpoint where the stakes justify one.
How does Trembit approach multi-agent coordination in production?
We start by trying to avoid it — and the clearest evidence is a system that looks multi-agent from outside and isn’t. For Chambiar, a workforce-automation platform, we built an AI assistant that acts across Gmail, Slack and Google Drive from natural-language prompts. A request like “find yesterday’s meeting notes, summarize them, and post the summary in the project channel” reads as delegation to a team of specialists; architecturally it is one LangChain agent with typed tool-calling — a single decision-maker invoking typed tools across three platforms. That’s why partial failure across API boundaries — the case study names a Drive search succeeding while the Slack post fails as a core challenge the build had to handle — stayed a tractable problem rather than a coordination problem.
Two lessons from that build shaped how we approach coordination generally. First, validate the model’s output against the target schema before execution, not after. Early on we executed the agent’s tool calls directly, and a malformed call — a channel ID that didn’t exist, a bad permission scope — only surfaced as an API error the user had already watched a “sending…” spinner for. Moving validation in front of execution turned unpredictable failures into predictable pre-execution corrections the bot could ask about. In a multi-agent topology, the same validation stops one agent’s malformed output becoming another’s confident input.
Second, keep coordination logic separate from the things being coordinated. Each integration was a hot-swappable module behind a standard interface, so when Calendar became the next priority we added it by implementing one module that registered itself with the existing agent, with no changes to the orchestration layer. That also makes a later move to multiple agents cheap: the boundary already exists as an interface. Across 15+ years and 50+ production video, voice and AI builds, the pattern we see most is teams adding agents to solve what was really a state or validation problem.
Next step: book a free 30-minute architecture session. Bring your agent setup and the coordination problem you’re hitting — the cost curve, the loop you can’t cap, the failure you can’t attribute — and we’ll tell you plainly whether it needs a second agent or a better single one. No deck, no pitch. Email welcome@trembit.com.
FAQ
What’s the difference between an orchestration engine and multi-agent orchestration? The orchestration engine is the layer that routes model calls, manages state and retries, and enforces policy — it exists even with a single agent. Multi-agent orchestration is one workload running on it, where several independent LLMs must be coordinated and their outputs reconciled. See what an AI orchestration engine is.
Do I need a framework like LangGraph or the OpenAI Agents SDK, or can I hand-roll this? A framework earns its place when you need durable state, replayable runs and checkpointing — rebuilding those correctly takes weeks. Hand-rolling is reasonable for two or three agents in a fixed sequence. What you shouldn’t hand-roll is the boring part: ceilings, retries, idempotency keys, per-agent tracing.
Can two agents run in parallel and edit the same data? Not safely, unless the write goes through a single owner. Parallel reads are fine; concurrent writes across agent boundaries produce transactionally inconsistent data, because each agent assumes its update landed. Route writes through one owner, version the records, and use compare-and-set so conflicts are detected rather than silently resolved.
Does multi-agent orchestration replace human-in-the-loop review? No — it changes where the checkpoint sits. With one agent, review typically gates the final action. With several, the highest-value checkpoint is often at the boundary: before a delegated result becomes another agent’s input, or when two agents disagree. The checkpoint design doesn’t change; its placement does.