Skip to content

An onboarding agent provisions a new hire in under a minute. It reads the role, creates the accounts, orders the laptop, books orientation, assigns the first training. Four systems, one clean path, no human in the loop.

Then you point it at more departments and ten more systems, with contractor exceptions and approvals by cost center, and it starts granting the wrong access. Nothing crashes. It just imperceptibly degrades the more you ask of it.

That gap, between the demo and a system that holds up, is where most agent projects stall, and a better model will not close it. The shape of the problem decides what you build, so it is worth knowing before you scale.

How an agent works

Almost every agent runs the same loop.

The agent loop: plan, execute, remember, reflect

It plans the task into steps, executes by calling tools, writes what matters to memory, and reflects on what worked before the next step. For a narrow job the loop is cheap and reliable: a handful of steps, a few tool calls, a clear finish, a few seconds, a few cents. This is why demos land. The problem is bounded, so the agent looks brilliant.

The loop does not change when you scale. The cost of running it does, at every stage.

The same loop, more expensive at every stage as scope grows

Planning has more steps and more branches to weigh. Execution has more tools to choose between, so picking the right one gets harder. Memory grows, so each step carries more context and more of it is noise. Reflection has more history to weigh and less signal to find in it. Same four boxes, every decision more expensive, and the costs compound rather than add.

Two things people call “scaling”

When a demo works, the instinct is to scale it, and the word hides two different jobs.

One is throughput: more users, more requests, the same task. That is the familiar operations problem, solved the way you scale any service, by adding capacity. Behavior does not change, so a stateless service behind a load balancer just runs on more machines.

The other is scope: the same system doing more. The reflex, almost verbatim from every planning meeting, is “more tools, more domains, less supervision.” This is the job that breaks, because widening scope changes the work itself. You are enlarging the set of decisions the agent makes, and every decision is a judgment call that can be wrong.

What widening scope costs

Three things degrade as scope grows. Each has a number, and each has a fix worth knowing.

Tool selection collapses as the menu grows. Give a model 5 tools and it picks well. Give it 50 and it starts reaching for the wrong one, because it has to tell apart far more options that look alike. The problem starts before the model even chooses. In my own tool-retrieval test, the chance of the right tool landing in the top 5 search results fell as the catalogue grew, for keyword and semantic search alike. Keyword search (BM25) fell hardest, to 57% at 10,000 tools and 44% across the full 37,000. The strongest dense retriever I tested held up better but still dropped, to 72% at 10,000 tools and 57% across all 37,000. Even the best retrieval leaves the right tool out of the top 5 more than 40% of the time at full scale.

Tool-retrieval recall falls as the catalogue grows, on the ToolRet corpus

On top of that retrieval gap, the model’s own selection accuracy drops. One study swept a tool pool up to 11,100 entries and watched accuracy fall sharply as the menu grew. This is why GitHub Copilot caps a single chat request at 128 tools, and why Cursor sends only a limited set of active tools to the agent at all, on the order of dozens, and the rest never reach the model. The fix keeps every capability but stops showing the model all of them at once: retrieve the handful relevant to the current step and expose only those. In the same study, switching from dumping every tool into the prompt to retrieving only the relevant ones lifted tool-selection accuracy from 13.6% to 43.1%, more than triple, while cutting the prompt by over half.

Context degrades as it fills. Past a point, a fuller prompt just gives the model more to read through, and accuracy drops. Across 18 models tested, accuracy fell as input length grew, even on simple retrieval with the task held constant, and several models did better on shuffled context than on a clean, coherent document. Attention cost grows with the square of the input, so a fuller context is slower and less accurate at the same time. The fix is to manage context deliberately: summarize and compact long histories, keep each sub-task on its own clean context, and pull in only what the step needs.

Reliability compounds the wrong way. Model each step as 99% reliable and independent, and a task’s success is 0.99 raised to the number of steps.

A 40-step workflow at 98% per step succeeds less than half the time.

Compounding reliability across task length

At 99% per step, a 100-step task lands at 36.6%. At 99.9% per step, still only 90%. And real steps are not independent: a wrong early step poisons the ones after it, so failures correlate and the true numbers come in worse than the curve. Tell a travel agent to book a trip to Washington and it reads the District of Columbia when you meant the state. That assumption goes into the plan, drives the bookings, and gets written to memory as fact. Every later step is locally correct and globally useless, and because the agent runs on its own, no one is watching the moment it goes wrong.

Every tool and document you add to widen scope taxes every decision the agent makes after it. You pay twice, in the bill and in the failure rate. That is the scope tax: the wider the system, the more each choice inside it costs, and the further your cost per finished task drifts from the demo.

The root cause is ownership

The cause is structural, which is why a better model does not fix it. When one agent owns every decision, plan, tools, memory, and reflection, there are no boundaries and no separation of concerns. Every new capability widens the same context, lengthens the same plan, and crowds the same memory.

The same thing happens inside a company. When every decision, engineering, hiring, a customer refund, routes through one person, even small calls slow down as the company grows, because that person holds more context and switches between more domains. An agent behaves the same way. What limits the system is the span of responsibility you hand one agent, well before it is the raw capability of the model. That makes scaling a design decision, not a model-shopping one.

The fix is decomposition, and it runs two ways

Decompose into bounded roles, then scale horizontally or vertically

Scaling scope well means giving up the single agent that owns everything and splitting the work into bounded roles, each with less context, fewer decisions, and a narrower job. You can scale horizontally, adding agents for distinct jobs, which makes each capability easier to reuse but pushes cost into a coordination layer that has to keep them aligned. You can scale vertically, giving one agent more tools and sub-agents, which avoids coordination but concentrates complexity and cost in that agent. Either way the work moves to a new place.

And the new place has its own failures. The moment you have more than one agent, you inherit the problem of getting them to agree. A Berkeley study, Why Do Multi-Agent LLM Systems Fail?, hand-annotated 150 conversation traces from five open-source multi-agent frameworks and sorted what went wrong into three groups, in order of how often each appeared:

The two biggest groups have nothing to do with model capability. They are design and coordination problems. This is not abstract: CrewAI’s own documentation lists “agents delegate back and forth indefinitely” as a known failure to design around, and recommends turning delegation off on specialist agents to stop the loop. On one of the frameworks the study examined, a state-of-the-art system still came back correct as little as 25% of the time. The study’s conclusion is blunt: most of these failures are fixed by better system design, not a better model.

This is why I default against multi-agent fan-out for write-heavy work, and why the two most-cited results on it agree once you read past the headlines.

Anthropic’s multi-agent research system beat a single agent by 90.2% on their internal research eval, at about 15 times the tokens. The work that paid off was parallel and read-heavy: several agents searching independent strands, the results merged once. Anthropic name the limit in the same post. Tasks that “require all agents to share the same context or involve many dependencies between agents are not a good fit,” and “most coding tasks involve fewer truly parallelizable tasks than research.”

Cognition, who build a coding agent, ran straight into that limit. In June 2025 they argued against multi-agent systems because “actions carry implicit decisions, and conflicting decisions carry bad results.” Two agents writing in parallel make choices that collide at merge time. Ten months later, in April 2026, they published a revision that sharpened the rule instead of dropping it. Multi-agent works for them “when writes stay single-threaded and the additional agents contribute intelligence rather than actions.” Many agents read, review, and advise; one thread owns the writes.

Both camps reached the same boundary from opposite sides. Parallelize the thinking, serialize the writes.

Many agents read and advise in parallel; one writer thread owns every commit

In practice this looks like a code-review loop: several agents read the same change and flag issues in parallel, while one agent owns the single commit that results.

Multi-agent buys you parallelism and a coordination bill to go with it.

The 15 times multiple lands differently when you pay US-dollar token prices against APAC revenue rather than a venture subsidy. Whether the trade pays off comes down to whether your work is mostly reads or mostly writes.

When to split a capability, and when to embed it

My default is to embed, and make a capability earn its own agent.

A research assistant makes it concrete. Fact-checking deserves its own agent: it applies across the whole system, and you can verify its output on its own. Re-ranking retrieved results belongs inside the retrieval agent, glued to that agent’s context, where splitting it out would only buy you a handoff for nothing. Split a capability when it is reusable, independent, and verifiable. Embed it when it is coupled and context-dependent.

The number that should drive the call is cost per successful outcome, and it is easy to compute once you have a verifiable check. Count a run as successful only when it passes that check, then divide total spend across every agent, tokens, tool calls, and retries, by the number of passing runs.

So I built a controlled version and measured it. The agent reads a new hire’s role, department, employment type, seniority, and location, then has to produce the correct provisioning plan, the accounts, the access tier, the equipment, the approvals, by looking the rules up through tools. The rules carry the override traps a real policy has, like a contractor being capped at standard access no matter their seniority. I ran 60 cases two ways on the same model: one monolithic agent owning the whole job, against a planner with a separate policy-verifier agent.

The monolithic agent got 58% of plans exactly right. Splitting the verifier out raised that to 68%, and it cost 1.48 times as much per correct plan, because it ran up about 1.7 times the spend to buy those ten points. That is the bill for one extra agent; the fifteen-times figure from earlier is what a full fan-out costs. The split is the better build only when a correct onboarding is worth the premium.

A controlled onboarding agent over 60 cases: splitting a verifier out raised the pass rate and the cost per correct plan together

Decomposition bought reliability and a coordination bill in the same move, and the only way to see whether the trade paid was to divide spend by successes, not to read the pass rate alone. Every number here is from one controlled 60-case run, a directional result on a single task and one model.

And if you cannot define the pass check at all, that is the signal the workload is not ready to scale, the same verifiability test the table turns on.

What keeps it bounded

Decomposition decides where complexity lives. A few practices decide whether it stays bounded once it is there, and the agent frameworks have converged on the same short list.

Treat the context window as scarce. Compact long histories into summaries before they fill, give each sub-agent its own clean context so one agent’s noise does not leak into another’s, and pull in only what the step needs. These are first-class features now. The Claude Agent SDK runs subagents that “use their own isolated context windows, and only send relevant information back to the orchestrator,” and ships a compact step that “automatically summarizes previous messages when the context limit approaches.” Anthropic’s own work on long-running agents is honest that “compaction isn’t sufficient” on its own, and pairs it with state written outside the window: a progress file, a checklist of features that start marked failing, git commits at each step.

Reintroduce determinism where you can. This is the clearest pattern across the frameworks right now. Each of them has spent the last year stepping back from free-form, model-driven coordination toward fixed, inspectable control flow. CrewAI tells you plainly that “for any production-ready application, start with a Flow,” its deterministic, event-driven primitive, and to reach for an autonomous agent team only inside a step that genuinely needs one. The OpenAI Agents SDK puts it the same way: “orchestrating via code makes tasks more deterministic and predictable, in terms of speed, cost and performance.” Microsoft’s successor to AutoGen, its Agent Framework, is built around a graph-based workflow engine with checkpointing. The lesson under the version notes is the same one: if a task has a known shape, a fixed code path runs it more cheaply and reliably than a model deciding the route each time. Reserve agentic freedom for the steps that actually need judgment.

Put the human at the boundary, not on every step. A checkpoint earns its keep exactly where a wrong autonomous action is expensive, and lets the rest run. LangGraph makes this concrete with an interrupt() call: it pauses the run, checkpoints the state, surfaces the decision to a person, and resumes from the same point on their reply.

from langgraph.types import interrupt, Command

def approve_high_risk_action(state):
    # pause the run here and wait for a human to decide
    decision = interrupt({"action": state["proposed_action"]})
    return {"status": "executed" if decision == "approve" else "halted"}

# after the person replies, resume; the resume value becomes interrupt()'s return
graph.invoke(Command(resume="approve"),
             config={"configurable": {"thread_id": thread_id}})

The agent runs untouched right up to the risky action, blocks for a person, then continues with their answer. The call also carries the kind of honest caveat that marks real engineering: on resume the paused node re-runs from its start, so any side effect before the pause has to be safe to repeat. One good checkpoint at the risky boundary beats a person babysitting every step.

Read straight off the major frameworks, the pattern is hard to miss. Each gives you a way to coordinate, a way to make the flow deterministic, and a place to put a human:

The vocabulary differs across them, and every one lands on the same three levers: bound the context, make the flow deterministic where you can, and keep a human at the risky edge.

Instrument the one number. You cannot manage cost per successful outcome if you are not measuring it. Track pass rate against a defined check and spend per passing run, per workflow, and watch both as you widen scope.

In production, the agent is a coordination layer

In an operating business, the agent holds context, orchestrates actions across systems, applies the rules, and hands control to a human exactly where the rules require it. The hard part is rarely the reasoning; it is doing this reliably across systems that were never designed to be driven by one.

A production agent as a coordination layer, with a human at the high-risk boundary

Four patterns cover most of what enterprises actually automate. What separates a working deployment from a broken one is where the human boundary sits, so the column that earns its place is what breaks when you put it in the wrong spot.

In every one, the hard part is orchestrating reliably across systems while respecting policy and timing, and knowing which decisions to keep for a person. The reasoning is rarely what breaks. The winning design is narrow scope, rules applied consistently, a human at the boundary, built to integrate with the systems around it. It reads like a well-designed part of a larger system, which is the point.

The decision that matters

Capability is rising fast. The length of task a model can finish reliably has roughly doubled every seven months since 2019, and possibly faster since 2024. That number measures the human time for tasks the model clears about half the time, not hours of unattended autonomy. It still moves the frontier outward. The shape of the curve underneath stays the same, and an agent that owns too much still gets expensive and fragile at the edge.

So the real decision is where you let complexity accumulate, and what you refuse to let scale. Decide it deliberately, measure cost per successful outcome, and widen only where that number holds.

—Viz

Next: We deleted fifty years of information retrieval and called it RAG

Read next