SOCIAL AIBusiness Solutions

  Notebook

Ed. 19September 28, 2026 · 7-min read· For builders

I Shipped a 13-Agent Pipeline. Eight Agents Never Fired.

Multi-agent earns its keep on independent, parallel work. On one coherent change it just fragments the context.

I routed a single-section landing-page redesign through a 13-step multi-agent pipeline, and afterward the logs showed eight of those thirteen agents never fired. The main session did the copy, the design, and the implementation itself, because the scope was one section and one section doesn’t need a committee.

This was one of my own sites, not a client project, which is exactly why I could afford to look at it honestly. I’d built the pipeline the way I build most things now: a plan-gate up front, domain experts for SEO, GEO, copy, design, and implementation, then performance and regression agents, then reviewers on the back end. It looked disciplined on paper. It looked like the kind of setup a careful operator runs.

What did the routing logs actually show?

The logs showed the plan-gate approved the change, the SEO and GEO experts flagged nothing worth acting on, and the design and implementation work happened inside one continuous session because there was no natural handoff point for a single section. Five agents fired. Eight sat idle. The pipeline wasn’t wrong, exactly. It was oversized for the job.

That gap between routed and fired is the kind of thing you only see if you go back and check. I wrote about a similar audit of my own agent roster in what I measured when I compared agents installed against agents used, and the same pattern showed up there: most of the roster exists for scope you don’t hit on any given day. Having it available isn’t the problem. Routing every task through all of it is.

AGENTS ROUTED vs AGENTS FIRED, TWO PROJECTS OF SIMILAR SIZE051013single-sectionredesignnext project,same size5 fired8 idle6 fired0 idlefiredrouted but never fired
Fig. 1: On the single-section redesign I routed 13 agents and 5 fired, so 8 never ran. The next project of the same size I routed 6, and all 6 fired. Counted from my own routing logs.

Why does more agents feel safer than it is?

More agents feels safer because delegation looks like diligence, and diligence is what you want to feel when you’re shipping something client-facing or public. But delegation has a cost that doesn’t show up until something breaks: context fragments across handoffs, and now you’re tracing a bug through five agents instead of one continuous train of thought. Cognition made this case directly in their post on why they don’t build multi-agent systems, arguing that splitting a task across agents splits the context that would otherwise let one process catch its own mistakes.

I’d built an orchestration layer before this project that tried to solve that same tension by adding more structure on top, more gates, more explicit handoffs. It didn’t work either, and I ended up tearing most of it back out, which I covered in the post about the framework I deleted. The instinct both times was the same: when something feels risky, add process. Sometimes the fix is the opposite. Cut the process back to the size of the actual task.

Is multi-agent ever the right call?

Multi-agent is the right call when the subtasks are genuinely independent, parallel, and well-scoped, not when a single change is being routed through stages out of habit. Anthropic’s own research team found the opposite of Cognition’s result for their use case: in their write-up on building a multi-agent research system, a lead agent plus parallel subagents beat a single Claude Opus 4 agent by 90.2% on their internal research eval. That’s their eval, their task shape, not a universal law. Research fan-out is independent by nature: each subagent can go chase a different thread without stepping on the others. A one-section redesign isn’t that. It’s one thread.

Anthropic’s own guidance on this, in their post on building effective agents, is to find the simplest workable pattern first and only add complexity once it demonstrably improves the outcome. I’d read that piece before I built the 13-step pipeline. I didn’t follow it. I built the complex version first because it felt more thorough, then found out afterward that thoroughness and step count aren’t the same thing.

What did I get right on the same project?

The one place the multi-agent structure earned its keep on that same over-built project was review, not building. I ran three independent passes on the finished work: my own self-review, an adversarial subagent instructed to try to break it, and an external model doing a separate pass. Zero overlap between what each one caught. The adversarial pass alone found four defects my self-review missed entirely, and the external model caught two more that neither of the first two passes touched.

That result held up because review is genuinely parallel work. Each reviewer looks at the same finished artifact from a different angle and doesn’t need to know what the others are thinking. I wrote the full breakdown of that in the post on what one agent approved that I ended up rejecting. Separating reviewers worked because the task was already independent. Separating builders on a coherent single change did not, because the task wasn’t independent to begin with, it just got treated like it was.

What’s the actual rule now?

The rule I run now: multi-agent earns its keep on the review axis and on genuinely independent, parallel research or build work, not on a single coherent change. The very next project of similar size, I routed six steps instead of thirteen, gates and reviewers, no domain-expert fan-out, and it held. No idle agents in the log. No handoff where context got lost. The work still got checked three independent ways at the end. It just didn’t get built by committee in the middle.

That distinction, knowing which axis to fan out on and which to leave as one continuous pass, is most of what I actually get paid for. Not the orchestration itself. The judgment about what not to build into it. That’s the part of the engagements I run that doesn’t show up in a deliverables list, but it’s the difference between a pipeline that looks careful and one that actually is.

I still default to checking the logs after every project now, routed against fired, every time. Eight idle agents on one project was enough to change how I plan the next thirteen.

This is a field note, not a case study. If it maps to a problem you’re staring at, bring the actual problem.

Book a 30-minute call ← All notes