← All articlesAgents IAProduitQualité6 min read

A chatbot answers, a swarm delivers: what parallel work actually changes

Why twelve specialised agents produce a better document than one model writing everything — and why the real gain is attention budget, not intelligence.

Ask a conversational assistant for a twelve-slide deck. You will get twelve slides. Look at the twelfth: it is visibly thinner than the first. That is not a comprehension failure, it is a budget failure — attention and output tokens are finite, and a model front-loads them.

That observation is the starting point of the architecture we use: instead of asking one model to do twelve things, ask twelve agents to do one thing each.

The production line

Analyste angle Chercheur faits datés Architecte design + plan Rédacteur · 1 Rédacteur · 2 Rédacteur · n ‖ en parallèle · 1 par slide Vérificateur + Relecteur ↺ renvoyée tant qu'elle ne passe pas
Each stage feeds the next; the writers all run at the same time, one per slide.

The Analyst frames the argument. The Researcher queries the web for dated facts and usable numeric series. The Architect designs the visual system and the slide-by-slide plan. Then one writer per slide starts, all at once, each with the full research dossier.

The measurable result is not "it is smarter". It is that the twelfth slide gets the same treatment as the first. The same effect holds for a regulatory audit: one expert per obligation means the twentieth obligation is handled as carefully as the first — which never happens when a single model is asked to "cover all applicable obligations".

A quality gate that cannot be charmed

A critic agent judges prose very well and facts very badly. We learned this on a flawless slide: three KPIs, perfect layout, and the third figure invented — plausible, right order of magnitude, right unit. The critic approved it, of course: nothing in the composition betrayed the fabrication.

So we split the review in two.

1 · MÉCANIQUE extrait TOUS les chiffres affichés vs dossier exhaustif · jamais fatigué mais plein de faux positifs liste courte 2 · JUGEMENT « lesquels sont inventés, lesquels sont des calculs ? » verdict slide bloquée ou validée
Extraction is exhaustive by construction; the model only rules on a shortlist.
  • Mechanical pass: every figure on the slide is extracted, every figure in the research dossier too, and the difference is the shortlist. A regular expression does not get bored and does not decide that twenty checks are enough.

  • Judgement pass: one model call, on that shortlist only — "which of these are inventions, and which are honest arithmetic?"

Our first version shipped the mechanical pass alone. It was unusable: it flagged every derived total, every rounding, every legitimately computed percentage. The mechanical pass narrows the question; it does not deliver the verdict.

If a check can be written as code, write it as code. Use the model for the question that remains.

What it means for the deliverable

A figure absent from the research blocks the slide. Content overflowing its box is measured in the browser and scaled down. An unmet mandatory requirement in a tender blocks delivery. None of those three rules is an instruction given to a model: they are checks, and they have tests.

You can watch the swarm work live, each agent in its own panel, and take over at any point.

You can try the swarm online, no credit card.

Launch my first mission →
Read next