Same model, same tools, same total context. A supervisor coordinating nine domain specialists produces materially deeper investigations than one generalist doing everything. That result decided the architecture — and three months later, measuring it properly turned up a problem with the axis I had specialised along.
The shape
A supervisor receives the alert, frames the investigation, and writes down the competing explanations before anything is queried. It then dispatches specialists — identity, cloud control plane, endpoint, mail, network, web, Windows, Unix, containers — each with its own knowledge scope and its own slice of the toolset. Each specialist returns findings. The supervisor reads them, decides what is still unresolved, and dispatches again or stops.
No cycles, no negotiation between specialists, no shared scratchpad. A star, with the supervisor at the centre. That simplicity is what made the no-framework decision straightforward, and it is also what makes the coordination cost bearable.
Why the generalist goes shallow
The interesting part of the April experiments was not that specialists scored better. It was the failure mode of the generalist, which was consistent: it stopped early, and it stopped confidently.
A generalist with full tool access asks the obvious questions and then concludes. It does not know what it has not checked, because knowing what to check next is domain knowledge, not reasoning ability. A specialist that has been told it owns identity knows there is a difference between a sign-in log and an audit log, knows that a service principal does not have interactive sign-ins, and knows which absence is meaningful and which is structural.
That last one matters more than any of the rest. In investigation work, a clean result is only informative if you know the query could have returned something. "I looked and found nothing" and "I looked in a place that never has anything" produce identical output and mean opposite things, and telling them apart is domain knowledge.
The coordination cost is real
This is not free, and the honest accounting matters. Every dispatch carries a brief. Every finding comes back through a contract that has to be parsed and validated. The supervisor's context grows with each round. On a typical investigation the supervisor alone consumes a substantial share of the token budget before a single specialist runs.
Early on, the specialists also returned prose, which meant the supervisor was reading nine essays and re-deriving structure from them. Replacing that with a structured finding contract — what was queried, what came back, what direction it points, and the tool call IDs behind it — was one of the higher-value changes in the project. It cut the token cost and, more importantly, it made specialist output auditable rather than merely readable.
The problem: I specialised along the wrong axis
The roster maps human SOC knowledge domains onto agents, because that is how analyst expertise is organised and it made obvious sense. Months later, when the system started fusing specialist findings into a single verdict, a question surfaced that the original design had no answer for: if sound evidential reasoning requires independent sources, are nine specialists nine sources?
They are not. I measured both axes against the actual code. Nine expertise domains resolve to roughly four distinct evidence groups, because several specialists read the same underlying tables. Five of the nine are strict subsets of another specialist's table access. Two analysts querying one table are one observation, not two — and treating them as two inflates confidence exactly when they agree.
| Axis | Unit | Good for | Fails at |
|---|---|---|---|
| Expertise | Knowledge domain | Knowing what to ask and what silence means | Counting independent evidence |
| Sensor | Telemetry source | Counting independent evidence | Far too coarse — six of nine collapse into one endpoint agent |
The resolution: stop using the agent as the unit
The tempting fix is to redraw the roster along sensor lines. I costed that and it is worse: it collapses most of the roster into a single endpoint agent, destroys the domain knowledge that made specialists useful in the first place, and still does not deliver independence, because the shared apparatus — same query builder, same schema assumptions, same model — survives any partition of the agents.
So the answer was to separate two things that had been conflated. Specialists stay on the expertise axis, because that is what produces good questions. Evidence gets grouped on the sensor axis, independently, and the fusion arithmetic counts groups, not agents. A finding from the identity specialist and a finding from the cloud specialist that both rest on the same audit table land in one group and are merged, not added.
That distinction — who investigates versus what counts as a witness — took three months to surface and is now one of the load-bearing rules of the system. The architecture from April survived; the naive reading of it did not.
What I would tell someone starting this
Specialise, and do it along whatever axis produces the best questions. Then, separately and before you build any confidence arithmetic, work out what your independent sources actually are. They will not be your agents. If you discover this after the fusion maths is built, you will be reworking the thing that decides verdicts, which is the most expensive place in the system to be wrong.