← Back to blog

Four experiments in five days that decided the architecture

In April I spent five days running four experiments against the same incidents, changing one variable at a time. They decided the architecture of the whole system, and five months later I have not revisited any of the conclusions. Here is what each one tested and what it showed.

The setup was deliberately cheap. Same incidents, same model, vary one component, read the outputs the way a senior analyst reads a junior's case notes: is this complete, is it correct, would I have to redo it?

I want to be precise about what that method can and cannot support, so the limitation is stated up front rather than buried at the end: these were qualitative practitioner reads on a small number of cases with no ground-truth labels. They establish direction, not magnitude. Any number you could extract from them would be false precision.

Experiment one: does tool access change the verdict?

The same supervisor agent, with and without access to the SIEM query tools. Everything else identical.

Tool access changed verdict accuracy, which is unsurprising. What was more interesting was how the failures differed. Without tools, the model produced confident, coherent assessments derived entirely from the alert description and its own priors — and they were wrong in a specific way, defaulting to the textbook interpretation of the alert class rather than what had actually happened in that tenant. With tools, it was sometimes wrong too, but the wrongness was traceable to a query that returned something misleading.

That distinction turned out to be load-bearing for the entire project. A wrong answer you can trace to a specific piece of retrieved evidence is a bug. A wrong answer derived from priors is a personality trait, and you cannot debug it.

The finding in one line. A model reasoning over evidence it retrieved is a different kind of system from a model reasoning over an alert description — not a better version of the same system.

Experiments two and three: does method matter?

The same supervisor, with and without an explicit methodology prompt — investigative order of operations, the discipline of writing hypotheses before querying, how to treat a null result.

The output that changed most was the remediation guidance, not the prose. Without method, recommendations were generic and correct-sounding: review the account, check for persistence, monitor. With method, they were specific to what had been found and what had not, and they distinguished between "this was checked and was clean" and "this was not checked".

That distinction — between a clean result and an absent one — later became one of the two or three most important ideas in the system, and it took months to implement properly. It first showed up here, as a difference in the quality of advice.

Experiment four: does specialisation matter?

The supervisor dispatching specialist agents, with and without domain-specific prompts for those specialists. One arm had nine specialists who were told what they were experts in; the other had nine identical generalists with the same tool access.

Specialist depth determined how complete the investigation was, and the effect was strongest exactly where you would hope: the specialists knew which tables to look in, which fields carried the signal, and which negative results were meaningful. A generalist with the same tools asked shallower questions and stopped earlier, because it did not know what it had not yet checked.

ExperimentVariableWhat changed
1SIEM tool accessVerdict accuracy, and the failure mode became debuggable
2 & 3Methodology promptRemediation specificity; checked vs unchecked distinction appears
4Specialist domain promptsInvestigation completeness and depth of questioning

The conclusion that shaped everything after

The prompt architecture is the product. Not the model, which was a commodity choice then and is more of one now. Not the orchestration code, which is thin by design. The thing that determines whether the output is worth reading is how the problem is framed, what knowledge is attached to each agent, and what method they are held to.

And specifically: a supervisor coordinating narrow specialists produces deeper findings than one generalist with the same total context available to it. Same model, same tools, same tokens. Different architecture, different results.

What I would do differently

Build the labelled corpus first. Every one of these findings is directionally right and none of them is measurable, which was fine for deciding an architecture in April and became a real constraint in September, when the question moved from "is this better" to "how much confidence justifies closing a case without a human" — a question that cannot be answered without a set of incidents where you know the right answer.

I did not have one then. I still do not have one now, and it is currently the single largest limitation of the entire system. Five days of lab work in April to establish direction was the right call. Not starting the slow, boring work of building a labelled attack corpus in parallel was not.

Who is writing this

I am Ivan Melekhin. Twenty-five years in cybersecurity, most of the last decade running security operations — building and operating distributed SOC and MSSP teams across Asia-Pacific, with a long detour through OT and maritime environments. This log is the build record for an autonomous SOC investigation agent I started in January 2026, written as the decisions happened rather than tidied up afterwards. I am on LinkedIn if you want to argue with any of it.