← Back to blog

The model reasons, the code decides

On the fifteenth of September the language model stopped writing the verdict. A deterministic harness fuses the specialists' evidence and decides the route; the model writes the explanation for a decision it no longer makes. Thirty incidents, and it disagreed with the model on twelve of them.

The division of labour

WhoProduces
Supervisor modelThe competing activity hypotheses, and what each predicts would be observable — written before evidence is queried
Specialist modelsFindings: which hypothesis, which evidence group, supports or contradicts, how strongly, and the tool calls behind it
Harness (code)The route, the verdict class and the confidence figure, by fusion over those findings
Supervisor modelThe prose explaining the decision the harness made

The models still reason. They do not decide. Generating competing explanations and reading a query result into a direction are things they do well and a program cannot do at all. Turning a set of readings into one routing choice is arithmetic, and the calibration run measured what happens when a model is asked to do arithmetic.

Why there is no final model pass

The obvious objection: shouldn't the model get the last word, in case the arithmetic is wrong?

No, and the reason is structural. A final pass that can overrule the harness gives back everything the harness exists for. The decision becomes unreproducible again, and the audit record stops explaining the outcome — it would say what the evidence implied, and then a verdict that does not follow from it. If the harness decides badly, the fix is better inputs, not a model with a veto.

Stopping rules, not confidence adjustments. A case only closes when the fusion is usable and the evidence is decisive. Too few sources carried weight, sources in genuine conflict, too much belief uncommitted, or a harmful explanation leading — each routes to a question or a human. The harness is allowed to decline to decide. It is not allowed to decide quietly on thin evidence.

The run

Thirty incidents. $3.92 total, $0.1305 per case, 129 seconds mean. No errors. Eighteen closed benign, ten asked the customer a question, two held for human review.

This was also the first run where the harness declined rather than the model: one case led toward a harmful verdict at belief 0.308 and was held; another fused a single group at belief 0.000 with total ignorance and was held. Both correct, and neither required a model to notice.

The twelve disagreements

The harness differed from the model's own proposal on twelve of thirty, and eleven of those fall into two shapes.

Nine cases where the model wanted to ask the customer and the arithmetic closed the case benign. All nine had the customer's own previously-recorded facts in the fusion, at belief between 0.625 and 0.833. The model read those facts and still wanted to ask. The harness accumulated them across the investigation and concluded they were enough.

That is the single most interesting result in this project. A senior analyst does exactly this: recognises that the customer already answered this question in a different form last week, and closes rather than asking again. The model, reading everything, kept wanting reassurance. The arithmetic, counting evidence, did not need it. That is the analyst behaviour the whole system set out to encode, and it emerged from counting rather than from prompting.

Three cases going the other way, all toward caution. A harmful explanation leading, held rather than reported. A case with no usable evidence, held. And one the model proposed closing, overridden to a question because the evidence sources were in genuine conflict.

One fix, visible in the numbers

The previous run had a defect where the same clean lookup was counted twice, once by a specialist filing a finding and once by the absence machinery synthesising a contradiction from the same tool calls. It inflated conflict and sent cases to the customer for no reason.

After the fix: mean conflict 0.1115 to 0.069, and cases above the routing threshold from nine to five — while adding a fourth evidence group to most fusions. Belief values also de-quantised, from about six distinct values across 24 cases to about fourteen across 30, spanning the full range. A decision system whose confidence takes six values is not really measuring; that spread is the arithmetic starting to discriminate.

What changed about testing

Reproducibility stopped being a lab report and became a unit test. The same findings must produce the same route, every time, and that is now asserted in code rather than observed in a run.

The models remain non-deterministic and the evidence they gather still varies between runs — that has not been fixed and is listed honestly as unmeasured. But the decision layer no longer adds variance of its own on top, and when a verdict is questioned, the answer is a derivation you can read rather than a number a model wrote.

Who is writing this

I am Ivan Melekhin. Twenty-five years in cybersecurity, most of the last decade running security operations — building and operating distributed SOC and MSSP teams across Asia-Pacific, with a long detour through OT and maritime environments. This log is the build record for an autonomous SOC investigation agent I started in January 2026, written as the decisions happened rather than tidied up afterwards. I am on LinkedIn if you want to argue with any of it.