On the fifteenth of September the language model stopped writing the verdict. A deterministic harness fuses the specialists' evidence and decides the route; the model writes the explanation for a decision it no longer makes. Thirty incidents, and it disagreed with the model on twelve of them.
The division of labour
| Who | Produces |
|---|---|
| Supervisor model | The competing activity hypotheses, and what each predicts would be observable — written before evidence is queried |
| Specialist models | Findings: which hypothesis, which evidence group, supports or contradicts, how strongly, and the tool calls behind it |
| Harness (code) | The route, the verdict class and the confidence figure, by fusion over those findings |
| Supervisor model | The prose explaining the decision the harness made |
The models still reason. They do not decide. Generating competing explanations and reading a query result into a direction are things they do well and a program cannot do at all. Turning a set of readings into one routing choice is arithmetic, and the calibration run measured what happens when a model is asked to do arithmetic.
Why there is no final model pass
The obvious objection: shouldn't the model get the last word, in case the arithmetic is wrong?
No, and the reason is structural. A final pass that can overrule the harness gives back everything the harness exists for. The decision becomes unreproducible again, and the audit record stops explaining the outcome — it would say what the evidence implied, and then a verdict that does not follow from it. If the harness decides badly, the fix is better inputs, not a model with a veto.
The run
Thirty incidents. $3.92 total, $0.1305 per case, 129 seconds mean. No errors. Eighteen closed benign, ten asked the customer a question, two held for human review.
This was also the first run where the harness declined rather than the model: one case led toward a harmful verdict at belief 0.308 and was held; another fused a single group at belief 0.000 with total ignorance and was held. Both correct, and neither required a model to notice.
The twelve disagreements
The harness differed from the model's own proposal on twelve of thirty, and eleven of those fall into two shapes.
Nine cases where the model wanted to ask the customer and the arithmetic closed the case benign. All nine had the customer's own previously-recorded facts in the fusion, at belief between 0.625 and 0.833. The model read those facts and still wanted to ask. The harness accumulated them across the investigation and concluded they were enough.
That is the single most interesting result in this project. A senior analyst does exactly this: recognises that the customer already answered this question in a different form last week, and closes rather than asking again. The model, reading everything, kept wanting reassurance. The arithmetic, counting evidence, did not need it. That is the analyst behaviour the whole system set out to encode, and it emerged from counting rather than from prompting.
Three cases going the other way, all toward caution. A harmful explanation leading, held rather than reported. A case with no usable evidence, held. And one the model proposed closing, overridden to a question because the evidence sources were in genuine conflict.
One fix, visible in the numbers
The previous run had a defect where the same clean lookup was counted twice, once by a specialist filing a finding and once by the absence machinery synthesising a contradiction from the same tool calls. It inflated conflict and sent cases to the customer for no reason.
After the fix: mean conflict 0.1115 to 0.069, and cases above the routing threshold from nine to five — while adding a fourth evidence group to most fusions. Belief values also de-quantised, from about six distinct values across 24 cases to about fourteen across 30, spanning the full range. A decision system whose confidence takes six values is not really measuring; that spread is the arithmetic starting to discriminate.
What changed about testing
Reproducibility stopped being a lab report and became a unit test. The same findings must produce the same route, every time, and that is now asserted in code rather than observed in a run.
The models remain non-deterministic and the evidence they gather still varies between runs — that has not been fixed and is listed honestly as unmeasured. But the decision layer no longer adds variance of its own on top, and when a verdict is questioned, the answer is a derivation you can read rather than a number a model wrote.