After the calibration result I did the obvious next experiment and the one I had been avoiding: leave the verifier running, recording everything, and let it route nothing. Eleven incidents. The outcomes were indistinguishable from the runs where it was in charge.
The setup
A single environment flag put the verifier into advisory mode. It still received the full evidence log, still scored belief across hypotheses, still fused, still computed every derived quantity and wrote all of it to the audit record. The only change was that its output no longer influenced routing. The supervisor's own conclusion shipped.
That flag existing at all is the only reason this experiment was possible, and it was not foresight — it was there for debugging. It turned out to be the most valuable line of configuration in the system.
The result
| Run | Verifier routing? | Outcomes (benign / TP / held / asked) |
|---|---|---|
| Earlier run | Yes | 5 / 1 / 3 / — |
| Earlier run | Yes | 5 / 0 / — / 5 |
| Earlier run | Yes | 4 / 0 / — / 5 |
| Control | No — advisory only | 5 / 0 / 0 / 6 |
Nothing distinguishes the control run from the three where the verifier decided. The distribution of outcomes sits inside the ordinary variation between runs.
Looking case by case: the two layers agreed on seven of eleven. Of the four disagreements, the supervisor was the more cautious one on three — asking a question where the verifier would have shipped a verdict. On the single case where the verifier was more cautious, it wanted to hold a correct benign conclusion for human review, which on a tenant with no attacks means it would have sent a correct answer to a person for no reason.
What it cost to keep
A full additional model call on every investigation — the largest single cost line per case — to produce a branch decision that, measurably, changed nothing. Plus the maintenance surface: a second prompt to keep aligned, its own set of open defects, and a scoring layer whose numbers had just been shown to be unreproducible.
The general lesson
Build the off switch first, and use it before you are confident.
A component you cannot disable is a component you cannot evaluate. Every argument for the verifier was structural and sound — a model should not grade its own homework, independent scoring is better than self-assessment — and every one of those arguments survives the measurement intact. The measurement did not say the reasoning was wrong. It said the component was not producing the effect the reasoning predicted, on this system, on this corpus.
You cannot reason your way to that. You can only run it with the thing off.
What happened to the code
Two days later the verifier was switched off in the default build: roughly four thousand lines across fourteen files, moved behind a build tag. Excluded from every normal build, called from nowhere. Still compiles if you ask for it explicitly.
Not deleted, deliberately. The corpus is thirty alerts from a quiet tenant with no attacks in it. It is entirely possible that on a tenant with real intrusions, an independent check on the supervisor's framing earns its cost — the control run says the verifier added nothing here, not that a second opinion is worthless in general. Keeping it compilable means that question can be reopened by flipping a flag rather than by reconstructing a month of work from memory.
The parts that survived are the parts that were never really about the verifier: the hypothesis discipline, the ask gate, the structural caps, the customer context store, and the evidence grouping. What went was the single idea that a model should write the number.