On the thirteenth of September I ran the same incidents three times on byte-identical input, changing nothing, to calibrate the thresholds that decide when the agent acts without a human. The run produced a result I was not looking for: most of the quantities being thresholded move more, run to run, than the thresholds are apart.
Why run the same thing three times
The system had roughly a dozen numeric constants — how much belief is needed to close a case benign, how much separation between the leading hypothesis and the next, how much ignorance is tolerable before a case is capped. Every one was a judgement value. I picked them by reasoning about what felt defensible, and they decided real routing.
The plan was calibration: run a fixed incident set repeatedly, look at where the quantities actually land, and set thresholds at defensible points in the observed distribution. Not proper calibration against labelled outcomes — that needs ground truth I do not have — but enough to stop the constants being arbitrary.
Three passes over the same incidents, with the customer fact store checksummed before and after to confirm it did not change between runs. Identical input, three times.
What came back
| Quantity | Derived from | Mean run-to-run range |
|---|---|---|
| Summed belief | Verifier model's stated beliefs | 0.428 |
| Hypothesis separation | Verifier model's stated beliefs | 0.336 |
| Frame ignorance | Verifier model's stated beliefs | 0.296 |
| Interval width | Mixed | 0.223 |
| Fusion conflict | Specialist evidence | 0.086 |
Read the separation row against its thresholds. That quantity had two cut-offs, at 0.15 and 0.30 — fifteen hundredths apart. It moves 0.336 between identical runs. The noise is more than twice the entire span the thresholds were meant to divide.
Frame ignorance moves 0.296 against a ceiling of 0.30. Of the eight cases present in all three runs, the routing branch matched across runs on four, and the confidence band on three.
Where the variance lives
This is the part that made the result actionable rather than merely depressing.
The arithmetic is deterministic and its own reproducibility tests pass — the same inputs produce the same outputs every time. The variance is entirely upstream: it is in what the models write. The verifier reads the same evidence log three times and states materially different beliefs each time. The specialists, running against the same tenant, produce somewhat different evidence.
And the split in that table is not random. Every quantity derived from a model's free-written belief moves by roughly a third of its own range. The one quantity computed from what the specialists actually retrieved — how much the evidence groups disagree with each other — moves four times less, and is the only constant whose threshold sits wider than its own noise.
What this actually said
Not "language models are unreliable", which is too general to act on. Something narrower and more useful: asking a model to turn a body of evidence into a number is asking it to do arithmetic, and it is not doing arithmetic — it is generating a plausible number. Plausible numbers are not stable under repetition, and anything downstream that thresholds them inherits that instability.
Generating competing hypotheses and reading a query result into a direction — those are things models do well and programs cannot do at all. Combining a set of readings into one routing decision is arithmetic, and code does arithmetic reproducibly.
That sentence is the whole design change that followed. The system was routing on the noisier of two available numbers, and the steadier one had been computed and discarded on every single run.
A note on method
This result cost one afternoon and required no new code — just running an existing thing three times and diffing the outputs.
I had been operating the system for weeks without ever doing that. Every run was a new incident set, so every difference in output had an available explanation in the input, and I never saw the variance because I never held the input still. If you are building anything that routes on a model-derived number, run the identical input three times before you tune anything. It is the cheapest experiment available and mine invalidated a month of work.