Every investigation asked: is this alert a true positive, a false positive, or benign? Replaying one case's recorded evidence under a different framing produced a number that stopped me. Conflict between evidence sources: 0.2129 under the old frame, 0.0000 under the new one. Same evidence. The disagreement was manufactured by the question.
The old frame
The set of competing hypotheses was the verdict taxonomy: true positive, false positive, benign. Every specialist finding was scored against those three, and the fusion arithmetic combined them into a distribution and a measure of how much the evidence sources disagreed.
That disagreement measure is load-bearing. High conflict means two sources are pointing opposite ways, which is not something another query fixes — so it routes to asking a human. It is the most stable quantity in the system, as the calibration run showed, and it opens the ask branch.
The replay
I took one real investigation, pulled its recorded evidence out of the audit file, and re-ran the arithmetic on paper under a different frame. The harness reproduced the shipped fusion exactly, digit for digit, so the comparison was against production rather than against a model of it.
Under the shipped frame, conflict was 0.2129 — above threshold, which is why that case had gone to the customer with a rationale reading "the evidence groups disagree, so this is not a gap another query can fill".
Under a frame of competing activities — what actually happened here, rather than what class of verdict this is — the identical bearings produced conflict of exactly 0.0000.
Why the old frame produced phantom conflict
Because the three classes are not three explanations of the world. They are one explanation of the world and two statements about other things entirely.
Across 132 live hypothesis sets, the third slot — the mandatory benign-or-false-positive option — was an evidence gap 74 times and an alert defect 37 times. Only 21 times was it an actual third account of what happened. And "false positive" was selected as the leading hypothesis zero times in 132 reports, with a mean mass of 0.055.
Look at what was being mixed:
| Kind of statement | Example | What it is really about |
|---|---|---|
| Activity | "The admin performed authorised subscription setup" | The world. Belongs in the frame. |
| Alert defect | "The detection misfired; no role assignment occurred" | The rule's reliability. Detection engineering's problem. |
| Evidence gap | "Telemetry is insufficient to tie the change to a sign-in" | Our own visibility. That is ignorance, not a hypothesis. |
Fusing a claim about the world with a claim about our own blind spots is a category error, and the arithmetic faithfully reported the resulting incoherence as conflict between sources. Some sentences were even both at once — "the alert is a false positive caused by rule logic or incomplete telemetry" is two different claims sharing a slot.
The new frame
The investigation asks what activity produced this alert. Not what class of verdict this is.
The hypotheses are competing accounts of what actually happened, each carrying a flag for whether it is harmful. The alert becomes the first piece of evidence rather than the subject of the enquiry — a detection rule is written wide on purpose, catching an attack pattern along with a lot of ordinary work, so the alert's job is to say which activities are worth considering.
Alert defect moves out of the frame and becomes a reliability claim about the rule, owned by whoever maintains detections. Evidence gap moves out and becomes ignorance, which the system already measured separately and correctly.
The customer still sees a true positive, false positive or benign verdict. That is an output taxonomy, joined to the frame by an explicit mapping rather than by sharing a field with it. Leading activity is harmful, so the verdict is a true positive. Nothing survives and the rule is claimed to have misfired, so it is a false positive. Those are different objects that had been conflated into one.
What happened live
Thirty incidents under the new frame. Specialist findings landing on non-activity hypotheses fell to one in 140, from roughly a third of 769 under the old frame — with no prompt iteration, which suggests the models had been straining against the old frame rather than failing at the new one.
It also broke something immediately. Ignorance more than doubled and started capping the confidence band on ninety per cent of cases, because the intended exit for alert-defect claims was now dumping mass somewhere new. One constant, one day to find, one line to fix.
And one prediction I had made from the replay turned out to be wrong. I had generalised from that single case that the new frame would reduce conflict generally. It did not — a two-member frame puts contradictions on single hypotheses, which makes empty intersections more likely, not less. The replay case had bearings that all pointed one way. One case is one case, even when the number it produces is dramatic.
The transferable bit
When an agentic system produces a confusing signal, check the frame before you check the arithmetic. The maths was correct throughout. It was correctly computing the consequences of a question that mixed three kinds of claim, and no amount of tuning would have found that, because tuning assumes the quantity means what you think it means.
The specific version, for anyone building investigation systems: your hypothesis set should contain accounts of what happened. Claims about your instrument's reliability and claims about your own blind spots are real and important and belong somewhere else.