← Back to blog

Dempster-Shafer, and why we measured before switching calculus

Combining evidence from multiple sources into one belief needs a calculus. I use Dempster-Shafer. There is a well-argued alternative without its most questionable assumption, and the tempting move was to switch. Two measurements over 769 pieces of real evidence said don't — and, more usefully, said exactly what would have to change for that to be wrong.

What the calculus has to do

Several specialists investigate. Each returns findings pointing for or against the competing explanations, with a strength. Some rest on the same underlying telemetry, so they are not independent witnesses. Something has to turn that into one decision.

Dempster-Shafer gives three things that matter here. Mass can be assigned to "I don't know" rather than being forced onto a hypothesis, which is essential when your telemetry has holes. Fusing sources produces a conflict measure when they disagree, which is the most stable quantity in the system and drives the ask gate. And belief and plausibility give an interval — how much is committed, and how much is still open — instead of a single number that hides its own ignorance.

The assumption worth questioning

Dempster-Shafer requires hypotheses to be mutually exclusive and collectively exhaustive. With two competing activities, that has a strong consequence: evidence supporting one automatically becomes disbelief in the other, because in a two-member frame the complement of one is the other.

Subjective Logic does not do that. It tracks belief, disbelief and uncertainty separately, adds a base rate for what to assume absent evidence, and never forces support for one thing into opposition to another. Both explanations can be well-supported simultaneously, which is the real situation when an attacker operates inside a legitimate change window.

Run on real data in a two-member frame, the two produce identical numbers on seven of ten per-group rows. Every mismatch is that one automatic transfer.

The two questions that decide it

Rather than arguing from first principles, I measured the two properties that determine whether the automatic transfer helps or hurts, across 769 findings and 176 derivations already on disk.

Are the activity hypotheses actually exclusive in practice? Of 167 derivations with two cleanly identified activities, both scored above 0.5 belief on 18, above 0.6 on 4, and above 0.7 on none. Mean leader 0.678, mean runner-up 0.316. At the specialist level, 29 of 122 findings supported both activities — but only 3 did so citing genuinely separate evidence. In this record, exclusivity holds.

Is disbelief actually being elicited? Zero disbelief on 236 of 333 hypothesis pairs — 71 per cent. On the benign-activity slot, which leads most derivations, 85 per cent. And not because specialists are timid: when they do contradict, the mass is comparable to when they support. Contradiction is used less often, not more weakly.

MeasurementResultConsequence
Exclusivity holds?Yes — both above 0.7 on 0 of 167The automatic transfer is not doing damage
Disbelief elicited?No — absent on 71% of pairsThe automatic transfer is doing most of the work
Base rate estimable?No — one incident per rule on this tenantThe alternative's fallback is unavailable
The chain that settles it. Disbelief is absent on 71% of hypotheses. Under the alternative, those can only be pushed down by a base rate. The base rate is not estimable on this tenant. Therefore the alternative cannot eliminate them at all — it would trade a mechanism that works on 85% of cases for two inputs the system does not have.

The part that generalises

The instinct when a formalism feels wrong is to replace it. The more useful move is to measure which of its properties is currently load-bearing.

Here, the questionable assumption turned out to be the thing carrying the system. Exclusivity was compensating for the fact that specialists rarely state what evidence argues against — so removing it would have exposed a weakness rather than fixing one. The theoretically cleaner choice was the practically worse one, and only measurement could tell me that.

Parked, not closed

The decision is recorded with three named conditions that would reopen it, so this does not have to be re-argued from taste later:

The second one is directly actionable, and the work that moved it is covered later in this log: it went from 71 per cent to 51 per cent, the largest movement on disbelief the project has measured, and still not far enough to fire the trigger.

The honest limit

All of this rests on one tenant, twenty-nine incidents, and exactly one true positive ever recorded. The case where exclusivity fails — an attacker riding along inside a legitimate change window — is precisely the case that has not occurred here.

So the measurement says exclusivity is currently safe, not that it is safe. That is a materially weaker claim than it appears, and the first trigger exists to check it against the first tenant that has real incidents in it.

Who is writing this

I am Ivan Melekhin. Twenty-five years in cybersecurity, most of the last decade running security operations — building and operating distributed SOC and MSSP teams across Asia-Pacific, with a long detour through OT and maritime environments. This log is the build record for an autonomous SOC investigation agent I started in January 2026, written as the decisions happened rather than tidied up afterwards. I am on LinkedIn if you want to argue with any of it.