The system works. It investigates thirty incidents for four dollars, decides from evidence you can trace, and explains itself. This final chapter is about everything it still cannot tell you — including the discovery that the scenario the entire product was imagined around cannot be tested on the tenant it was built in.
The thresholds are chosen, not measured
Two numbers decide when the agent may act without a person: the belief required to close something benign, and the belief required to report a true positive. Both are judgement values. Neither is fitted to data.
They cannot be fitted yet. Fitting requires a body of alerts with known right answers, and the only corpus available is thirty alerts that the SIEM raised while a tenant was being set up. There are no attacks in it. With no errors to minimise, there is nothing to fit against.
On the last run, four of the nine cases the agent closed sat between 0.625 and 0.651 against a 0.60 line. Move the line to 0.70 and five of eighteen closes become questions. The routing behaves sensibly; the line's exact position is a preference, and it sits in configuration for exactly that reason.
The tenant cannot produce the scenario
This is the finding that reframed everything else, and it came from an offline study that was asking a completely different question.
I was testing whether hypotheses that name a specific mechanism discriminate better than a generic authorised-versus-unauthorised pair. The answer was no — and the reason was the lab:
| Telemetry | In the lab tenant |
|---|---|
| Interactive sign-in records, 30 days | Zero rows |
| Endpoint telemetry | None |
| Mail telemetry | None |
| Service principal sign-ins | 94,729 rows |
| Directory audit records | 118 rows |
Every table with data in it is service-principal, non-interactive, managed-identity or control-plane activity. Two of the seven declared evidence groups have produced zero findings across 769.
The first true positive was wrong
The first harmful verdict the system ever shipped was a false positive. It closed against two standing customer facts that said the opposite, and the rule written in response — customer context contradicting the leading explanation blocks the close — is still live.
Since then: zero true positives, on a corpus with no attacks in it. The harmful-verdict path is barely exercised, which is why the bar for reporting one is set higher than the bar for closing one.
Ten ways this can drift
Not hypothetical. Several are visible in current data.
- Benign drift. Every alert in the corpus is benign, the customer confirms they are benign, and the thresholds were chosen while watching that corpus. Nothing pushes back toward suspicion. An agent tuned on a quiet tenant learns that quiet is correct.
- Testimony dominance. On a thin tenant the customer may be one of only two or three sources with anything to say. The less telemetry a customer has, the more the agent leans on asking them — exactly backwards from where you would want the reliance.
- Onboarding camouflage. Tenant setup produces a burst of privileged activity by one actor from one location. An attacker in that window looks identical. Neither the agent nor a human can separate them — but a human feels the risk and the agent does not.
- Absence inflation. "We looked and found nothing" argues against the harmful explanation. A narrow query that runs clean is indistinguishable from a thorough one that runs clean.
- Framing bias. The supervisor writes both the harmful and the benign explanation. If it writes a weak harmful one, the evidence refutes it easily and benign wins on a technicality. Nothing audits the quality of the opposition.
- Indistinguishable explanations. If both hypotheses predict the same observable, no evidence separates them and everything lands on "unsettled". Hypotheses are now required to name different queries — but that requirement is instructed, not enforced.
- Threshold cliffs. Belief values cluster. A small threshold change does not move one case, it moves five or six at once.
- Selection bias upstream. The maths is sound about the evidence it receives. What evidence exists is decided by a model choosing which queries to run. A systematic blind spot there is invisible to everything downstream and looks like confident agreement.
- Source-count asymmetry. More sources agreeing raises confidence. How many sources exist is a property of the customer's licensing, not of the truth.
- The conflict-to-ask-to-close loop. Sources disagree, the agent asks the customer, the customer says it was expected, confidence rises, the case closes. Correct most of the time, and precisely the path an attacker with access to the answering account would walk. The only brake today is that a harmful leading explanation never goes to the customer at all.
What else is not known
Reproducibility of the evidence gathering is unmeasured — the decision code is deterministic by construction, but whether specialists gather the same evidence across repeated runs has not been tested. A failed query is not retried within an investigation, so a transient failure becomes a permanent blind spot for that alert. There is no per-asset criticality, so a domain controller and a test VM need identical belief to close. Evidence is not carried between related incidents, so a customer confirming one alert does not inform the other eight from the same actor in the same window.
What would have to be true
In priority order, and none of them is code:
| # | Needed | Unblocks |
|---|---|---|
| 1 | A corpus with real or simulated attacks | Every unmeasured threshold. Nothing else unblocks calibration. |
| 2 | A tenant with endpoint and mail telemetry | Two of seven evidence groups stop being theoretical |
| 3 | Repeated runs over a fixed incident set | Whether evidence gathering is as stable as the decision code |
| 4 | Analyst-confirmed outcomes on closed incidents | Any scoring at all. Without a record of what was right, it cannot improve. |
Why publish this
Because the honest version is more useful than the confident one, and because the trust problem has no other solution. A vendor listing their own system's ten drift modes is doing something a marketing page structurally cannot do.
And because item two is a request. The limitations above are not solved by more engineering. They are solved by tenants with real telemetry and real incidents, which means service providers willing to run this in shadow mode — investigating and posting verdicts, closing nothing — and compare it against their own conclusions for a few weeks. That costs them reading a chat channel. It is the only thing that converts any of the unknowns above into knowns.
That is the whole ask, and it is the honest end of a ten-month build log. The decision is made by code you can read, from evidence you can trace to the query that produced it, and every verdict carries a line naming the numbers behind it. What it has not yet done is meet a real attack.