← Back to blog

The corpus is the constraint: what this system still cannot tell you

The system works. It investigates thirty incidents for four dollars, decides from evidence you can trace, and explains itself. This final chapter is about everything it still cannot tell you — including the discovery that the scenario the entire product was imagined around cannot be tested on the tenant it was built in.

The thresholds are chosen, not measured

Two numbers decide when the agent may act without a person: the belief required to close something benign, and the belief required to report a true positive. Both are judgement values. Neither is fitted to data.

They cannot be fitted yet. Fitting requires a body of alerts with known right answers, and the only corpus available is thirty alerts that the SIEM raised while a tenant was being set up. There are no attacks in it. With no errors to minimise, there is nothing to fit against.

On the last run, four of the nine cases the agent closed sat between 0.625 and 0.651 against a 0.60 line. Move the line to 0.70 and five of eighteen closes become questions. The routing behaves sensibly; the line's exact position is a preference, and it sits in configuration for exactly that reason.

The tenant cannot produce the scenario

This is the finding that reframed everything else, and it came from an offline study that was asking a completely different question.

I was testing whether hypotheses that name a specific mechanism discriminate better than a generic authorised-versus-unauthorised pair. The answer was no — and the reason was the lab:

TelemetryIn the lab tenant
Interactive sign-in records, 30 daysZero rows
Endpoint telemetryNone
Mail telemetryNone
Service principal sign-ins94,729 rows
Directory audit records118 rows

Every table with data in it is service-principal, non-interactive, managed-identity or control-plane activity. Two of the seven declared evidence groups have produced zero findings across 769.

The motivating scenario cannot be tested here. "A user signs in from an unusual country" — the example that shaped the product — requires interactive sign-in telemetry the lab tenant does not have. Every hypothesis the system generates crowds onto the identity domain because that is the only place with anything in it.

The first true positive was wrong

The first harmful verdict the system ever shipped was a false positive. It closed against two standing customer facts that said the opposite, and the rule written in response — customer context contradicting the leading explanation blocks the close — is still live.

Since then: zero true positives, on a corpus with no attacks in it. The harmful-verdict path is barely exercised, which is why the bar for reporting one is set higher than the bar for closing one.

Ten ways this can drift

Not hypothetical. Several are visible in current data.

What else is not known

Reproducibility of the evidence gathering is unmeasured — the decision code is deterministic by construction, but whether specialists gather the same evidence across repeated runs has not been tested. A failed query is not retried within an investigation, so a transient failure becomes a permanent blind spot for that alert. There is no per-asset criticality, so a domain controller and a test VM need identical belief to close. Evidence is not carried between related incidents, so a customer confirming one alert does not inform the other eight from the same actor in the same window.

What would have to be true

In priority order, and none of them is code:

#NeededUnblocks
1A corpus with real or simulated attacksEvery unmeasured threshold. Nothing else unblocks calibration.
2A tenant with endpoint and mail telemetryTwo of seven evidence groups stop being theoretical
3Repeated runs over a fixed incident setWhether evidence gathering is as stable as the decision code
4Analyst-confirmed outcomes on closed incidentsAny scoring at all. Without a record of what was right, it cannot improve.

Why publish this

Because the honest version is more useful than the confident one, and because the trust problem has no other solution. A vendor listing their own system's ten drift modes is doing something a marketing page structurally cannot do.

And because item two is a request. The limitations above are not solved by more engineering. They are solved by tenants with real telemetry and real incidents, which means service providers willing to run this in shadow mode — investigating and posting verdicts, closing nothing — and compare it against their own conclusions for a few weeks. That costs them reading a chat channel. It is the only thing that converts any of the unknowns above into knowns.

That is the whole ask, and it is the honest end of a ten-month build log. The decision is made by code you can read, from evidence you can trace to the query that produced it, and every verdict carries a line naming the numbers behind it. What it has not yet done is meet a real attack.

Who is writing this

I am Ivan Melekhin. Twenty-five years in cybersecurity, most of the last decade running security operations — building and operating distributed SOC and MSSP teams across Asia-Pacific, with a long detour through OT and maritime environments. This log is the build record for an autonomous SOC investigation agent I started in January 2026, written as the decisions happened rather than tidied up afterwards. I am on LinkedIn if you want to argue with any of it.