← Back to blog

A second AI to check the first one's confidence

A model that writes a verdict and then scores its own confidence is grading its own homework. The number arrives after the conclusion, so it justifies rather than measures. The fix looked obvious, I built it over several weeks, and it was the most reasonable-sounding wrong thing in this whole project.

The problem, precisely

The supervisor produced a verdict and a confidence score in the same output. Read the order of operations and the flaw is structural rather than incidental: the model commits to a conclusion, then writes a number describing how sure it is about the thing it has just committed to.

That number cannot be a measurement. It is a post-hoc justification generated by the same process that produced the claim, subject to the same framing, the same context, and the same commitment bias. It is also unauditable — nothing in the system could explain why a given case scored 70 rather than 55, because nothing computed it.

And this number was load-bearing. It decided whether a case closed automatically, went to a human, or triggered a question to the customer. An unauditable number was choosing when the system acted without a person.

The design

Three ideas, each of which I still think is correct.

Analysis of competing hypotheses. Borrowed from intelligence analysis, and the discipline it imposes is simple: write down the full set of candidate explanations before you go looking for evidence, then judge every piece of evidence against every hypothesis rather than against your favourite. Choose the next query for what it would discriminate between. Eliminate hypotheses; do not crown one. Writing hypotheses first is what stops an investigation becoming a search for confirmation of whatever the alert suggested.

An independent verifier. A second model instance that never sees the supervisor's verdict. It receives the hypothesis set and the complete tool-call log — every query, every result — and independently scores belief across the hypotheses. Separating the scorer from the concluder was the whole point.

Deterministic code on top. The verifier's scores go into code, not more prose. Normalise, compute the separation between the leading hypothesis and the runner-up, measure how much belief sits on "none of these", and apply structural caps — rules that force human review when specific conditions hold regardless of the score.

RouteTaken when
ShipOne hypothesis clearly leads and the evidence is decisive
Re-queryThe gap is something another query could close
Ask the customerThe missing fact is human-held, and only after evidence is exhausted
Re-synthesiseSupervisor and verifier disagree — one round, both versions kept
Human reviewA structural cap fires, or nothing else applies

The ask gate

One sub-problem deserves its own note, because the fix generalises.

The agent had a habit I came to think of as alert-proxy behaviour: when uncertain, ask the customer. That feels safe and is corrosive. The value proposition is that the customer stops having to think about alerts; an agent that forwards its uncertainty to them has rebuilt the original problem with extra steps and a friendlier tone.

So asking became a gated action rather than an available one. Evidence must be exhausted first — if a query could answer it, run the query. The question must be about something genuinely human-held, the kind of fact no log contains: was this expected, did you authorise this, is this contractor supposed to have that access. And never, under any circumstance, ask the identity currently under investigation.

The rule that survived everything. A question to a human is an admission that the evidence could not settle it. That makes the ask rate a symptom worth reading, not a number worth steering — which is why it was deliberately given no target.

Why no target

There was an open question in the design about what the right ask rate should be. It was closed by deciding not to answer it.

A target rate is a false optimisation goal. Give the system a number to hit and the number becomes reachable in ways unrelated to investigating: asking less on cases that deserve a question, or more to pad a quota. Both look like success against the metric and are failures at the work. The ask rate is a consequence of how much the evidence closed, so it is recorded in every lab report and read as a symptom. A jump gets explained by finding which rule fired more often, never corrected toward a line.

What happened next

All of the above is sound, and most of it is still in the system today.

The verifier is not. In September I measured it three ways — a calibration run, a control run, and a replay — and the measurements did not support it. It was switched off on the fifteenth, after weeks of work, and what replaced it keeps the hypotheses, the ask gate and the structural caps while throwing away the part where a model writes the number.

That sequence is the most useful thing in this log, and it only exists because the component was built in a way that could be switched off and measured against.

Who is writing this

I am Ivan Melekhin. Twenty-five years in cybersecurity, most of the last decade running security operations — building and operating distributed SOC and MSSP teams across Asia-Pacific, with a long detour through OT and maritime environments. This log is the build record for an autonomous SOC investigation agent I started in January 2026, written as the decisions happened rather than tidied up afterwards. I am on LinkedIn if you want to argue with any of it.