Field Notes

Notes from the SOC.

A build log for an autonomous SOC investigation agent — ten months of decisions, pivots, experiments and dead ends, across business, methodology, implementation and lab results. Read it in order, or filter to the track you care about.

The build log

Ten months from idea to working system, in order. 25 chapters.

The security stack nobody investigates

Every Microsoft 365 Business Premium tenant ships enterprise-grade detection, and almost nobody investigates what it produces. The gap is an economics problem, not a tooling one.

Build a business, not a startup: funding R&D on cloud credits

R&D has to be paid for by something. The choice is an investor's money or someone else's infrastructure — what $5,000 in credits actually buys, and the three weeks of support tickets nobody mentions.

Why there is no agent framework in this codebase

No LangGraph, no LangChain, no CrewAI. Thirty lines and five dependencies — the reasoning, the security argument that actually decided it, and the conditions for revisiting.

Hub and spoke: why nine narrow analysts beat one broad one

Same model, same tools, same context budget — a supervisor coordinating nine domain specialists produces deeper investigations than one generalist. And the axis you specialise along turns out to matter more than the specialisation.

The “case backend” was three different problems

Microsoft Sentinel and Defender XDR are telemetry estates separated by a licensing wall. Jira, SharePoint and Teams are communication channels. I spent weeks calling all of them the backend, and they answer to completely different constraints.

We are not the SOC: the pivot from service to engine

Repositioning from an AI-augmented SOC delivering the service to the investigation engine an MSP resells under their own brand — who signs, why deployment is free on purpose, and what the moat actually is.

Data sovereignty as an architecture, not a checkbox

What in-tenant deployment actually means, the two residuals that make the absolute claim false, and why telling a prospect how to audit you is the strongest move an unknown security vendor has.

The dependency the entire unit economics rests on

Near-zero marginal cost per customer is the whole business model, and it is true only if inference runs on the customer's own cloud bill. Three conditions have to hold, and one of them is not fully proven.

Twenty touches, zero conversions: a funnel autopsy

Two proofs of concept that went dark, two ad campaigns with zero leads, a content plan that never published and a permanent ban from the largest MSP community. What zero does and does not tell you.

Everyone is shouting the same sentence

Every vendor in this category claims the best AI SOC, from every angle, constantly. When every message is identical, does architecture still decide anything — or does spend? I do not know, and this chapter does not resolve it.

Reversing a decision I had argued well: Python to Go

In June I wrote a careful assessment recommending we stay in Python and revisit Go at twenty customers. In August I rewrote it in Go. The assessment was not wrong — the weights changed.

A second AI to check the first one's confidence

A model that scores its own confidence after writing its verdict is grading its own homework. The fix looked obvious: an independent verifier scoring belief across competing hypotheses from the evidence log.

The control run: switching it off and watching nothing change

The verifier scored, fused and derived on eleven incidents and routed nothing. The outcomes were indistinguishable from the runs where it decided. Where the two layers disagreed, the supervisor was usually the more cautious one.

Dempster-Shafer, and why we measured before switching calculus

The obvious next move was to swap the evidence maths for something without forced exclusivity. Two measurements over 769 pieces of evidence said keep it — and named the conditions that would change the answer.

Coverage is a ceiling on confidence, and that is arithmetic

A tenant with one sensor cannot produce high-confidence verdicts on multi-domain explanations. Not because analysis is hard — because belief accumulates only through corroboration, and a sensor you do not own contributes ignorance.

The model reasons, the code decides

Thirty incidents routed by a deterministic harness instead of a language model. It disagreed with the model on twelve of them — and in nine, it closed cases the model wanted to escalate to a human.

No posts in this category yet.

Who is writing this

I am Ivan Melekhin. Twenty-five years in cybersecurity, most of the last decade running security operations — building and operating distributed SOC and MSSP teams across Asia-Pacific, with a long detour through OT and maritime environments. This is the build record for an autonomous SOC investigation agent I started in January 2026, written as the decisions happened rather than tidied up afterwards. I am on LinkedIn if you want to argue with any of it.