← Back to blog

The security stack nobody investigates

Every Microsoft 365 Business Premium tenant already owns an enterprise-grade detection stack. Endpoint detection, identity protection, email security — the customer is paying for all of it, and it fires around the clock. Almost none of it is ever investigated. That gap is an economics problem, not a tooling one, and it is the reason this project exists.

This is the first entry in a build log covering ten months, from an idea in January to a working system in September. It is written in four tracks — business and go-to-market, methodology, implementation, and experiments — because those are the four places the decisions actually happened, and they did not happen in a tidy sequence. Everything is published at once, so you can read it in order or jump to the track you care about.

Start with the problem, because it is the only part of this story that has not changed.

The signals fire. Nobody reads them.

A 50-person company on Business Premium has Defender for Business on the endpoints, Entra ID P1 on identity, and Defender for Office on mail. That is a real detection estate. It generates alerts continuously. And in the overwhelming majority of those tenants, an alert's entire lifecycle is: fire, sit in a portal nobody has open, age out.

The instinct is to call this a tooling failure. It is not. The tools work. What is missing is the thing between an alert and a decision, and that thing has always been a person.

Why the gap persists: the arithmetic

A minimum viable analyst team — tier one, tier two, someone senior, with enough coverage to matter — costs somewhere between $400,000 and $700,000 a year before you buy a single tool. Spread that across clients and the floor for a credible managed detection engagement lands around $25,000 per client per year.

Below that number, the arithmetic does not close. This is not a secret, and it is not for lack of trying: essentially every managed security provider has attempted to move downmarket, and the ones that got there did it by cutting what the investigation actually involves until the service was a dashboard and a monthly PDF.

LayerAnnual costWhat it means downmarket
Analyst team$400–700k before toolingFixed, and it does not scale down gracefully
Credible engagement floor~$25k per clientRoughly 10x what a 50-seat SMB will pay for security
What the SMB already paysBusiness Premium licensingThe detection is bought; the investigation is not

The second problem: the output has nowhere to land

Even where someone does investigate, the shape of the output is wrong. Enterprise security operations assume a security team on the receiving end — a queue, a triage process, people who read "possible AiTM phishing, medium confidence, recommend session revocation" and know what to do next.

A 50-person business has no CISO and no queue. The recipient is the business owner, or the three-person IT shop that manages their tenant. Send enterprise-style analyst output to those recipients and you have not delivered security operations. You have delivered a document nobody reads.

Two constraints, not one. The cost of investigation puts it out of reach, and the format of investigation makes it useless when it arrives. Solving only the first one produces a cheaper report that still goes unread.

Who actually feels this

The sharpest version of the pain sits with small managed service providers — three to ten people, generalist IT, running Business Premium tenants for a portfolio of small clients. They have no security practice and no realistic path to staffing one. They carry the reputational risk when a client gets breached. They have nothing to show at insurance renewal. And they are sitting on top of a detection stack that is already generating the signals, already licensed, already theirs to use.

That is an unusual shape for a market gap: the capability is present and paid for, and the only missing component is the labour to interpret it.

Why it is worth attempting now

The honest answer is that the constraint moved. Structured investigation reasoning — working through evidence, holding competing explanations, deciding what to query next based on what would discriminate between them — is now something a frontier model can do at a useful standard. Not perfectly, and, as later entries in this log get into at length, not without a great deal of scaffolding that turned out to matter more than the model choice. But the binding constraint of the last decade is technically breakable in a way it was not three years ago.

What remains is not intelligence. It is infrastructure: specialist agents that know their domains, a real interface to the telemetry, a method the reasoning follows, and output written for someone who has never opened a SIEM. Most of this log is about building those four things, getting several of them wrong, and measuring which parts were load-bearing.

What this log is, and is not

It is not a product pitch. Where numbers appear they come from lab runs on a single test tenant, and where the results are unflattering — a first true positive that turned out to be wrong, a component built over weeks and then switched off, a go-to-market that produced zero conversions — those entries exist too, at the same length as the wins.

The most useful thing I can offer anyone considering this problem is not the conclusion. It is the sequence of decisions, and which ones survived contact with measurement.

Who is writing this

I am Ivan Melekhin. Twenty-five years in cybersecurity, most of the last decade running security operations — building and operating distributed SOC and MSSP teams across Asia-Pacific, with a long detour through OT and maritime environments. This log is the build record for an autonomous SOC investigation agent I started in January 2026, written as the decisions happened rather than tidied up afterwards. I am on LinkedIn if you want to argue with any of it.