The security stack nobody investigates
Every Microsoft 365 Business Premium tenant ships enterprise-grade detection, and almost nobody investigates what it produces. The gap is an economics problem, not a tooling one.
A build log for an autonomous SOC investigation agent — ten months of decisions, pivots, experiments and dead ends, across business, methodology, implementation and lab results. Read it in order, or filter to the track you care about.
Ten months from idea to working system, in order. 25 chapters.
Every Microsoft 365 Business Premium tenant ships enterprise-grade detection, and almost nobody investigates what it produces. The gap is an economics problem, not a tooling one.
R&D has to be paid for by something. The choice is an investor's money or someone else's infrastructure — what $5,000 in credits actually buys, and the three weeks of support tickets nobody mentions.
The first build forwarded alerts into a ticket system. The second summarised them with a language model. Neither automated anything — and the reason why decided the architecture.
Controlled lab runs testing whether tool access, methodology prompts and domain specialisation actually change investigation quality — plus the limitation I should have fixed in April.
No LangGraph, no LangChain, no CrewAI. Thirty lines and five dependencies — the reasoning, the security argument that actually decided it, and the conditions for revisiting.
Same model, same tools, same context budget — a supervisor coordinating nine domain specialists produces deeper investigations than one generalist. And the axis you specialise along turns out to matter more than the specialisation.
The original design split the agent, the SIEM toolset and the internet toolset into separate containers talking over HTTP. Collapsing them into one process removed an attack surface instead of adding one.
Microsoft Sentinel and Defender XDR are telemetry estates separated by a licensing wall. Jira, SharePoint and Teams are communication channels. I spent weeks calling all of them the backend, and they answer to completely different constraints.
Repositioning from an AI-augmented SOC delivering the service to the investigation engine an MSP resells under their own brand — who signs, why deployment is free on purpose, and what the moat actually is.
What in-tenant deployment actually means, the two residuals that make the absolute claim false, and why telling a prospect how to audit you is the strongest move an unknown security vendor has.
From a vendor's own endpoint to cloud-hosted GPT-5.4 to 5.6-terra. Data residency drove every migration, a batch-only SKU blocked one of them outright, and here is the honest quality comparison.
Near-zero marginal cost per customer is the whole business model, and it is true only if inference runs on the customer's own cloud bill. Three conditions have to hold, and one of them is not fully proven.
Two proofs of concept that went dark, two ad campaigns with zero leads, a content plan that never published and a permanent ban from the largest MSP community. What zero does and does not tell you.
Selling security to a service provider requires trust in data safety, in the vendor, and in the verdicts. Architecture answers the first one — which is also the cheapest one to claim.
Every vendor in this category claims the best AI SOC, from every angle, constantly. When every message is identical, does architecture still decide anything — or does spend? I do not know, and this chapter does not resolve it.
In June I wrote a careful assessment recommending we stay in Python and revisit Go at twenty customers. In August I rewrote it in Go. The assessment was not wrong — the weights changed.
The prompts are the product, and the product ships inside customer-controlled infrastructure. Runtime delivery, a leaked container image, and an honest assessment of what obfuscation actually buys.
A model that scores its own confidence after writing its verdict is grading its own homework. The fix looked obvious: an independent verifier scoring belief across competing hypotheses from the evidence log.
The same incidents, three times, on byte-identical input. Every quantity derived from a model's stated belief moved further between runs than the thresholds reading it. One quantity did not.
The verifier scored, fused and derived on eleven incidents and routed nothing. The outcomes were indistinguishable from the runs where it decided. Where the two layers disagreed, the supervisor was usually the more cautious one.
The same recorded evidence produced conflict of 0.21 when framed as true positive versus false positive versus benign, and 0.00 when framed as competing activities. The disagreement was manufactured by the frame.
The obvious next move was to swap the evidence maths for something without forced exclusivity. Two measurements over 769 pieces of evidence said keep it — and named the conditions that would change the answer.
A tenant with one sensor cannot produce high-confidence verdicts on multi-domain explanations. Not because analysis is hard — because belief accumulates only through corroboration, and a sensor you do not own contributes ignorance.
Thirty incidents routed by a deterministic harness instead of a language model. It disagreed with the model on twelve of them — and in nine, it closed cases the model wanted to escalate to a human.
Two thresholds decide when the agent acts alone and neither is fitted to data. The lab tenant has no attacks, no endpoint telemetry and zero interactive sign-ins. Here is every limitation, and the ten ways it can drift.
No posts in this category yet.