← Back to blog

From alert-forwarder to analyst: what a chatbot could not do

The first version forwarded every alert into a ticket system. The second pointed a language model at the alert and asked it to explain itself. Neither of them automated anything, and understanding exactly why is the single most useful thing that happened in the first two months.

Version zero: plumbing

The first working thing was not intelligent at all. An open-source detection platform running in a container, a custom integration script hooked into its alerting framework, and every alert forwarded into a ticket project — summary line, structured field table, timestamp, rule ID, agent, source, user, and a deep link back into the dashboard at exactly the alert that caused it.

It worked. It was also, on reflection, a pure restatement of the problem. Alerts were now in a place where someone could look at them, which had never been the constraint. Nobody was short of places to look at alerts. What was missing was the hour of analyst attention each one needed, and forwarding does not produce attention. It produces a longer queue with better formatting.

Version one: the chatbot

The obvious next move, and the one nearly everyone makes: put a language model in front of the alert and ask it what it means. This reads impressively. You get fluent, structured, plausible-sounding prose about credential access techniques and recommended next steps, and the first few outputs feel like a solved problem.

Two things were wrong with it, and only one was obvious.

The obvious one: a summary still requires a human to act on it. If the output of your system is a better-explained alert, the person downstream still has to open the console, run the queries, correlate the identity against the sign-in log, decide, and close. You have improved the input to the work. You have not done the work. Augmentation is a legitimate product category and it is not the one that closes the economic gap, because the cost was never in reading the alert.

The non-obvious one took longer to see: the model had nothing to reason about. It was working from the alert text and its own training, which means every confident sentence it produced was a prior, not a finding. It could tell you what this class of alert usually means. It could not tell you what happened in this tenant, because it had never looked.

The unit that matters is the closed investigation. Not the enriched alert, not the summary, not the recommendation. If a human still has to run the queries, nothing structural has changed.

What actually changed it was a document

The thing that moved this forward was not a model upgrade or a prompt technique. It was sitting down and writing out, in tedious detail, what a SOC analyst actually knows.

That became a grade ladder: three analyst grades, ten knowledge topics, twelve security domains, described as capabilities rather than job titles. What does someone at this grade know about identity attacks, about Windows internals, about cloud control planes, about mail flow? What can they do unaided, and what do they escalate?

Once that existed, the mapping onto a system fell out of it almost mechanically: twenty-two knowledge areas resolving onto nine specialist agents — network, Windows, Unix, web, email, container, DevOps, identity, cloud — each with a defined scope and a defined toolset.

VersionWhat it didWhy it was not enough
ForwarderEvery alert into one queue with contextThe queue was never the constraint
ChatbotFluent explanation of the alertNo evidence, no method, human still does the work
Framed specialistsDomain knowledge + method + tool access, separatedThis is the thing that can actually close a case

The insight, stated plainly

An analyst is not "a clever person who reads alerts". An analyst is three separable things: a body of domain knowledge, a method for applying it, and access to the tools that produce evidence. A chatbot has a diffuse version of the first, none of the second, and none of the third.

Separating them is what makes the problem tractable, because each one is independently buildable and independently testable. Domain knowledge goes into specialist prompts. Method — the investigative discipline, the order of operations, the rule that you write your hypotheses before you go looking — goes into the supervisor that frames the case. Tool access is engineering: query builders, schema knowledge, an interface to the telemetry that does not lie about what it found.

That three-way split is the architecture, and it survived everything that came afterwards, including a language rewrite and a complete change to how verdicts get decided. Nearly every later entry in this log is a refinement of one of those three components rather than a challenge to the split itself.

The obvious next question

Writing all this down produces a satisfying diagram and exactly zero evidence. The claim embedded in it — that domain specialisation and explicit method produce materially better investigations than one capable generalist with the same context — is testable, and at that point it was completely untested.

So before building any of it properly, I spent five days testing it. Those results are in the next entry in this track, and they are the reason the architecture looks the way it does.

Who is writing this

I am Ivan Melekhin. Twenty-five years in cybersecurity, most of the last decade running security operations — building and operating distributed SOC and MSSP teams across Asia-Pacific, with a long detour through OT and maritime environments. This log is the build record for an autonomous SOC investigation agent I started in January 2026, written as the decisions happened rather than tidied up afterwards. I am on LinkedIn if you want to argue with any of it.