I had assumed trust was one problem and that the architecture solved it. Writing up the failed quarter forced a more precise view: it is three distinct problems, the architecture answers one of them, and that one is the cheapest to claim and the least persuasive on its own.
Layer one: will my customer's data leak?
This one the architecture genuinely answers. The agent runs inside the customer's own tenant, permissions are overwhelmingly read-only, and the write surface is limited to incident comments and chat messages. Nothing can be modified, disabled, or deleted by the agent, because it holds no permission to do any of those things.
Two residuals have to be volunteered rather than waited for: inference, and the outbound threat intelligence flow. Volunteering them is not humility, it is self-interest — a prospect who finds an unstated exception discounts everything else you claimed.
Layer two: who are you, and what is that binary doing?
Here the architecture is not just silent, it actively works against me.
Read-only access sounds reassuring until you say it to a security buyer, who hears something else: read access to the entire security log estate of every client we manage. That is not a small blast radius. That is precisely the blast radius they worry about, because it is everything an attacker would want to read before doing anything.
And the container is deliberately opaque. The prompts are the product, so they are delivered at runtime and never sit at rest on disk. Every intellectual-property control that protects the business makes the thing harder to inspect, which is exactly the wrong direction for a buyer asking what it does with their data.
No permission model fixes this, because the question is not about permissions. It is "why should I believe you". Only three things convert that: a reference from someone they already trust, verifiability they can exercise themselves, and a visible credible human attached to the thing. I had none of the three.
| Layer | The buyer's question | Does architecture answer it? |
|---|---|---|
| 1. Data safety | Will my client's data leak or be damaged? | Mostly yes — with two residuals stated up front |
| 2. Vendor | Who are you, and what is this opaque binary doing? | No — and IP protection makes it worse |
| 3. Judgment | Do I trust its verdicts enough to close a case? | No, and it cannot — only observed performance does |
Layer three: do I trust the verdict?
This is the one that cannot be argued at all, by anyone, ever.
The service provider resells the output. When the agent closes an alert as benign, that decision is delivered to their client under their brand. If it was wrong, the breach conversation is theirs — not mine, not the model vendor's. They are being asked to stake a client relationship on the judgement of a system they have watched work for zero hours.
No amount of architecture addresses this. No permission model, no audit trail, no explanation quality. The only thing that converts judgment trust is observed performance over time, which means the first customer has to take a risk that cannot be argued away.
The one mechanism that actually addresses layer three
Shadow mode, and it should have been the default entry from day one.
The agent investigates every alert and posts its verdict to a chat channel. It closes nothing. It changes nothing. For two to four weeks the provider compares what the agent concluded against what their own people concluded, on their own tenant, with their own alerts.
That converts the ask from "risk your client relationship on an unproven system" to "read a channel for a month". It generates exactly the evidence layer three requires, on their data rather than my lab's. And graduating from shadow mode to closing verdicts is a natural moment to start charging, because by then the value is a thing they have watched rather than a thing I have claimed.
It costs the provider almost nothing and it is the only mechanism I have found that builds judgment trust without requiring someone to go first on faith.
What layer two actually needs
Three things, in rough order of cost. Publish the verifiability material — exact permissions with per-permission justification, exact egress endpoints, and instructions for watching the container's outbound traffic from their own side. Publish real artefacts: redacted reports from actual investigations, real cost and latency numbers, a recording of an attack chain being investigated end to end. And attach a visible human with a track record to the whole thing.
The third one is why this log exists, and why it is signed rather than published under a brand. A company page saying "we built the best AI SOC" is the least distinguishable sentence in the category. A named practitioner writing down what was measured, including what failed, is at least a different kind of object.