The production model backend has moved twice. Neither move was decided by a benchmark, a leaderboard, or a quality comparison. Both were decided by where the inference was allowed to happen — and one of them was blocked for weeks by a deployment SKU rather than anything technical.
Where it started
Everything was built against one model vendor's own API endpoint, using their Sonnet 4.6 model. That is the shortest path from idea to working agent: best documentation, most mature SDK, first-class tool use, prompt caching that works. If you are trying to find out whether an idea is viable, this is the correct choice and I would make it again.
It was also, in hindsight, a decision with a business constraint hidden inside it.
What ended it
Data sovereignty. The product's central claim is that customer data stays inside the customer's tenant. Every other component honoured that. Inference did not: the investigation content — including text derived from the customer's own logs — went to an endpoint I controlled, on my key.
That is the one hole in an otherwise clean claim, and it is exactly the hole a diligent security buyer finds. It is also worth understanding precisely: placing a model resource in a particular cloud region does not necessarily mean inference is processed there. Some routing arrangements send the prompt to the model vendor's own infrastructure regardless of where you provisioned the resource. Reading the routing documentation carefully, rather than assuming the region selector means what it appears to mean, was one of the more consequential hours of this project.
The move to the cloud vendor's own models
The residency-compliant option was to run inference on models hosted by the same cloud provider the customer is already using, provisioned inside their own subscription. Same trust boundary as the rest of the deployment, same billing relationship, same audit surface. The customer can see the resource, because it is theirs.
This also has a large commercial consequence, covered in its own chapter: the inference cost moves onto the customer's bill, which is the difference between a per-tenant cost of a hundred-odd dollars a month and a per-tenant cost of nearly nothing.
The blocker nobody warns you about
The target model, GPT-5.4, turned out to be provisionable in that subscription only as a global batch deployment — an asynchronous job queue, with no synchronous endpoint.
For most workloads that is a cost optimisation. For an agent it is a wall. A tool-use loop is a conversation: the model asks for a query, you run it, you hand back the result, it asks for the next one based on what it just saw. Each turn depends on the last. You cannot run that over a job queue with no interactive endpoint — not slowly, not with worse latency, but structurally, because the shape of the interaction does not fit the shape of the deployment.
The fix was a quota request and a wait. Nothing technical, nothing clever: submit the request, wait for synchronous capacity to be granted, then proceed. Worth knowing if you are planning an agent on a specific hosted model — check that a synchronous deployment of that exact model is available in your subscription and region before you design around it, because model availability and model deployment shape availability are different things and only one of them is on the marketing page.
| Stage | Model | Why |
|---|---|---|
| Build | Sonnet 4.6, vendor's own endpoint | Fastest path to a working loop; best SDK and tool use |
| Residency | GPT-5.4, cloud-hosted in tenant | Inference inside the customer's own boundary; cost moves to their bill |
| Current | GPT-5.6-terra | Successor on the same hosting path |
| Development | Still the original vendor | Better iteration experience; no customer data involved |
The honest quality comparison
Offered as what it is: a practitioner's subjective read from working with all three on the same investigations, not a benchmark, with no scoring rubric and no labelled corpus behind it.
Sonnet 4.6 was better than GPT-5.4 at this task — noticeably so on multi-step reasoning over evidence, and on the level of specific detail that survived into the final write-up. The gap was not enormous but it was consistent enough that I noticed it without looking for it. GPT-5.6-terra is better than 5.4, which closes part of that gap.
And the reason the production backend is the one that was not my first choice on quality is that quality was not the deciding variable. A slightly better model on infrastructure I cannot let the customer's data reach is worth less than a slightly worse model inside their own boundary. That is an uncomfortable trade to write down, and it is the right one for this product.
What made the moves cheap
A thin provider adapter: an interface, one implementation per vendor, a factory, and callers that know nothing about which one is active. A few files, written in an afternoon, which is roughly the scope this abstraction deserves.
Note the difference between that and adopting a framework for the same benefit. The swappable-backend argument is the strongest one frameworks make, and it is real — but the version of it you actually need is about a hundred lines, and writing those hundred lines does not bring thirty packages and someone else's upgrade cadence with it.
The transferable rule: do not pick a model. Pick a swap cost, keep it low, and expect the thing that eventually forces the swap to be a contract, a region, or a deployment SKU rather than a capability.