org.tech/Compare/Self-hosted modelsGlass box.The strongest alternative
This is the option we argued with hardest internally, and the only one where the weights themselves sit inside the perimeter you control.
Weights {inside} the perimeter, and a team to keep them running.
If you self-host an open-weight model, no request ever leaves your account, not even to a managed service. That is a stronger statement than anything we can make. What follows is why we still chose a managed route, and the conditions under which you should not.
Our own claim has a boundary in it: data at rest stays in your account, and a managed model service still processes the content of a request. Self-hosting removes that sentence entirely. For a small number of organizations that is decisive, and no amount of private networking is a substitute for it.
01
Nothing leaves, at all.
The weights, the request and the answer are all inside infrastructure you operate. There is no provider term to read, no retention election to make, no abuse-monitoring process to disclose to a reviewer. The data-flow diagram is one box.
One boxNo third party in the path
02
No per-token meter.
You pay for capacity rather than for calls. At high, steady, predictable volume that arithmetic can turn firmly in your favour, and it stops being a question of governance and becomes a question of utilisation.
Capacity, not callsUtilisation decides
03
Nothing can be withdrawn.
A model you host cannot be deprecated out from under you, repriced, or made unavailable in your region. You hold the artifact. In a market that moves as fast as this one, that is a genuine form of continuity.
No deprecationNo repricing
Why we did not build on it.
OUR DECISION · 02
An architecture decision, recorded
Three reasons, none of them about ideology.
Self-hosting behind the same governance layer was the strongest alternative we considered, and the decision against it was made on operating cost rather than on principle. It has not been reversed, and the conditions that would reverse it are written down.
01
It creates a standing function. Provisioning accelerators, patching the serving stack and optimizing inference is not a project. It is a rota, and a small practice should not be staffing one per client.
02
Quality at hostable sizes still trails. On synthesis, the models most organizations can realistically host do not match the current frontier tier. That gap narrows, and when it closes for a given workload the calculation changes.
03
Hardware has a shelf life. Accelerators bought for a three-year horizon are competing with a market that reprices every few months. Somebody carries that risk, and on a managed route it is not you.
Solid zone: what you own. Red zone: what you must staff. The retrieval, identity and policy work is identical on both routes, which is the part people underestimate on either side of this argument.
Self-hosted topology·The left box is the appeal, the right box is the cost
When it is right.
THE CONDITIONS · 03
Three conditions make self-hosting the correct answer rather than the romantic one. If any of them holds for you, we will say so in the assessment and we are not the supplier for that workload. We would rather lose the engagement than deploy something that cannot meet the requirement.
01
An air-gapped requirement.
If the environment genuinely has no route to a managed service, the argument is over before it starts. No private endpoint helps, because the requirement is not about privacy of the path. It is about the absence of a path.
No route outNo managed option
02
Residency with no compliant route.
If your obligation pins inference to a jurisdiction and the models you need have no route inside it, self-hosting is the only way to have both. Our answer in that case is a refusal with a reason rather than a quiet reroute, which is still a no.
Hard residencyWe refuse, in words
03
Very high, steady volume.
Per-token pricing is generous at moderate volume and unkind at scale. If your workload is large, constant and narrow enough for a smaller model to serve well, hosted capacity can win on cost alone. Do the arithmetic before the architecture.
Constant loadUtilisation carries it
Open weights are in the catalogue too.
BOTH, ACTUALLY · 04
The choice is often framed as open models against closed ones. On a managed route it is not: the catalogue generated from the provider's live list on 5 September 2026 held 104 active models from 18 labs, and the open-weight families are in it alongside the frontier tiers. Choosing a managed service does not mean choosing a single vendor's model.
A
What is available in Canada.
Of those 104 models, 13 run in-region in Canada, 1 is routed within Canada only, 18 are reachable but processed outside Canada, and 72 are served from US and global regions. The in-Canada generation set is previous-generation: Claude 3 Sonnet and Haiku, Llama 3 8B and 70B, Mistral Large 24.02, Mixtral 8x7B and Mistral 7B. Open weights are most of that list.
Generated 5 Sep 2026Never retyped into copy
B
What that costs you.
Every current frontier model routes outside Canada. So a Canadian residency pin, available on the Strict tier, is a real choice with a real capability cost, and we will not tell you it is free. The second inference plane offered alongside it, with an alternative API, has no Canadian region at all, so a Canada-pinned deployment cannot serve the models on it.
Strict tierResidency is a trade
What "composable" means here, precisely
Switching the generation model behind the application layer is a configuration change and an evaluation, not a migration: your data, your index, your prompts and your audit trail do not move. The honest exception is the embedding model, where a change forces a re-index. And composable refers to models, not clouds. Moving the platform to a different cloud provider would be a rebuild. The full statement is here.
Fine-tuning is not an answer to a knowledge problem.
NON-ANSWERS · 05
Both self-hosters and managed-route buyers reach for the same two shortcuts when the question is "how does it know our material". Both are good techniques used for the wrong job here, and it is worth being precise about why.
Approach
What it is genuinely good at
Why it does not answer this
Fine-tuning on your corpus
Style, format, tone, narrow task behaviour that is stable over time
✕Bakes knowledge into weights that are stale the day a policy changes, and cannot cite the document an answer came from
A very long context window
Reasoning across one large document, or a working set somebody has already chosen
✕An organization's corpus exceeds any window, and the cost of every request scales with what you paste into it
Keyword search alone
Exact identifiers, names, clause numbers, anything you can spell
△Real questions are situational and the answer usually spans several passages
Retrieval with citations
Changing knowledge, because the index changes without touching the model
△What we run, with limits stated: one shared knowledge base per deployment, citations returned but not independently validated, PDF only today
The last row is ours. Everything loaded into a deployment's knowledge base is readable by every user of that deployment, so you load only what everyone may see. Chat and knowledge states the limits next to the capability.
What people ask next.
QUESTIONS · 06
01Could you deploy a self-hosted model for us?
Not today. What we can do is tell you, during the assessment, whether your requirement actually needs self-hosting or whether a residency pin and an open-weight model on the managed route meets it.
02Is "the model runs inside our network" ever true?
With a self-hosted model, yes. With a managed model service, no. Our claim is narrower and checkable: the platform runs in a cloud account you own, data at rest stays in your account, and the connection to the model service can be a private network endpoint on the Standard and Strict tiers rather than the public internet. The boundary defines the term.
03What if a provider changes its terms?
Then the model behind the application layer changes, through a review and a configuration change, and your data, index and audit trail stay where they are. That is the argument for keeping the control point in infrastructure you own rather than in a subscription. It is also why the catalogue is generated from the provider's live list rather than typed into a page.
04Do the major providers train on our prompts?
The major cloud providers all publicly document that prompts and outputs are not used to train foundation models and are not shared with the third-party model makers, and all of them document narrow abuse-monitoring or safety processes as well, so nobody should write "nothing is ever retained". On our deployments the one named exception is disclosed and elected explicitly: the newest frontier tier from one vendor requires provider-side data sharing with retention up to 30 days, and those models show as unavailable until you choose otherwise.
⎯⎯ Book the strategic assessment ⎯⎯
Your private AI, inside your control ·
If your requirement needs self-hosting, the assessment will tell you that. You keep the topology and data-flow map either way.