Why your models run inside your cloud, not ours.
Every private-AI vendor will tell you your data is safe. Almost all of them mean it is safe on their computer. The choice that decides every downstream question is whose cloud account the platform runs in, and it is worth understanding before you sign anything.
- Private has two meanings. Most vendors mean they hold your data privately. The other meaning is that the account holding it is one you opened and are billed for.
- The frontier labs now sell through the big clouds. All three major providers publish the same two commitments in their own documentation: prompts and outputs are not used to train foundation models, and are not shared with the model makers.
- Inside your cloud is a claim about the account, not the network. The model does not run inside your network. Nobody's does, and any vendor telling you otherwise is describing something else.
- Owning the account buys four concrete things: the compute bill, the encryption keys, the audit records, and the exit.
- Residency is the part that costs. In our generated catalogue of 104 models, 13 process inside Canada, and every current frontier model routes outside it.
A question I get asked in almost every first call is some version of "is it private?" It is the right question and it is almost impossible to answer honestly in one word, because the word private is doing two completely different jobs in the market at the same time. One of them is about how carefully a vendor looks after data they hold. The other is about who holds it. They sound alike in a sales deck and they lead to entirely different systems.
Private has two meanings
Ask a typical AI vendor whether their product is private and you get an accurate, well-rehearsed answer. Your tenant is logically separated. Data is encrypted in transit and at rest. Your content is not used to train the model. There is a SOC 2 Type II report, available under NDA. Every one of those can be true, and none of them changes the arrangement underneath: your material sits in a database the vendor owns, governed by controls the vendor operates, evidenced by an audit the vendor commissioned.
That last point deserves precision, because attestations get quoted as though they settle the question. SOC 2 is an attestation report, not a certification: a licensed accounting firm examines a service organization's controls against the Trust Services Criteria and issues an opinion about that organization's controls.1 It tells you the vendor is a serious operator. It does not tell you what happens to your data if the vendor is acquired, changes its terms, or stops answering the phone.
Most "private AI" means they hold your data privately. Ours means you do.The line the whole practice hangs from
The second meaning of private is structural rather than procedural. The platform runs in a cloud account that your organization opened, that your organization is billed for, and that your people can sign into without asking anyone's permission. Isolation is not a tenant identifier in someone else's database. Isolation is the account itself: one organization, one account. Everything else in this article is a consequence of that one decision.
What actually changed
This arrangement was not practical a few years ago. If you wanted a frontier model you called the model vendor's own API, which meant your prompts went to the model vendor, full stop. Self-hosting was the only alternative, and the models you could realistically self-host were a long way behind.
What changed is that the frontier labs now also distribute their models through the large cloud providers' managed model services. The model becomes a service you call from inside an account you already control, with your own identity, your own audit trail and your own invoice. That is the development the category rests on, and it is worth reading what each provider commits to in writing rather than taking anyone's summary of it.
What the providers publish
On Amazon Bedrock, the service FAQ states: "No, AWS and the third-party model providers will not use any inputs to or outputs from Amazon Bedrock to train Amazon Nova, Amazon Titan, or any third-party models," and "Users' inputs and model outputs are not shared with any model providers."2 The user guide explains the mechanism rather than just asserting the outcome: each model provider's weights run in a deployment account owned and operated by the service team, and "Because the model providers don't have access to those accounts, they don't have access to Amazon Bedrock logs or to customer prompts and completions."3
Google Cloud's documentation says "Gemini doesn't use your prompts or its responses as data to train its models,"4 and its published privacy commitment says "Google won't use your data to train or fine-tune any AI/ML models without your prior permission or instruction."5
Microsoft's page for models sold by Azure is the most explicit of the three. Your prompts, completions, embeddings and training data "are NOT available to other customers," "are NOT available to OpenAI or other providers of Models sold by Azure," "are NOT used by providers of Models sold by Azure to improve their models or services," and "are NOT used to train any generative AI foundation models without your permission or instruction." The same page adds that "The models are stateless: no prompts or completions are stored in the model."6
All three also document a way to reach the service over private network connectivity rather than the public internet.7 Three different companies, competing hard, publishing the same two commitments and the same connectivity option. That convergence is what makes a governed private deployment a buildable thing rather than a marketing position.
They are about training and about sharing with the model makers. They are not a zero-retention claim, and you should be suspicious of anyone who reads them that way. All three providers also document narrow abuse-monitoring or safety processes. Microsoft's page states that when its abuse-monitoring system flags potential misuse, "a sample of customer's prompts and completions may be selected for review... with additional reviews by human reviewers as necessary," and that approved customers can apply for modified abuse monitoring under which that storage and human review are not performed.6 The supportable claim is the one the documentation supports: not used to train foundation models, not shared with the model makers.
What inside your cloud does and does not mean
Here is where most copy in this category quietly overreaches, so let me draw the line carefully in both directions.
What it does mean
The account is the boundary. The dashboard, chat handler, document store, knowledge index, user directory, usage records and audit configuration all live in an account your organization owns, and the cloud provider bills that account directly. We hold no client data on our own infrastructure and we do not resell compute. To find out what is running, you look at your own account, not a status page we control.
What it does not mean
The model does not run inside your network. A managed model service is a service: the request leaves your subnet, is processed by the provider's model-serving infrastructure under the terms quoted above, and the answer comes back. Nobody serving a current frontier model is running it on a box in your office either, and the open-weight models you genuinely could run in your own network are not the same models. Anybody implying otherwise is describing self-hosted open weights, which is a legitimate but different choice, or being careless.
It also does not mean "your data never leaves your boundary." That phrase gets written a lot and it is not true of any managed model service. The accurate version is narrower and easier to defend: data at rest stays in your account, and the inference request is processed by the managed model service under published terms that neither train on it nor share it with the model maker. If a vendor will not draw that distinction for you unprompted, they either have not thought about it or they are hoping you will not.
What owning the account buys
Ownership is not a feeling. It is four specific, checkable things.
The bill
The cloud provider invoices your organization directly, at list price, for the compute you used. We never sit between you and that bill, so there is nothing for us to mark up. It also means the cost signal is real: if a team's usage triples you see it on the invoice that month, attributed by per-person metering rather than inferred.
The keys
Encryption posture is a choice, and here is the whole of it. On the Strict tier, data at rest is encrypted with customer-managed keys with rotation enabled. On Standard, provider-managed keys. At Baseline, provider-owned keys. Baseline is the low-cost entry posture and it suits low-sensitivity use and a first evaluation. A security review will usually want Standard.
The audit records
Model invocation logging runs at every deployment: model, caller identity, latency, token counts. It deliberately carries no prompt or answer text, so it is an accountability record rather than a transcript archive, and I would rather you know that now than discover it during an investigation. Account-level audit logging, multi-region with file validation and one-year retention, is available on Standard and Strict. An administrator reading someone's chat transcript is itself written to the audit record.
The exit
If your AI vendor disappeared tomorrow, would anything still run? With us, you log into your own cloud and keep running.What ownership is actually for
The platform and its data live in your account. The monthly fee buys the service, not access. If the service stops, you keep the deployed platform and its source and can run it or hand it to another provider. The honest qualifier: running it afterwards needs a competent cloud engineer, which is why runbooks and infrastructure-as-code are contract deliverables rather than favours. And it is deeply built on one cloud provider, so moving to a different provider would be a rebuild, not a migration. What is composable here is the models, where switching is configuration.
It is worth noting that Canada's own government uses a three-part test for sovereignty, not a one-part one. Its sovereign compute infrastructure programme requires a "Canadian-located, Canadian-governed system that ensures data residency, operational control, and decision-making authority and agency remain in Canada."8 Residency is the first of the three. Operational control and decision-making authority are the other two, and they are precisely what owning the account gives you and what a tenancy on somebody else's platform does not.
A deny, not a setting
This is the part that changes a security review, and it is the least glamorous thing on the page.
A content policy configured inside an application is a setting. Someone with the administrator password can change it. A new integration can be written that calls the model a different way and skips it. A code path added in a hurry can forget it. The control exists as long as everyone behaves, which is another way of saying it is a promise with a checkbox next to it.
Enforced as a deny in cloud identity policy, the same control behaves differently. A request that does not carry the content policy is refused before it reaches a model, no matter which piece of code made it, no matter who is signed in, no matter what the application thinks. Four denies apply at every deployment: an inference call without the content policy attached, inference outside the configured endpoint regions, inference that bypasses the governed route, and any widening of the account's data-retention election. Tests fail the build if any of those denies is weakened, which is one of the reasons the suite has grown to roughly 2,700 automated test cases across about 170 test files.
Two related properties matter to a reviewer. The content policy is a pinned, numbered version rather than a mutable draft, so changing it forces a new version and what applied to an answer last quarter stays identifiable. And administrators, deliberately, cannot change the model allowlist, the policy, the posture tier, the residency setting, or admit a new module. Those are deploy-time parameters changed through a release. Nothing in the running system can create a role.
Why this matters more than it sounds: among organizations that reported AI-related breaches in IBM's 2025 breach study, 97% said they lacked proper access controls, and 63% of the breached organizations studied lacked AI governance policies.9 The failure mode in that data is not exotic model attacks. It is ordinary missing controls around ordinary access. A deny in identity policy is boring, and boring is the point.
The same discipline is applied to what we tell you about your own deployment. The page inside the product that describes your resolved controls is generated from the same typed object that builds the infrastructure, and a test fails the build if the page and the exceptions table disagree. Your standing exceptions, the things we have not yet closed, are shown to you in the product rather than discovered by you later.
The residency tradeoff
Residency is where this design has a real cost, and it is the thing I most often have to talk a prospective client down from, not up to.
Our model catalogue is generated by script from the provider's live catalogue and classified by where inference actually runs, never typed into copy by a person. At the 5 September 2026 generation it held 104 active models from 18 labs. Classified from the point of view of a Canadian deployment: 13 process in-region in Canada, 1 is routed within Canada only, 18 are reachable but processed outside Canada, and 72 are served from United States and global regions.
Read those numbers carefully, because the second one is the whole story. The in-Canada set is previous-generation: Claude 3 Sonnet and Haiku, Llama 3 8B and 70B, Mistral Large 24.02, Mixtral 8x7B, Mistral 7B. Every current frontier model routes outside Canada. If you pin a deployment to Canada you are choosing a genuinely enforced control and you are giving up the models most people mean when they say AI. Anyone who sells you residency as free is either not measuring it or not telling you.
Buyers feel this tension in the data too. In Cisco's 2025 privacy benchmark study, 90% of respondents said data would be inherently safer if stored within their own country or region, 88% said localization adds significant cost, and 91% said global providers can protect data better than local providers.10 The workable answer is not one position for the whole organization but a decision per workload, which is what a configurable pin is for. Residency is configurable, pinned to your region when you need it, with the price on the label. There is a whole article on it: data residency is a choice with a price.
What it costs
Three separate numbers get conflated in this market, so here they are separately.
The platform's own infrastructure. On your cloud bill, excluding any inference, roughly $5 to $10 a month at Baseline and about $110 a month on the private-network tiers. The gap is almost entirely private networking. These are figures from deployments we operate, not a quotation.
Inference. Billed by your provider at list price. As an estimate, a knowledge question of about 10,000 tokens in and 1,000 out costs about 4 to 5 cents at a mid-tier frontier model's list price, on the order of $3 per million input tokens and $15 per million output. Light team usage across a whole deployment has run an estimated $40 to $120 a month all-in on the cloud bill. Again: what we have seen, not a quote.
Us. The strategic assessment is $2,500 and bounded, and you keep the deliverable whether or not you proceed. Deployment starts from $2,500 and typically lands around $5,000 once the solutions pipeline is scoped, invoiced by milestone with nothing at signing. Managed operations are $500 a month, cancel anytime. Compute is billed by your cloud provider directly to you, at their list price, which is why there is nothing for us to mark up.
What this does not fix
A page that only makes claims is a sales page. Here are the limits that belong next to them, and there is a longer list at what we do not claim.
- The managed model service processes the request. The platform runs in your account and data at rest stays there, but the model does not run inside your network. That is the sentence to hand your reviewer, in place of "data never leaves your boundary".
- Residency costs capability. A jurisdiction pin is genuinely enforced, and pinning to Canada costs you every current frontier model. Thirteen of the 104 in the catalogue process in-region.
- The knowledge base has no per-document permissions today. Everything loaded into a deployment's knowledge base is readable by every user of that deployment, so load only what everyone may see. Citations are returned but not independently validated. PDF only today.
- We hold no third-party attestation. The platform is designed to align with SOC 2 and ISO/IEC 27001, and we provide the technical controls that support your programme, mapped as 17 technical control rows with an enforced, partial or not-enforced verdict against each. We do not hold a report of our own, and we say so.
- Baseline has no private networking. No private network path to the model service and no account-level audit trail. It suits low-sensitivity use and a first evaluation, and a security review will usually want Standard.
Three hosting models, six questions
Most decisions in this category are really a choice between three arrangements. The differences only become visible when you ask operational questions rather than feature questions.
| The question | Vendor-hosted SaaS | Single tenant, in the vendor's account | Platform in an account you own |
|---|---|---|---|
| Whose account holds the data at rest | ✕The vendor's | ✕The vendor's, dedicated to you | ✓Yours |
| Who receives the compute invoice | ✕The vendor, priced per seat | ✕The vendor, repriced to you | ✓You, at provider list price |
| Where the content policy is enforced | △Application settings you configure | △The vendor's configuration of your instance | ✓A deny in your account's identity policy |
| What your auditor can read directly | ✕The vendor's attestation report | △Attestation plus instance documentation | ✓Your own account's configuration and logs |
| Can inference be pinned to a jurisdiction | △Only if the vendor offers it | △Usually, at the vendor's discretion | △Yes, at a stated capability cost |
| What still runs if the vendor stops | ✕Nothing | ✕Nothing you can reach | ✓Everything, if you have an engineer |
The middle column is the arrangement most often described as "private" in this market, and it is a real improvement on the first. The last column is a different kind of claim, and the last row is where the difference is decided. Cells marked with a triangle are judgments about how these arrangements are typically sold, not measurements.
There is a fourth option, which is to build it yourself. That is legitimate for an organization with a platform team, and it is worth knowing what the evidence says about the odds. In MIT NANDA's 2025 study, pilots built with an external partner reached deployment about 67% of the time against about 33% for internally built tools, with the report's own caveat that this reflects an interview sample of 52 organizations and the correlation "does not necessarily prove causation."11
Ten questions to ask any private-AI vendor
You do not need to be an infrastructure engineer to run this list. Ask them in order and write down the answers. A vendor who answers all ten crisply is worth your time even if the answers are not the ones in this article.
- Whose cloud account holds our data at rest, and whose name is on the invoice for the compute? If the answer needs a paragraph, it is their account.
- Can you show me, in writing from the model provider, that our prompts are not used for training and not shared with the model maker? Not your summary of it. The provider's page.
- Where is the content policy enforced, and what exactly happens if an application tries to call a model without it? Listen for "it is refused by policy" rather than "it is configured to".
- Can an administrator inside the running system switch that policy off? The right answer is no, with an explanation of what would have to happen instead.
- Is the policy versioned, and can you tell me which version applied to an answer given last quarter?
- What is in your audit record, and what is deliberately not in it? A vendor who cannot tell you what is missing has not looked.
- Which models can be pinned to our jurisdiction, by name, and what do we lose by doing it? If the answer is "all of them", ask them to generate the list.
- What happens to the running system if we stop paying you, and what would it take for someone else to operate it?
- What have you built that is not yet proven at customer scale, and what is not built at all? The quality of this answer tells you more than any of the others.
- Who is the person who will actually do the work, and what happens when they are unavailable?
Nine: the module runtime is the most complete subsystem and client modules in use today run on the staging channel, with the production promotion path built and exercised on our own deployment. Client-authored modules are in pilot. Not built at all: per-user or per-document permissions on retrieval, and a tunnel into your office network.
Ten: org.tech is a small senior practice, one principal with a bench of contractors, and the leverage comes from the platform and the deployment factory rather than headcount. That is a real answer to question ten and it is a real constraint, which is why continuity is written down rather than implied.
What to do with this
If you take one thing from this article, make it the distinction at the top: a vendor telling you your data is private is answering a different question from the one about whose account it is in. Both answers can be good. Only one of them survives the vendor.
And the recommendation at the end of it is not the one you might expect from a page like this.
The recommendation is not "trust Sam". It is: own the environment, inspect the evidence, and keep the exit open.How to buy anything in this category
That holds whether or not you ever work with us. If you own the account, you can change your mind. If you do not, you are relying on a vendor continuing to deserve your trust, indefinitely, with no way to check. The whole design follows from preferring the first arrangement. To see it against your own environment rather than in an article, the strategic assessment is the bounded way to do it.
Sources
- AICPA and CIMA, "System and Organization Controls: SOC Suite of Services," accessed 20 September 2026. aicpa-cima.com ↩
- Amazon Web Services, "Amazon Bedrock FAQs," accessed 20 September 2026. aws.amazon.com/bedrock/faqs ↩
- Amazon Web Services, "Data protection," Amazon Bedrock User Guide, accessed 20 September 2026. docs.aws.amazon.com ↩
- Google Cloud, "How Gemini products in Google Cloud use your data," last updated 18 September 2026. docs.cloud.google.com ↩
- Google Cloud Blog, "Google Cloud unveils AI and machine learning privacy commitment," accessed 20 September 2026. cloud.google.com ↩
- Microsoft Learn, "Data, privacy, and security for Foundry Models sold by Azure in Microsoft Foundry," last updated 5 June 2026. learn.microsoft.com ↩ ↩
- Private connectivity, each provider's own documentation: Amazon Web Services, "Protect your data using Amazon VPC and AWS PrivateLink," docs.aws.amazon.com; Google Cloud, "VPC Service Controls with Vertex AI," docs.cloud.google.com; Microsoft Learn, "Securing Azure OpenAI inside a virtual network with private endpoints," 26 November 2025, learn.microsoft.com. All accessed 20 September 2026. ↩
- Innovation, Science and Economic Development Canada, "AI Sovereign Compute Infrastructure Program," page modified 1 June 2026. ised-isde.canada.ca ↩
- IBM, "Cost of a Data Breach Report 2025," summarized at IBM Think, 12 November 2025. Figures are for organizations reporting AI-related breaches within the study, not for all organizations. ibm.com ↩
- Cisco, "The Privacy Advantage: Building Trust in a Digital World, Cisco 2025 Data Privacy Benchmark Study," 2025, Figures 1 and 2. Survey of 2,600+ security and privacy professionals in 12 countries; Canada was not among them, and Cisco sells in this market. cisco.com (PDF) ↩
- MIT NANDA (Aditya Challapally, Chris Pease, Ramesh Raskar, Pradyumna Chari), "The GenAI Divide: State of AI in Business 2025," July 2025, section 6, pp. 19 to 20. The report describes itself as preliminary findings from 52 organizational interviews and 153 survey responses. mlq.ai (PDF) ↩