August 6, 2026 · KeyCaliber
Does Your SaaS Vendor Train on Your Data?
Finding shadow AI is the easy half. The question your DPO, your auditor, and your board ask next is whether the AI in use — sanctioned or not — feeds your data back into someone else's model.
Most organizations have now run the shadow AI exercise. Pull the proxy logs, grep for the provider domains, produce a list of AI services in use. It is a useful list. It is also the easy half of the problem.
The question that follows is the one that actually reaches legal, privacy, and the board: for each of those services, does the vendor train on our data? And the harder companion to it: which of the SaaS apps we already approved have quietly turned an AI feature on for us?
Neither question is answered by a domain list.
Three postures, and only one of them is fine
Every AI vendor sits in one of three places, and the difference is material:
- Trains by default. Customer prompts and uploads are used to improve the model unless someone explicitly opts out. Common on consumer and free tiers.
- Opt-out available. Training happens unless a setting is changed or an enterprise agreement is in place. The setting exists. Whether anyone flipped it is a separate question.
- No training on customer data. Contractually excluded, usually on enterprise or API tiers.
Two employees can use what looks like the same product and land in two different postures, because one is on a personal free account and the other is on the corporate tenant. A list of domains cannot tell those apart. Neither can a list of applications. You need the account and the tier behind the traffic.
And these postures move. A vendor changes its terms, adds an enterprise tier, or reverses a default, and last quarter’s assessment is wrong. A policy answer without a source and a date is not evidence — it is a memory. Any record worth presenting to an auditor cites where the policy came from and when it was last reviewed.
The AI you never approved, inside the software you did
The uncomfortable category is not the rogue chatbot. It is the AI feature your sanctioned vendor shipped into a product you already run.
The application went through procurement. The AI assistant inside it did not. Depending on the vendor, it arrived in one of three ways:
- Enabled by default — on for every user the day it shipped, with no admin action.
- Admin-controlled — off until someone with tenant rights turns it on, which may already have happened without a review.
- Opt-in per user — each user chooses, which means your exposure is however many of them chose.
This is the gap most AI governance programs miss entirely, because every control they built points outward at unsanctioned services. The vendor is on the approved list. The data flow is new.
Why no single tool answers this
Answering “which AI is in use, whose is it, and does it train on our data” requires three things at once:
- Discovery — network, identity, endpoint, and cloud signals correlated into an actual list of AI services and the assets and users behind them.
- A curated vendor catalog — the data-training posture for each provider, with a source URL and a review date, maintained as terms change.
- Identity context — which account, which tenant, which tier, so the posture attaches to the right instance rather than the brand.
Your proxy has the first ingredient and none of the others. Your CASB knows sanctioned apps but not the AI feature inside them. Your DLP sees one upload pattern. A vendor questionnaire captures a point in time and goes stale on signature. The answer only exists in the correlation, which is exactly why it stays unanswered in most environments.
Honest evidence beats a confident guess
There is a temptation, when you build this kind of inventory, to fill the blanks. Traffic goes to a provider, so infer the likely model. Infer the tier. Infer the posture.
Don’t. An AI inventory that guesses is worse than one with gaps in it, because the gaps are at least visible. A model name should appear only when a gateway or API log actually observed it. Everything else should read as unknown, and unknown should be a finding you can work — not a blank someone filled in to make the report look complete.
That distinction — observed evidence versus derived inference, clearly marked as which — is the first thing a privacy officer or an auditor tests. It is also what makes the rest of the inventory trustworthy.
From inventory to a decision someone owns
The output of this work is not a list. It is a set of decisions: this service is sanctioned, this one is prohibited, this one is under review. Each decision recorded with who made it and when.
That is the artifact that satisfies a DPO, survives an audit, and lets a CISO tell the board something more useful than “we found 47 AI tools.” It converts an observation into a position the organization has actually taken.
How KeyCaliber approaches it
KeyCaliber connects by API to the tools already in your environment — network, identity, EDR, cloud, and SaaS — and correlates their signals rather than adding another sensor. For AI specifically, that means:
- Services and models actually observed, including the major hosted providers and gateways, with the model name shown only when a log recorded it.
- Per-vendor data-training posture — trains by default, opt-out available, or no training on customer data — each carrying the source it came from and the date it was last reviewed.
- Embedded AI in sanctioned SaaS, classified by how it arrived: on by default, admin-controlled, or opt-in.
- Sanction, prohibit, or review recorded against each service, with the person and the timestamp attached.
And because every AI signal ties back to a real asset with a computed business impact, the AI integration reading your finance systems does not sit in the same undifferentiated list as a designer’s image tool.
Banning AI is not a strategy, and neither is a spreadsheet of blocked domains. Knowing which AI is in use, whose it is, what it does with your data, and who decided that was acceptable — that is a program you can defend.
← All articles