Free AI is not free. You pay with data - yours and your customers'. The difference between "your data trains someone else's model" and "your data stays yours" often comes down to a single pricing plan. We went to the source and read the data policies of OpenAI, Anthropic, Microsoft and Google - with dates, quotes and links, because these documents can change over a single August weekend.

One question before we start: when did someone at your company last paste something into ChatGPT? Exactly - you don't know. That's what this article is about.

Does AI train on your company's data?

It depends on one thing: which account your team uses. Consumer accounts - ChatGPT Free and Plus, Claude Free, Pro and Max - use your conversations for model training by default. Business plans - ChatGPT Business and Enterprise, Claude Team and Enterprise, the API, Microsoft 365 Copilot - don't, by default. The question is not "is AI safe". The question is which door you use to let your data in.

In practice this means one simple, uncomfortable truth: if your team uses private free accounts for work, the content of those conversations - customer data included - can end up training someone else's models. Legally, and in line with terms of service everyone accepted without reading.

Training on consumer-account conversations is not an option you consciously turn on. It's a setting you have to consciously turn off. At both OpenAI and Anthropic, the "improve the model" toggle is on by default. Classic opt-out by inertia: the consent box was ticked for you.

Anthropic introduced these rules on August 28, 2025 - Free, Pro and Max users had until October 8, 2025 to make a choice. Anyone who clicked nothing is training. Anthropic's consumer privacy policy says it plainly: "We will train new models using data from Free, Pro, and Max accounts when this setting is on".

There is one more detail almost nobody knows. At Anthropic, opting out only works going forward: your data "will still be included in model training runs that are already in progress". You switched it off today - you didn't undo yesterday. Retention depends on the same choice: with training consent your data lives up to 5 years, without it - 30 days.

What's different about a business plan?

Business plans flip the logic: training on your data is off by default. OpenAI states on its Enterprise privacy page: "By default, we do not use your business data for training our models" - and explicitly lists ChatGPT Business, Enterprise, Edu and the API. Anthropic's commercial terms work the same way: no training on your prompts or code unless you actively join their partner program.

Microsoft 365 Copilot goes a step further because it operates under an agreement with your organization. Per Microsoft's enterprise data protection documentation: "the prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation models". Microsoft formally acts as a data processor under a Data Processing Agreement (DPA) - the exact legal construct that privacy law expects when you hand customer data to a third party.

For API-scale use there's also Zero Data Retention: OpenAI can refrain from storing request content at all - on request, for eligible endpoints.

Watch out for one naming trap: "Microsoft 365 Copilot" (the business one, under contract) and "Microsoft Copilot" at copilot.microsoft.com (the consumer one) are two different worlds. The consumer version trains on your interactions until you turn it off in account settings.

Comparison: what each tool does with your data

The dividing line doesn't run between vendors - it runs between account types. OpenAI, Anthropic, Microsoft and Google all follow the same pattern: a consumer account feeds your conversations to the models by default, a business account doesn't. Before you pick a tool for your company, don't look at the logo. Look at the table row your team actually sits in.

Tool / plan Trains on your data How to change it
ChatGPT Free / Plus YES by default Settings → Data controls → turn off "Improve the model for everyone"
ChatGPT Business / Enterprise / API NO by default - (training only on explicit opt-in)
Claude Free / Pro / Max YES by default (since Aug 28, 2025) Privacy settings → turn off "model improvement" (forward-only)
Claude Team / Enterprise / API NO by default -
Microsoft 365 Copilot (business) NO, Microsoft = data processor under DPA -
Microsoft Copilot (consumer) YES by default account settings
Gemini (consumer) YES by default + some chats are read by humans turn off "Keep Activity" (reviewed chats are kept up to 3 years anyway)
Google Workspace with Gemini / Vertex AI NO (without your permission) -

Sources and access dates: openai.com/enterprise-privacy, privacy.claude.com (formerly privacy.anthropic.com), learn.microsoft.com, support.google.com/gemini, support.google.com/a (Workspace admin) - July 2026. These documents change; check the date, not just the content.

What about Google Gemini?

Same split, plus one thing few people know: in consumer Gemini, some of your conversations are read by humans. Google employs reviewers (including external contractors) who read selected chats - regardless of your settings. Conversations that went through review are kept for up to 3 years, even if you turn off activity history and delete everything. Google says it outright in its Gemini Apps privacy hub: "Please don't enter confidential information that you wouldn't want a reviewer to see".

On the business side it's clean: Google Workspace with Gemini doesn't train on your company's data and doesn't show it to reviewers without permission, and Vertex AI (the API) carries a zero-training commitment. Workspace also lets you pin data storage and processing to a European region.

What is shadow AI and why does it affect a 20-person company?

Shadow AI is any AI tool your employees use without the company's knowledge or oversight - most often through private free accounts. According to the Netskope Cloud & Threat Report 2026 (telemetry from millions of users, not a survey), 47% of people using generative AI at work do it from personal accounts. A year earlier it was 78% - companies are catching up, but half of all usage still happens outside any oversight.

Picture an ordinary Wednesday: your bookkeeper gets a spreadsheet of client records to sort. She pastes it into a free chatbot, because it's faster. From that moment, tax IDs, addresses and amounts live on a server abroad and - under default settings - can end up in model training. Nobody at the company knows, because there's nothing to notice: no alert, no log.

The scale? The same Netskope report counts an average of 223 data policy violations involving AI per month per organization, and incidents of sensitive data being sent to AI apps doubled year over year.

How much does it cost when it goes wrong?

According to the IBM Cost of a Data Breach Report 2025 (a Ponemon Institute study of 600 organizations that went through real breaches), one in five companies had a breach involving shadow AI. Organizations with high levels of shadow AI paid on average $670,000 more per breach than those with low levels or none.

Two more numbers from the same report show where the problem sits: among companies that had an AI model or application breached, 97% lacked access controls for those tools, and only 37% of all organizations studied have any AI governance policies at all. This is not a technology problem. It's an open-door problem.

What does the law say - GDPR and beyond?

When your employee pastes customer data into an AI tool, the responsibility doesn't vanish on the other side of the screen - your company remains the data controller. If you serve customers in the EU, GDPR applies regardless of where your company sits. The AI tool acts as a data processor, and Article 28 of the GDPR requires a Data Processing Agreement (DPA) for exactly this relationship. US state privacy laws such as the CCPA impose comparable contractual duties when you share personal data with a vendor.

And here the split from the beginning of this article comes back. On business plans, the DPA simply exists: at OpenAI you execute it self-serve in the dashboard (Settings → Compliance), at Anthropic it's automatically part of the commercial terms, Microsoft acts as a processor in Microsoft 365 Copilot, Google does the same in Workspace. On consumer accounts there's nothing to sign - a private account trains on your data and gives you no Article 28 basis at all.

The European Data Protection Board added another layer in December 2024 (Opinion 28/2024): a vendor's assurance that "the model is anonymous" is not enough on its own - the assessment is always case by case. And a company deploying someone else's model has its own duty to check whether it was built on unlawfully processed data. The vendor's compliance declaration doesn't close the question.

As for the EU AI Act, the obligations for a company merely using ready-made AI tools are narrower than the headlines suggest: transparency (Article 50) and AI literacy for your team (Article 4). Data privacy is governed by the GDPR, not the AI Act.

Three things to do today

  1. Check the toggles. ChatGPT: Settings → Data controls. Claude: privacy settings → "model improvement". Five minutes per account, zero cost.
  2. If anyone at the company works with customer data - move to a business plan. This is not a convenience expense. It's the line between "our data stays ours" and "our data trains someone else's model".
  3. Agree on one team rule: what we never paste. Customer personal data, contracts, financials. One page, not a 20-page policy. The policy can come later - the rule works from tomorrow.

The closing question isn't "is AI safe for our data". It's: which door did you leave open, and who knows about it. Data is one side of the coin - the other is control over what AI does day to day, which we break down in Can you trust AI in your business?.

If you want this checked properly - we audit where your data actually goes, deploy AI on business plans and train teams on what can and cannot be pasted.

Frequently asked questions

Does ChatGPT train on my data?

On Free and Plus accounts - yes, by default. Your conversations can be used for model training until you turn it off under Settings → Data controls. On Business, Enterprise and API plans - no, by default.

Does Claude use my conversations for training?

On Free, Pro and Max accounts - yes, if the "model improvement" toggle is on, and it has been on by default since August 28, 2025. Opting out only works going forward. Team, Enterprise and API plans don't train on customer data.

What is shadow AI?

AI tools used at work without the company's knowledge - most often private free accounts. According to the Netskope Cloud & Threat Report 2026, 47% of people using AI at work do it from personal accounts. According to IBM, one in five companies had a data breach involving shadow AI.

Is it safe to paste customer data into AI tools?

Not into a consumer account: there's no data processing agreement, and the conversations can train models. On a business plan with a DPA - yes, within reason: OpenAI's DPA is executed in the dashboard, Anthropic and Google Workspace include it in their terms, Microsoft covers it in Microsoft 365 Copilot. Your company remains the data controller either way.

Does my data stay in the EU?

Depends on the tool and configuration. One trap worth knowing: Anthropic models inside Microsoft 365 Copilot are excluded from the EU Data Boundary - Microsoft keeps them off by default for EU customers, and enabling them moves processing outside the EU.

How much does an AI-related data breach cost?

According to IBM (2025), organizations with high levels of shadow AI paid on average $670,000 more per breach than those with low levels or none. An average organization sees 223 AI data policy violations per month (Netskope).

Find out where your company's data really goes

We audit data flows, move your team to business plans and train them on safe AI use. Explore AI Trust Layer and AI Training

Let's talk