Solari SolutionsIT

Articles · Consultingdraft

How far to trust ChatGPT and Claude with company data

Published · by Dario Solari · 7 min read

OpenAI and Anthropic promise not to train their models on companies' data and to delete it within thirty days. The promises are written down and certified. What they do not promise, and what no contract can cover, is what matters in the choice.

Anyone who entrusts documents to a cloud artificial intelligence service should ask three questions. Is the data used to train the models? How long does it stay on the provider's servers? Who can read it, and under which law? The answers are in the providers' contractual pages, read on 2 October 2026, and they vary a great deal from one plan to another.

OpenAI

Training. On ChatGPT's individual plans (Free, Go, Plus, Pro) conversations are used to train the models, unless this is switched off in the settings. On the business plans (Business, Enterprise, Edu) and on the API they are not, unless the company explicitly authorises it.

Retention. Deleted conversations are removed within thirty days, "unless we are required to retain it for legal or security reasons". The API keeps abuse-monitoring logs for thirty days; on request, and for qualifying use cases, OpenAI grants zero data retention, that is, no retention of content. On the Business plan, OpenAI staff and "specialized third-party contractors" may access conversations for abuse review.

The "unless legally required" clause put to the test. In May 2025 a New York court, in the New York Times case against OpenAI, ordered the preservation of all conversations that would otherwise have been deleted. The order covered the individual plans, the Team plan and the API without a zero data retention agreement; it did not cover Enterprise nor, going forward, conversations originating in the European Economic Area. It lasted until 26 September 2025. In January 2026 the judge upheld the order to hand over to the publishers a sample of twenty million de-identified conversations. The lesson is simple: what exists on an American server can be demanded by an American court.

European residency. Since February 2025 OpenAI has offered data storage in Europe for the API and for new Enterprise and Edu customers. The Business plan is not on the list. Even with in-region processing, some steps (text extraction from documents, routing, authentication) "may still occur outside the selected region". The contract is with OpenAI Ireland, but transfers to group companies outside the Union are provided for, under the standard contractual clauses.

Anthropic

Training. Since August 2025 users of Claude's individual plans choose whether their conversations may be used for training; those who accept see retention rise to five years, those who decline stay at thirty days. The commercial plans (API, Team, Enterprise) are excluded from training.

Retention. Thirty days, with exceptions of up to two years for breaches of the usage policy. Zero data retention exists, per organisation and through the sales team, but it does not cover the most recent models (the Fable 5 and Mythos 5 families), which require thirty-day retention.

Residency. On Anthropic's API and applications the only storage region available is the United States. Europe can be reached only through Amazon Bedrock or Google Cloud, where the cloud provider becomes the data processor.

For completeness: Microsoft Copilot and Google Workspace with Gemini both state that they do not train models on customer data and offer storage in Europe; Microsoft warns that at times of peak demand calls to the models may be served from other regions, and that Anthropic models are excluded from the EU Data Boundary.

Anonymisation: what exists

None of the providers anonymises what it receives. There is no function that strips names, tax codes or amounts from a document before the model reads it. The measures on offer are others: not training, not retaining, storing in region, encrypting with customer keys. Microsoft, on Azure, sells separately a personal-data detection service that can replace them with placeholders, but it is a component to be integrated, not a setting.

Anonymisation, if wanted, has to be done before sending, within the company: an intermediate piece of software recognises personal data, replaces it with placeholders, keeps the mapping table only locally and reapplies it to the answer. Open-source tools such as Microsoft's Presidio serve this purpose. Two limits must be stated.

The first is legal. Under the GDPR, pseudonymised data remains personal data for whoever holds the key, that is, for the company; the Court of Justice reaffirmed this in September 2025. Pseudonymisation is a security measure, not a way out of the obligations.

The second is practical. Automatic recognition removes the names, not the context: the description of a case identifies a person even without the name. And much of what a company wants to protect is not personal data: a price, a drawing, a price list, a piece of code. A study from April 2026 that compared eight redaction techniques found that the best combination let no personal data through, but let through 31% of proprietary code.

How far to trust them

OpenAI's and Anthropic's policies are serious, written down and certified (SOC 2, ISO 27001). In the incidents of the last three years the weak point has rarely been the provider's server. Company data has got out in three ways.

  • Employees pasting data into an individual service. The Samsung case, in 2023: source code and meeting minutes in ChatGPT, three times in twenty days.
  • Instructions hidden in documents. In June 2025 a flaw in Microsoft 365 Copilot (EchoLeak) made it possible, with a simple e-mail, to have internal files sent to an external server without any action by the user; in September 2025 a similar flaw in ChatGPT (ShadowLeak). Both fixed, neither exploited as far as is known; both possible by construction.
  • Features that publish by design. In summer 2025 thousands of conversations shared from ChatGPT ended up indexed by Google; in September 2026 OpenAI confirmed that its own agents had published images uploaded by users who had consented to training.

To these is added jurisdiction. In June 2025 the legal director of Microsoft France declared under oath before the French Senate that he could not guarantee that French customers' data would not be handed over to the American authorities. Data residency in Europe does not change the law the provider answers to.

What remains the company's responsibility

Whatever the provider, the company remains the data controller: it must have a legal basis, inform the people whose data ends up in the documents, sign a contract appointing the provider as processor (individual plans do not provide for one), and assess the impact when the data is sensitive or concerns employees. Professionals must also tell their clients which systems they use, as Italian law 132/2025 requires.

A working rule

Trust is not a matter of brand but of contract and of the kind of data. A workable criterion divides data into three classes.

  1. Non-confidential information (public texts, generic drafts, translations of material already in circulation): a cloud service with a business contract will do.
  2. Personal data of customers and employees: a business plan with a processor contract, zero retention where available, pseudonymisation before sending; or an internal system.
  3. Professional secrecy, intellectual property, health data: a system that runs on the company's premises and transmits nothing outside.

Most companies use all three classes every day. Deciding which tool each one needs is the work to be done before choosing the provider.

Sources

  • OpenAI, How your data is used to improve model performance; Enterprise privacy; Data controls (API); Data residency for ChatGPT; Response to NYT data demands; Data Processing Addendum v.010126.
  • U.S. District Court S.D.N.Y., In re OpenAI, 25-md-3143, order of 5 January 2026.
  • Anthropic, Updates to Consumer Terms and Privacy Policy (28 August 2025); API and data retention; Data residency; Data Processing Addendum.
  • Microsoft Learn, Data, Privacy, and Security for Microsoft Copilot; Azure AI Language, PII detection. Google, Workspace AI privacy.
  • Court of Justice of the European Union, case C-413/23 P, 4 September 2025. Regulation (EU) 2016/679, recital 26 and arts. 4, 28, 32, 35.
  • LLM-Redactor, arXiv 2604.12064, April 2026. Microsoft Presidio, GitHub.
  • Checkmarx and arXiv 2509.10540 on EchoLeak (CVE-2025-32711); The Register on ShadowLeak; TechCrunch and Bloomberg on the Samsung case; Search Engine Land on the indexed conversations; Reuters on the published images (26 September 2026).
  • Hearing before the French Senate, 10 June 2025 (Microsoft France).