Solari SolutionsIT

Articles · Consultingdraft

Who trusts what. How companies choose between cloud AI and AI in-house

Published · by Dario Solari · 8 min read

The big AI providers promise by contract not to use companies' data. Many companies trust them, others do not, and both are right. What Samsung, the banks, Bosch, Lockheed, Airbus and the German government actually did, and the rule they share.

"Why should we run it in-house, if the provider guarantees by contract that our data is not used to train its models?" It is the question we are asked most often, and it is a serious one. Many companies use cloud services, ChatGPT, Copilot, Claude, Gemini, on exactly that basis: the contract says their data is not used, and that is enough for them. Others, above all in aerospace and defence, trust nobody.

They sound like opposite philosophies. Looking at what companies that took a public position actually did, it turns out they are not.

The 2023 bans were about the free version

In spring 2023 the bans made headlines one after another. Samsung blocked ChatGPT and all other generative AI services on company devices after employees, within twenty days, pasted semiconductor source code, test code and a meeting recording into it. Apple told staff not to use ChatGPT or GitHub Copilot. JPMorgan, Goldman Sachs, Deutsche Bank and Verizon did the same. In 2024 the US House of Representatives removed Microsoft Copilot from members' computers, deeming it "a risk to users due to the threat of leaking House data to non-House approved cloud services".

Almost all these bans targeted consumer tools, where training on inputs was then the default. And almost all of them, followed over the years, ended the same way: with an enterprise contract. Samsung, alongside its own internal models, adopted ChatGPT Enterprise in June 2026. The House lifted its ban in September 2025, with a version of Copilot with "heightened legal and data protections". The big banks built their own interface, with their own access controls and logs, in front of models rented from the big providers.

So the question was never "cloud or not". It was "under which contract".

What a contract actually promises

Take GitHub Copilot itself, on GitHub's official pages. On the Business and Enterprise plans code is not used for training: that is a contractual commitment. Suggestions requested from the editor are not retained; those requested from the website or the command line are, for 28 days; usage data for two years. The intellectual-property indemnity applies only if the filter blocking suggestions copied from public code is switched on. European data residency has existed since March 2026, on a specific variant of the service and at a ten per cent premium.

And one detail says a lot: since 24 April 2026 Copilot's individual plans, the ones an employee can switch on alone, do use data for training unless the user opts out. The same product, two different promises. Whoever trusts Copilot trusts a specific contract, and needs to know which one.

Who trusts nobody, and why

Then there are the organisations that run models on their own machines, sending nothing out. There are not many, and they resemble each other.

Lockheed Martin in 2024 presented a "fully internal" AI platform running on its own NVIDIA supercomputer "while maintaining data security on premises". In May 2026 Airbus signed an agreement with France's Mistral providing "highly secure on-premise deployments" for military and "highly confidential" applications. The German federal government uses KIPITZ, a platform based mainly on open models, which runs "in secure federally owned data centres" and is cleared for restricted classified documents. The French defence ministry has a framework agreement with Mistral.

The reason is not generic distrust: it is the law. US rules on exporting military material treat even sending technical data as an export, unless it travels encrypted end to end. A cloud model has to read the question in clear to answer it: the condition, as written, cannot be met. The EU dual-use regulation treats merely making a document available electronically as an export too. For anyone working with data of this kind, "we trust nobody" is not an attitude, it is an obligation.

The real rule: classes of data

The most interesting thing is that even those who refuse the cloud do not refuse it entirely. Bosch has six thousand developers using GitHub Copilot and, for the technical know-how it considers its asset, developed its own model with Germany's Aleph Alpha. Korea's SK hynix, which had blocked everything in 2023, in June 2026 reopened to external tools "in areas that aren't related to national core technologies". The US Army uses the Pentagon's platform, on Google models, for unclassified data, and a system of its own for secret data. Airbus uses Mistral on-premise for defence and trusted clouds for the rest.

None of them picks one provider for everything. All of them divide data into classes and decide class by class. Those who trust the contract and those who trust nobody, then, do not disagree: they have different data. For a software company's code, a Copilot Business contract is a reasonable guarantee. For the drawings of an aircraft component, no contract is enough.

And a small company?

All the examples so far are large organisations. But for a small company the question is more pressing, not less.

On the cloud side a small company starts at a disadvantage. Samsung, JPMorgan or the US House negotiate contracts with providers, have a legal department to read them and build their own interface with access controls and logs. A ten-person firm accepts the standard terms, when it reads them. And it often ends up, without deciding to, on individual plans: the employee who subscribes alone, the free version used from a phone. Those are precisely the plans where data is used for training unless switched off: ChatGPT always, GitHub Copilot since April 2026. And they are the plans that do not include the data processing agreement the GDPR requires when documents contain personal data. Samsung's 2023 problem, employees pasting confidential material into a public service, is more likely where the company does not give everyone a tool of its own and each person makes do.

In Italy, the third class of data belongs above all to small businesses. Law firms and accountants are bound by professional secrecy and, since Law 132 of 2025, must tell clients which AI systems they use. Medical and dental practices handle health data, the category the GDPR protects most. Suppliers to aerospace, defence and automotive work under non-disclosure agreements that the large customers write with their own rules in mind: Airbus's or Lockheed's policies reach suppliers through contracts. And design studios have their drawings as their only asset.

But for most small businesses the cloud is enough. The sums, at October 2026 prices, are clear. Ten seats on a business plan, ChatGPT Business, Claude Team or Copilot, cost around twenty euros per person per month: over three years, with rollout and staff time, about ten thousand euros. A machine on the premises for the same ten people costs, on its own, about two years of subscriptions; but the hardware is barely a quarter of the bill. The rest is installation, maintenance, updates and continuity: over three years it comes to around twenty-five thousand euros, and even the yearly running costs alone exceed those of the ten subscriptions. Below twenty or thirty people, AI in-house does not pay for itself. Meanwhile hardware is getting dearer because of the memory shortage, while cloud seats have become cheaper.

Then there is quality. An open model on a desk machine is today at the level of the best cloud models of a few months ago: almost level on questions about one's own documents, clearly behind on general knowledge, expert reasoning and agents (we wrote about it in What runs on a machine in the company, and how far it is from ChatGPT and Claude). And the cloud updates itself: a new model every few months, at no extra cost. On the premises every update is work, even though the same machine, kept updated, runs better and better models. Web search, deep research and integration with Office or Google Workspace remain mainly in the cloud. Many offices, moreover, already pay for AI without using it: Copilot Chat is included in Microsoft 365, Gemini in every paid Workspace plan.

Finally, the legal duties are the same either way: controller obligations, informing clients as Law 132 requires, the training the AI Act asks for. The on-premise system removes one thing only, but not a small one: the third party the data is entrusted to.

According to ISTAT, among Italian firms that considered AI without adopting it, 43.2% cite privacy and data protection as an obstacle. The concern is there; the method is often missing.

The method is the large organisations', at small scale. First an inventory of the data: what is public, what is personal, what is secret by law, by contract with clients or because it is the company's edge. Then a choice for each class. For most small businesses, a design studio, an agency, most of an accountant's work, the honest answer is a business cloud plan, with the contract actually read, an account for each person so nobody has a reason to use their own subscription, and the free versions blocked on company computers. Some need, alongside it, an on-premise system for one class of data only: a law firm's case files, a medical practice's clinical notes, the drawings a subcontractor receives under a non-disclosure agreement. Very few, those under military export rules or with clients who forbid any cloud, need it for everything. AI in-house is not bought to save money. It is bought for confidentiality, and only where confidentiality requires it.

So who trusts? Almost everyone, with something. Nobody, with everything.