Solari SolutionsIT

Articles · Local infrastructuredraft

What runs on a company's own machine, and how far it is from ChatGPT and Claude

Published · by Dario Solari · 6 min read

The open models a company can run on a computer it owns are six months to a year behind the best models in the cloud. The gap, however, is not uniform. On questions put to one's own documents it almost disappears; on world knowledge and expert-level tasks it remains wide.

When a company considers an internal artificial intelligence system, the first question is always the same: how much worse is it than ChatGPT? The answer depends on two things: which machine one is prepared to buy, and what one wants to do with it. The figures that follow are the public ones as of 2 October 2026, with the limitation every leaderboard has: it measures what it measures.

Four machines, four models

Open models are downloaded and run without sending anything outside. What changes is the memory available.

MachineRecommended modelLicenceAA index
16–24 GB (Mac mini, laptop, gaming card)Gemma 4 12B, Ministral 3 14BApache 2.014, 6
a 32 GB card or a 36–64 GB MacQwen3.8-27BApache 2.034
128 GB (NVIDIA DGX Spark, Mac Studio)Qwen3.8-Flash-Next; gpt-oss-120bQwen Community 1.0; Apache 2.040; 12
256 GB (Mac Studio M5 Ultra)GLM-5.3-FlashMIT42

The index is Artificial Analysis's (version 4.3.2), an average of tests of reasoning, knowledge, coding and tool use. On the same scale the best models in the cloud score 58 today (Claude Opus 5.5) and 53 (GPT-6 Astra, Gemini 4 Argon, Claude Fable 5.1).

How far behind

The most honest way to read the gap is by date. Qwen3.8-27B, released in August, scores 34: the frontier of February–March 2026 scored between 32 and 39. At release it was therefore about six months behind; against today, 24 points. Gemma 4 31B (15) and gpt-oss-120b (12) sit below the GPT-5 of August 2025 (23): more than a year.

Two recent studies give the same measure. Epoch AI, in May 2026, puts the lag of the best open model overall at four months, but that model requires a data centre. Mozilla's report on open artificial intelligence, from September 2026, is more useful for a company: the best model that runs on a single graphics card is 23 points below the frontier, the best model on a single server 10 points below.

A caveat worth more than the leaderboard: the frontier has risen by more than twenty points in 2026 alone. A model installed today will be "a year behind" in a year's time, whichever it is. An internal system should be designed to change model, not to keep it.

Where the gap is

The detailed tests say where local models lose and where they do not. Comparing Claude Opus 5.5 with Qwen3.8-27B:

  • World knowledge. The Omniscience test, which rewards those who know and those who admit they do not know, gives +46 to Anthropic's model and −10 to Qwen; Gemma 4 and gpt-oss fall to −48 and −49. A local model is not an encyclopaedia and does not know recent facts: gpt-oss, for instance, stops at May 2024.
  • Autonomous coding. In the tests where the model works on its own in a terminal, 60% against 6%.
  • Expert reasoning. On academic-level questions, 61% against 34%.
  • Reasoning over supplied documents. Here the gap almost disappears: over roughly a hundred thousand words of documents, 85% against 82%. This is the typical use case of a company system, which answers by searching the company's files.
  • Faithfulness in summarising. In Vectara's leaderboard on fabrications in summaries of a given text (22 September 2026), small open models do better than the large ones: Gemma 4 26B fabricates in 5.2% of cases and Qwen3 8B in 4.8%, against 8.7% for GPT-6 Astra, 10.4% for Gemini 3.1 Pro and 10.9% for Claude Opus 4.5. Not all open models, though: gpt-oss-120b reaches 14.2%.

The reading is this: if the answers must come from the company's documents, a 27-billion-parameter model on a three-thousand-euro machine does the job. If they must come from the model's knowledge, or if the task is open-ended specialist reasoning, the frontier is needed.

And Italian

No independent evaluation from 2026 yet covers Qwen3.8 or Gemma 4 in Italian. The available tests, on models from a year ago, say something counter-intuitive: the small international models beat those trained in Italy. In the ITALIC test, published in May 2026, Ministral 3 8B scores 80–81%, Gemma 3 12B 76–79%, Qwen3 8B 73–77%; Velvet-14B 67% and Minerva-7B 35%. GPT-5 nano, OpenAI's smallest model, 86–87%. On the Italian part of the MMMLU test, gpt-oss-120b scores 85 against 91 for the closed model o3. The difference with the cloud exists, and it is a few points. For a company's documents the only test that counts is the one on the company's documents: it is the first step of every evaluation we do.

Licences and speed

Qwen3.8-27B, Gemma 4, gpt-oss, Mistral Small 4 and Ministral 3 are distributed under the Apache 2.0 licence; GLM under the MIT licence. Qwen3.8-Flash-Next has a licence of its own that allows internal use but requires an agreement from anyone offering it as a service to third parties. Meta's Llama 4 excludes European companies from the rights to the multimodal models: we do not recommend it.

Speed on machines one owns is adequate for a small office. Qwen3.8-27B generates 59–75 tokens (roughly words) per second on an RTX 5090 card and 48 on a Mac Studio M5 Ultra; on the DGX Spark 12 in the basic configuration, 31–58 with optimisations. gpt-oss-120b reaches 60 on the DGX Spark and 196 on an RTX PRO 6000. Reading the documents, which is the stage that weighs when an archive is queried, runs from 3,000 to 10,000 tokens per second on NVIDIA cards.

What it costs

A ChatGPT Business or Claude Team seat costs 20–25 dollars a month; Microsoft 365 Copilot for small businesses 18–25. A 27-billion-parameter model on two inexpensive cards costs, amortised, 13 cents per million tokens; electricity alone, one. For a company, though, the comparison is not made on the price per token: below thirty to fifty users the cloud costs as much as or less than a server over three years. The reason to keep models in-house is another, and we have described it elsewhere: the data that must not leave.

Sources

  • Artificial Analysis, Intelligence Index v4.3.2 and model comparison pages, read on 2 October 2026; Sub-32B open weights, 13 April 2026.
  • LMArena, text leaderboard, 2 October 2026.
  • Epoch AI, Open models lag state-of-the-art closed models by 4 months, 29 May 2026. Mozilla, The State of Open Source AI v1.1, 15 September 2026.
  • Vectara, Hallucination Leaderboard (HHEM-2.3), 22 September 2026.
  • Engineering, EngGPT2 benchmark, arXiv 2605.07731, 20 May 2026; OpenAI, gpt-oss model card, arXiv 2508.10925.
  • Model cards on Hugging Face: Qwen/Qwen3.8-27B, Qwen/Qwen3.8-Flash-Next (licence), google/gemma-4-31B, mistralai/Mistral-Small-4-119B-2603; Meta, Llama 4 Community License.
  • Speed measurements: llama.cpp, discussions #15396 and #16578; Kubesimplify, August 2026; MacStories, September 2026; LLMKube, 23 April 2026.
  • Prices: claude.com/pricing, 2 October 2026; ChatGPT Business and Microsoft 365 Copilot prices from secondary sources, to be re-checked.