Solari SolutionsIT

Articles · Local infrastructuredraft

An AI model on a blockchain. What really runs on ICP, and what does not

Published · by Daniil Tolmacov · 6 min read

The Internet Computer promises "on-chain" artificial intelligence, verifiable and free of any single cloud provider. We checked what has actually run inside the blockchain, what it costs, and where the prompt goes in the most advertised service.

Anyone who has worked in the Internet Computer ecosystem, as I have, developing in Rust for the Plug wallet, often hears the same question: can you run an artificial intelligence model on a blockchain? For ICP the short answer is yes, but only small models, slowly and at a high price. And the service that is now presented as "LLMs on ICP" does not run the models on the blockchain.

The question is not only for people who deal in cryptocurrencies. It concerns anyone who has to choose where to run artificial intelligence and whom to trust with their data. A blockchain and a server in the office are two different ways of not depending on a single cloud provider, with opposite consequences for confidentiality.

What has actually run

On ICP, programs are called canisters: WebAssembly code executed by every node of a subnet, all of which must reach the same result. In July 2024 DFINITY, the foundation that develops the network, completed the milestone called Cyclotron: deterministic floating-point arithmetic and SIMD instructions made canister computation about ten times faster. An image classifier (MobileNet), face recognition and GPT-2 then ran entirely inside the blockchain.

Since 2024 the open-source project llama_cpp_canister by onicai has brought llama.cpp into a canister, the same engine people use for local models on a laptop. These are the measurements published in its repository, updated in October 2026:

ModelParameters, quantisationTokens per call
Gemma 3 270M0.27 billion, 8-bit44
Qwen3 0.6B0.6 billion, 8-bit25
LFM2.5 1.2B1.2 billion, 4-bit9
Qwen3 1.7B1.7 billion, 4-bit6
LFM2.5 2.6B2.7 billion, 4-bit4

In July 2026 three independent researchers ran a 3.9-billion-parameter model at 2 bits, with byte-identical answers on all thirteen nodes of the subnet. It is the largest documented result. A 7–8-billion model, the minimum size of a useful office assistant, does not run on the blockchain.

Why so small

There are three limits. Each call can execute at most 40 billion instructions, about twenty seconds of computation; that is why the text comes out a few tokens at a time, and a short answer takes minutes. Usable memory is about 4 GB. And there are no GPUs: the next milestone, Gyrotron, announced in 2024 precisely to bring graphics cards into the network, has not been delivered as of October 2026.

Then there is the point that makes a blockchain a blockchain: every node of the subnet repeats the same computation. One answer is computed thirteen times. With the costs published by onicai, Qwen3 1.7B costs about 6,000 dollars per million generated tokens on the network, LFM2.5 1.2B about 1,900; the same million tokens with Llama 3.1 8B, a model five times larger, costs eight cents at an ordinary provider.

The service called "LLMs on ICP"

In February 2025 DFINITY introduced the "LLM canister", now also called the Intelligence Gateway: a few lines of code to use Llama 3.1 8B, later Qwen3 32B and Llama 4 Scout, from a canister. On the foundation's forum its own engineers explained how it works. The canister queues the requests; they are processed by "AI workers", machines outside the blockchain run by DFINITY, which in this first version used "an external provider". In June 2025 they added: "we're currently limited to OpenRouter's available models", OpenRouter being an ordinary commercial model broker. And in August, about calls to outside services: you are still off-chain, so "there's no consensus or verification" that the result was not tampered with.

The current documentation instead gives "Onchain (ICP nodes)" as the place where inference runs. We found no announcement of this change, the network has none of the GPUs an 8-billion model would need at a usable speed, and a public question in September 2026 about where the models are hosted went unanswered. Until shown otherwise, the prompt leaves the blockchain, is processed by a machine nobody on the network can check, and the blockchain records the answer without being able to verify it.

Verifiable does not mean confidential

When the model really runs inside a canister, ICP offers something a server does not: anyone can rebuild the code, check which model answered and know that no single node altered the result. It is a real guarantee, and it has a price that should be stated plainly.

DFINITY's security documentation says: "Node operators on standard application subnets can read canister memory." A prompt processed on the blockchain sits in plain text in the memory of thirteen machines, in thirteen data centres, run by thirteen operators the company does not know. The network's cryptographic keys (vetKeys) protect stored data, not computation on that data. Subnets with protected execution environments have existed since 2026: there are two, they have seven nodes and no GPUs.

For a contract, a payslip or a medical record, the question that comes first is not "which model answered?" but "who else has seen this data?".

Where it makes sense

A small model on the blockchain makes sense when verifiability matters more than confidentiality and speed: an automatic classification tied to a payment or a public record, a decision that anyone must be able to repeat and check, an agent that manages funds. For everything else, the same models of 0.5–2 billion parameters run much faster on an ordinary office server, and if proof of which model answered is needed, a fingerprint of the model and the input data can be published.

Three questions for any provider

The "on-chain" label deserves the same scrutiny as "private", "secure" or "in Europe" on any artificial intelligence service. There are three questions, and they apply to us too.

  1. Where does the model physically run? On which machine, whose, in which country.
  2. Who can read the prompt? The provider, its intermediaries, the node operators.
  3. Can you verify which model answered? With which version, and with which data.

A local system answers the first two in the simplest way: on a machine of the company, and only whoever the company authorises. The third is answered by logging versions and requests. The blockchain answers the third well and the second badly. Knowing which of the three matters most, case by case, is the first step of every choice.

Sources

  • DFINITY, "The Next Step for DeAI: On-Chain Inference Enabling Face Recognition", 15 July 2024 (Wayback Machine copy).
  • onicai, llama_cpp_canister, README and model sheets, GitHub, October 2026; DFINITY forum, thread "llama.cpp on the Internet Computer", 2024–2026.
  • J. Aerni, S. Fluck, D. Becker, "On-Chain LLM Inference Under Instruction Budgets", Zenodo, July 2026, doi 10.5281/zenodo.20607598; DFINITY forum, 21–22 July 2026.
  • DFINITY forum, "Introducing the LLM Canister", February–August 2025; repository caffeinelabs/llm; ICP documentation "AI inference", "Resource limits", "Cycle costs", "Security model", "VetKeys", read on 3 October 2026.
  • DFINITY forum, "Upcoming proposal: the first TEE-enabled subnet", February–September 2026.
  • OpenRouter, public model list, 3 October 2026.