Articles · Local infrastructuredraft
Two weeks local. ThePrimeagen and artificial intelligence at home
Published · by Dario Solari · 6 min read
One of the internet's best-known programmers, a former Netflix engineer, decided in January 2025 that offline models were "the way to go". He bought hardware, ran DeepSeek on six borrowed graphics cards and wrote his own assistant in four days. Then he stopped. Why he stopped matters to every company.
Michael Paulson is ThePrimeagen: former Netflix engineer, programmer, streamer. His YouTube channels together pass one and a half million subscribers, and his audience is mostly developers. He has also been, for years, one of the loudest sceptics of artificial intelligence applied to code. That is exactly why his story with AI at home is worth telling: if anyone has the skills, the means and the motivation to do it himself, it is him.
January 2025: "offline models are the way to go"
On 27 January 2025 the world discovers DeepSeek R1, an open Chinese model that can be downloaded and run on your own machines. Paulson publishes I Think I Love Deepseek R1: "You could spend $1,000, get a good enough piece of hardware to be able to run something small and local. It can be completely offline and you can have it just for you, without having to worry about all the other stuff, the data collection, all that." And then: "This is why I'm buying a couple of those Mac minis. I'm going to try to build my own little Mac Mini farm."
The reasoning applies to companies too, and he says so himself: "Even more so if you're at a company that doesn't allow AI, you could actually have your own offline one." He does not trust DeepSeek, either: "That's why I want to run it offline. I don't want to give it access to the internet because I don't know if it calls home or not."
Two days later he is even firmer: "Nothing is more important than having your own offline models. I am becoming more and more convinced that offline models are the way to go: your own little Raspberry Pi, Mac minis, GPUs going. That's the only way to not have to rely on these larger companies." And he announces: "I have boughten thousands of dollars of Mac Mini. I'm going to also buy a whole bunch of Raspberry Pis."
February 2025: the graphics card wall
On 2 February he goes live to buy what he needs. The video is called Why Buying GPUs Is a Disaster. "I genuinely thought I was going to come on today, buy a couple cards, buy a motherboard, buy some RAM, get a little bit of power supply, get a case, call it a day. I didn't know how wrong I was." RTX 4090s are only on eBay, mostly from Hong Kong and China, at around three thousand dollars each. The new 5090s sold out on the day they went on sale. He falls back on plan B: "I'm just going to try to get some 3090s, that's probably just the easiest right now." And he sees the next problem coming: "I'd get stopped by power long before. I'd have to go rent a professional place just to figure the power situation out."
Two days later he manages it, on someone else's machine. George Hotz, the founder of tinygrad, lends him a server with six RTX 4090s. "I can run DeepSeek R1, the full model, at a 1.5 quant, and I get about 10ish tokens per second on six 4090s. Thank you George Hotz." Ten tokens a second is roughly a person's reading speed: it works, but it is far from the ChatGPT experience.
Four days of code
Between 1 and 6 February he writes cockpit: an autocomplete for Neovim, his editor, in the style of GitHub Copilot but entirely local. The code is public. A small Go server takes the editor's requests and spreads them across several copies of llama.cpp, one per graphics card. Thirty commits in six days, with messages that tell the story better than any comment: "basic autocomplete finished", "slightly good server for spawning many llamas", "i hate models sometimes".
On 8 February he presents it in a video. One sentence at the start says a lot: "I've been using the shared tinybox George Hotz let me have access to, and as you can see it's currently being completely used, 100%, so I don't even have access to the hardware while I'm recording this video." He calls it "all proof of concept". After 6 February the project receives no further changes.
And then the cloud
From then on, everything Paulson builds with artificial intelligence runs on cloud models. In November 2025 he releases 99, a Neovim assistant that almost five thousand developers have starred on GitHub: by default it uses Claude, and it has no option for local models. "It's actually using a big model," he explains in the video. His 2026 agents run on Cursor's infrastructure and on paid models. In September 2026 David Heinemeier Hansson, the creator of Ruby on Rails, asks him to lead automated testing for Omarchy, his Linux distribution. Paulson built the tool; the models are hosted by others.
The Mac Mini farm, the Raspberry Pis and the 3090s appear in no later video. In March 2026, commenting on PewDiePie's ten-GPU system, he says he had extra power brought to the room he works in: "Maybe I should think about building my own GPU rack. I don't know, bro. I don't know if I want all of this."
The diagnosis, his own
He had already explained why on 4 February 2025, at the height of his enthusiasm: "Not a lot of people are going to run their own LLMs. It's expensive, it's difficult, uptime's not going to be easy, it's practically impossible to find GPUs." In another episode of his podcast he adds: "I think the money is in the integration. It's not necessarily in the AI." Sooner or later powerful models will run locally, and then "it's going to come down to who has the best integration".
He has not changed his mind about data. Commenting on PewDiePie in November 2025, he describes ChatGPT as "one of the single most powerful ad networks that are going to be built ever. Not only do they know what you've asked, they know what you've dwelled on." The conviction stayed. The work of keeping the machines running did not.
What it teaches a company
Paulson's story does not say that AI at home does not work: in two weeks, on borrowed hardware, he made it work. It says that making it work once is the easy part. The rest is buying the right hardware when it is scarce, sizing power and cooling, keeping the servers on when they are needed, updating the models, and connecting them to the tools people actually use. Those are exactly the items in his diagnosis: cost, difficulty, uptime, integration.
For a programmer working alone, giving up is a reasonable choice: his code is not a professional secret and his data is mostly his own. For a law firm, an accountant or a company with clients' projects in its archives, the sums are different. The reasons for keeping data in-house are the ones Paulson had in January 2025, only weightier. What makes the difference is who takes care of the rest.