Articles · Local infrastructuredraft
Large models on the desk. What Framework shows with Gorgon Halo, and what it doesn't
Published · by Dario Solari · 3 min read
Framework has presented on video its Desktop with 192 gigabytes of memory and the new AMD processor. It runs on a single machine models that until yesterday needed two connected computers. For an office working with long documents, though, the right question is a different one.
On 30 September Framework, the American company known for laptops you can repair and upgrade part by part, published on its YouTube channel a quarter-hour video about the new version of the Framework Desktop. It is a small computer, little bigger than a shoebox, with AMD's new Ryzen AI Max+ PRO 495 processor, codenamed Gorgon Halo, and 192 gigabytes of memory shared between processor and graphics. The video is the manufacturer's, and should be read as such: it shows the machine where it performs best. That said, it is carefully made, with real models and real figures, and it is a good starting point for seeing where hardware for AI at home stands.
What it shows
The news is the memory. The previous version, with the Strix Halo processor, stopped at 128 gigabytes: enough for many open models, not for the largest. To run a model like DeepSeek V4.1 Flash you had to connect two machines, or read part of the model from disk, very slowly. With 192 gigabytes the whole model fits in memory on one computer. The video shows it with three recent open models, DeepSeek V4.1 Flash, Xiaomi's MiMo V2.6 Flash and GLM-5.3 Flash, compressed to fit in memory and run with free software such as llama.cpp.
The speed at which they write their answers is roughly the speed at which a person reads: not instant, but usable. And it holds steady as the conversation grows.
The second part is the more interesting one. With all that memory you can keep two models running together: a fast one that does most of the work and, when it gets stuck, asks a slower and more capable one for help. It is a pattern we will see more and more: not one huge model for everything, but different models for different tasks, handing work to each other.
Finally a practical detail: the expansion slot is open at the end, so you can add a separate graphics card to share the work, or a fast network card to join several machines into a small cluster.
What it doesn't show
The video is made for one person at a desk, or for jobs left running overnight. It says nothing about what happens when five or ten people ask questions at the same moment, nothing about power draw or noise, and it does not compare the machine with the alternatives, Apple's Mac Studio and NVIDIA's DGX Spark.
There is also something the video measures but does not comment on. Before answering, a model has to read everything it is given: the question and, if it works on the company's documents, the relevant pages of contracts, files, manuals. On this machine, reading long texts is the slow part. With a few lines the answer starts at once; with dozens of pages of context the wait before the first word is measured in minutes. It is not a flaw of Framework's: it applies, to different degrees, to every machine that bets on large memory, Macs included. NVIDIA cards, at a similar price, read much faster and cope better with several people at once.
Who it makes sense for
For one person or a small group who want to run the largest open models in-house, with short questions or jobs that can wait, the Framework Gorgon Halo is today one of the most interesting options: a single machine, quiet enough for a desk, open, running Linux and not tied to one vendor. For a professional firm where several people query long archives every day, the size of the model matters less than how fast documents are read and how many requests run at once. There the choice is usually a different one.
That is the question to ask before any purchase: not "what is the largest model I can run", but "how many people will use it, with how many documents at a time, and how long can they wait". The answer decides the machine, and often saves a good deal of money.
Framework's video: 192GB Framework Desktop: Local AI on Gorgon Halo.