Higher-end AI PCs landed this month at prices that make people ask the wrong question first: which one is worth the money?
The number that actually decides what you can run is not the price, the core count, or the TOPS figure. It is how much unified memory you can give the model, and how much of it survives the operating system.
Here is the sizing table I wish more hardware pages published.
What each model class actually needs
- Small models (7–8B parameters): about 5 GB at 4-bit quantisation
- Mid-size models (30B): about 15–20 GB at 4-bit, before you add context
- Large models (70B): about 35–40 GB at 4-bit, plus context
Context is the part people forget. A long conversation, a big system prompt, a code repository in the window — all of it lives in the same memory as the weights. Budget for it or you will meet the wall mid-session.
The unified-memory catch
On these machines the CPU and GPU share one pool. A 128 GB machine was demonstrated reporting roughly 110 GB of GPU-addressable memory — around 79.9 GB dedicated plus 30 GB shared. That is not the full 128 GB, and it is the honest number to plan against.
The practical consequence: 24 GB gets tight as soon as the operating system takes its share. 64 GB and 128 GB configurations are not marketing padding — they are the difference between running a 70B model and not.
The three questions to ask before spending
- Which models do you actually intend to run? Pick the largest one honestly, then size from that table.
- How much context will you really use? A code assistant with a repo in context needs multiples of the bare weight figure.
- Do you need to own it at all? Renting inference in the cloud costs nothing up front and reaches models far beyond what any of these machines hold. Owning wins on privacy, latency, offline work and long-run economics — not on raw capability.
Where this leaves the price tags
The current generation runs from $2,599 to $5,999. Spread across that range, most of the gap between the cheapest and most expensive configuration is a memory decision — 24 GB at the bottom, 128 GB at the top. Paying more for cores while staying at 24 GB buys you a machine that still cannot hold the model you wanted.
So: pick the model, derive the memory, then compare prices. Doing it in that order usually reduces the bill.
Originally published on SelfHost Pilot, where I test self-hosted software and AI stacks on real hardware and publish the measured numbers.
If you want to see what is actually rising in self-hosted software — ranked by commit momentum rather than stars — I maintain a tracker with 1,700+ projects: selfhosted-tracker.
Top comments (0)