A 125-billion parameter model that usually demands a server farm just moved into your gaming rig.
Meet Strata, an open source project that runs Qwen3.8-Flash-Next on a regular PC, Windows or Linux, with nothing more exotic than an NVIDIA or AMD GPU packing 12 GB of VRAM or more.
The headline number is the speed: around 100 tokens per second. Think of a token as roughly three quarters of a word, so the model is generating text faster than most people can read it.
And here's the part that actually matters: nothing leaves your machine. No cloud round trip, no server dependency, everything runs locally.
This isn't a toy chatbot either. It writes code, reads images, and plugs into the coding agents and apps you already use daily.
This is what happens when the open source community decides frontier scale AI shouldn't be locked behind data center walls. Not a replacement for big infrastructure, a redistribution of power.
The real question isn't whether this is possible anymore. It's how soon even bigger models follow the same path home.
🔗 Original Source & Reference: https://github.com/Niko1221/Strata
Published automatically via FeedMind AI Content Pipeline.

Top comments (0)