DEV Community

Cover image for Inkling 975B: The Open-Weights Model Almost Nobody Should Self-Host (2026)
Rohit Raj
Rohit Raj

Posted on • Originally published at rohitraj.tech

Inkling 975B: The Open-Weights Model Almost Nobody Should Self-Host (2026)

Originally published on rohitraj.tech

Thinking Machines released Inkling on July 15, 2026 — 975B params, 41B active, Apache 2.0, 1M context, weights on Hugging Face. Every writeup tells you how to run it. None tells you whether to. The BF16 checkpoint needs 2 TB of VRAM; NVFP4 needs 600 GB. The 8x H200 box they name is an AWS p5en.48xlarge at $63.296/hr — $46,206/month always-on. Against the $4.68/M output API, self-hosting breaks even at 9.87 billion output tokens a month. Here is the VRAM ladder, the real cost math, the July 17 price hike everyone missed, and the quant trap that will eat your agent.


Read the full version with code samples, diagrams, and architecture details: Inkling 975B: The Open-Weights Model Almost Nobody Should Self-Host (2026)

More engineering notes: rohitraj.tech/en/notes

Top comments (0)