The clean way to compare inference providers was to run the same open model on both and measure. Llama 3.3 70B was the standard yardstick: available on Groq, Cerebras, Together, everywhere.
That methodology quietly died. On Groq, llama-3.3-70b-versatile (and 3.1-8b) now carry an Enterprise badge with "Contact Sales" where the price used to be — not callable on a self-serve key. So a like-for-like Llama benchmark against Groq is no longer something you can reproduce.
What's left: benchmark on openai/gpt-oss-120b, which is on Groq's free tier and on most competitors, or accept that you're comparing different models and say so. Any speed comparison that doesn't name the exact model per provider is now unfalsifiable.
My last reproducible run, with the script and the model substitutions: https://toolfreebie.com/groq-vs-cerebras-vs-gemini/
Top comments (0)