DEV Community

Tania Chakraborty
Tania Chakraborty

Posted on

Blind-testing models with Netlify AI Gateway

In the latest episode of Netlify community livestream...

Karthik Puvvada (Head of Community) and I (Senior Community Manager, Programs) dove into Netlify AI Gateway: what it is, models you get access to, along with how to use it in your everyday workflow. We put AI Gateway to the test with a live blind taste test called “Which AI Cooked This?”

The Current Landscape

New models are released weekly, and for builders like me who want to use the best model for whatever I’m building, keeping up is a lot. How can we use the best models without constantly overhauling our existing apps and solutions?

"You can’t be building a new app for every new model." — KP

What is Netlify AI Gateway

Netlify AI Gateway is one possible solution to this rapidly changing AI landscape. Rather than managing multiple subscriptions and keys, AI Gateway acts like a universal travel adapter, allowing builders to connect to more than 220 of the latest frontier and open models from 38 labs. You don't need to open an account, keep a credit balance, or copy an API key for each provider, and the gateway doesn't store your prompts or model outputs.

It's on by default for credit-based plans (Free, Personal, and Pro). You can invoke it by asking Agent Runners in plain English, or in code through the official Anthropic, OpenAI, Gemini, and OpenRouter SDKs. (Docs)

"Imagine having a meta adapter that takes care of plugging into the specific country based on where you are... that's the way to think about AI Gateway." — KP

Practical Applications

One thing I wanted to dig into was the utility of AI Gateway for developers who might not want to engage with every available model. The ability to invoke different models for different tasks makes AI Gateway useful without the complexity of managing multiple services.

"I can actually use Netlify AI Gateway specifically to use the different models for the different tasks." — Tania

KP walked through community use cases: a tool that turns daily Discord threads into a digest, drafting thoughtful DMs to community champions, and sifting through hundreds of NPS survey responses for patterns. I shared one of mine: a personal shrimp-keeping blog I can publish to from my phone. The site uses open models for the grunt work and a frontier model as the reviewer.

Which AI Cooked This?

To show AI Gateway in action, KP built a game-show site with Agent Runners. It runs the same task across four to six hidden models, shows each one's speed and real dollar cost, then reveals which model was which. On a "what you missed" digest, the response we were sure was Claude turned out to be Mistral Large 3, at a fraction of the cost. Mistral won again when chat picked a social task, GPT-OSS 20B and Qwen held their own on a customer-win post, and Claude Haiku 4.5 won the round on rewriting a stiff corporate sentence. Mid-stream, an agent run even added a horse-race animation to the results.

The stream also hit rate limits and a deprecated model, a useful reminder of what working with many models looks like in practice.

Conclusion

The big lesson: every model has strengths its launch announcement won't tell you about. A cheap model might draft as well as the frontier model you're subscribed to already, and AI Gateway makes side-by-side testing possible without new accounts or keys. So you can choose which to bring in as your editor, and which to do all the grunt work.

"The only way to know this is by doing these kind of experiments, right?" — KP

"I can use AI to compare AI." — Tania

Start with the AI Gateway quickstart, and watch the full episode to see the horses race.

Top comments (0)