DEV Community

Cover image for The Day All Our AI Agents Stopped Working 💀
Frederick Peñalo
Frederick Peñalo

Posted on

The Day All Our AI Agents Stopped Working 💀

Spanish Version Here
September 24th.

Kanvas is right in the middle of a presentation.

The tests from the day before? All successful.

Our PO was happy with the results.

Everything was ready.

I had my protein shake that morning. I was prepared for the presentation.

Then, out of nowhere, I get a call:

— What’s going on with our agents? None of them are working.

Okay...

I check.

He’s right.

Not a single one is working. 💀

What the hell is going on?

The codebase hasn’t changed in the last 24 hours.

Our core is still operational.

The APIs are working.

The business logic is working.

Everything looks fine.

Except for our agents.

So I check our monitoring system and find the problem:

429 RESOURCE_EXHAUSTED

💀

Okay.

After that unnecessarily dramatic introduction, let me explain what actually happened.

Our AI models experienced a usage spike right in the middle of a presentation.

And that forced us to ask ourselves a question we hadn't seriously considered before:

How do we make AI agents highly available?


The problem

Kanvas Agent has become one of our flagship products.

And with Kanvas, we’ve always had something very important:

Control.

Our core is under our control.

Business rules, resources, APIs, infrastructure — all of it goes through systems we manage.

If something fails, we can investigate it.

We can scale it.

We can deploy another instance.

We can change the infrastructure.

We have redundancy, autoscaling, health checks, blue/green deployments, and years of knowledge about building highly available systems.

But now we're living through this new wave of AI.

And AI introduced a new problem.

You can have your infrastructure working perfectly.

Your API can be healthy.

Your database can be healthy.

Your servers can be running without a problem.

Everything can be green.

But if the provider powering your agents stops responding...

Your agents are dead.

When part of your infrastructure depends on a third-party AI provider, there's something you simply don't control.

At some point, all you can do is trust them.

And this time...

our faith wasn't enough. 😂


So... what do we do?

Fortunately, we managed to save the presentation thanks to our PO, Estrella.

But the incident left us with a problem we needed to solve.

Our Lead came up with a simple idea:

If a request to an AI model fails, the request shouldn't die with that model.

There should always be another route.

If Model A isn't available, try Model B.

If the entire provider is having problems, move to Provider B.

And that's when we started implementing something we internally call Routing.

More specifically, one of the main strategies behind our router is fallback.

Our agentic framework can now receive multiple models and multiple providers.

A request might start with our primary model.

If that model fails, the request doesn't die.

The router tries the next available model.

And if the problem affects the entire provider, we can move to another provider completely.

Something like:

Model A ❌ → Model B ❌ → Provider B → Model C ✅

Same mission.

Different model.


What did we learn?

Every production problem leaves you with a lesson.

Ours was pretty simple:

You can't build highly available AI agents while depending on a single model or a single provider.

Today, we can assign multiple models from the same provider and also configure multiple providers with different models.

And that opened another interesting door

Top comments (0)