DEV Community

Cover image for My Google AI API 500 errors stopped being scary when I stopped retrying the whole workflow
Lars Winstand
Lars Winstand

Posted on Originally published at standardcompute.com

My Google AI API 500 errors stopped being scary when I stopped retrying the whole workflow

I knew something was wrong when one flaky Gemini call turned into:

  • 3 duplicate HubSpot updates
  • 1 duplicate Jira issue
  • 4 Slack messages
  • and a very annoyed ops channel

At first it looked like a normal "Google AI API is throwing random 500s" problem.

It wasn’t.

The real bug was that we were retrying the entire automation instead of retrying the model call.

That distinction matters a lot once your workflow has side effects.

If you’re seeing intermittent 500, 502, 503, or 429 errors from Gemini or Vertex AI, the fix usually is not “add more retries everywhere.” The fix is to move the retry boundary.

The first thing that surprised me: Google was already retrying

If you use the Gemini Python SDK, Google already retries transient failures by default.

That includes transient 429 and 5xx responses, with exponential backoff.

So if your production automation is still blowing up after “we added retries,” one of these is probably true:

  1. You’re using raw HTTP or an automation HTTP node, so you don’t actually have safe retry behavior.
  2. You wrapped the entire workflow in retry logic, which replays every side effect.

That second one is the expensive mistake.

If your workflow does this:

  1. fetch lead
  2. build prompt
  3. call Gemini
  4. write to PostgreSQL
  5. update HubSpot
  6. send Slack message
  7. retry whole thing on failure

...then one transient model failure becomes a duplicate generator.

The model error is annoying.

The replay damage is worse.

What a “random 500” usually means in practice

A lot of teams treat 500 as "Google is broken."

Sometimes that’s true.

But with Gemini and Vertex AI, 500 can also mean overload, dependency failures, shared-capacity pressure, or quota-related behavior that doesn’t show up as a clean 429.

That’s why these incidents feel spooky in production:

  • one run succeeds
  • the next one fails
  • the retry half-works
  • then downstream systems are in a weird state

If you only watch for explicit rate limits, you miss the actual pattern.

Your "random 500s" may really be burst traffic, project-wide contention, or spend throttling wearing a different mask.

My opinionated take: workflow-level retry is lazy engineering

If a workflow has side effects, retrying the whole thing is the wrong default.

The retry boundary should sit around the model call, not around everything before and after it.

This is the rule I trust now:

  • collect data once
  • call the model in an isolated step
  • write side effects once
  • replay only the failed step

If you don’t do that, you get the usual mess:

  • duplicate HubSpot or Salesforce writes
  • duplicate Gmail or SendGrid emails
  • duplicate Slack or Discord alerts
  • duplicate Zendesk or Jira tickets
  • half-finished database state

At that point it’s not really an LLM problem anymore.

It’s a workflow design problem.

The pattern that fixed it for us

The cleanest version of this is to isolate the LLM call into its own sub-workflow or worker.

Instead of this:

  1. fetch records
  2. build prompt
  3. call Gemini
  4. update CRM
  5. send Slack alert
  6. retry whole workflow if anything fails

Do this instead:

  1. fetch records
  2. persist a replay-safe payload
  3. call a sub-workflow or worker that only handles Gemini
  4. return structured output
  5. perform downstream writes once
  6. route failures to an error handler

That one change removes most of the blast radius.

If Gemini throws a transient 5xx, only the model step retries.

Your CRM write doesn’t happen twice.

Your Slack alert doesn’t fire twice.

Your upstream data fetch doesn’t get repeated for no reason.

A concrete n8n version

n8n is actually pretty good at this if you use the primitives it gives you.

Useful pieces:

  • Error Workflows
  • Error Trigger
  • sub-workflows
  • execution data
  • log streaming

A decent shape looks like this:

Main workflow

  1. Fetch records
  2. Normalize payload
  3. Save payload + execution ID
  4. Call sub-workflow for Gemini
  5. Receive structured result
  6. Write downstream side effects once

Error workflow

  1. Start with Error Trigger
  2. Capture failed execution ID
  3. Log prompt + payload
  4. Alert Slack
  5. Queue replay of only the Gemini step
  6. Escalate after max attempts

That turns a noisy crash into something debuggable.

If you use raw HTTP, you own the retry logic

This is where a lot of teams accidentally create their own outage.

If you call Gemini through direct REST, an n8n HTTP Request node, Make, Zapier, or a custom worker, you need to implement retry policy yourself.

Minimum bar:

retry_on = [408, 429, 500, 502, 503, 504]
max_attempts = 4
base_delay_seconds = 1
max_delay_seconds = 60
use_jitter = True
Enter fullscreen mode Exit fullscreen mode

And the retry should wrap only the model request.

Not the whole business process.

Example: isolated retry wrapper in Python

import random
import time
import requests

RETRY_ON = {408, 429, 500, 502, 503, 504}
MAX_ATTEMPTS = 4
BASE_DELAY = 1
MAX_DELAY = 60


def call_gemini_with_retry(url, headers, payload):
    attempt = 0

    while attempt < MAX_ATTEMPTS:
        attempt += 1
        response = requests.post(url, headers=headers, json=payload, timeout=60)

        if response.status_code < 400:
            return response.json()

        if response.status_code not in RETRY_ON:
            response.raise_for_status()

        if attempt == MAX_ATTEMPTS:
            response.raise_for_status()

        delay = min(BASE_DELAY * (2 ** (attempt - 1)), MAX_DELAY)
        jitter = random.uniform(0, delay * 0.25)
        time.sleep(delay + jitter)
Enter fullscreen mode Exit fullscreen mode

That’s still not enough by itself.

You also need idempotency around whatever happens after the model returns.

The retry boundary I trust now

This is the design I’d recommend to anyone running AI automations in production.

1. Make the model call idempotent

Give each model request a stable operation ID tied to the business event.

Not the execution attempt.

For example:

lead_enrichment:hubspot_contact_12345
support_triage:zendesk_ticket_98765
invoice_review:invoice_2026_00412
Enter fullscreen mode Exit fullscreen mode

If the same job replays, your system should recognize it as the same operation.

2. Persist enough context to replay one step

Store:

  • prompt input
  • normalized source data
  • model name
  • temperature
  • candidate count
  • execution ID
  • business object ID
  • timestamp

If you don’t log the exact request shape, replay becomes guesswork.

3. Cap retries and fail deliberately

Do not let a worker spin forever because one model is having a bad hour.

After max attempts, route to one of these:

  • dead-letter queue
  • delayed retry
  • human review
  • fallback model path

Fallback routing is not cheating.

It’s production engineering.

4. Smooth traffic before blaming the model

A lot of “random instability” is really bursty traffic.

If your cron job wakes up and slams Gemini with a huge batch, shared-capacity systems can get weird fast.

Paced workers beat spiky workers.

Queues beat bursts.

Are you debugging one API key when the whole project is overloaded?

This is another easy trap.

Gemini and Vertex AI limits are not always about a single request or a single API key.

They can be project-wide.

So if you have:

  • an n8n workflow
  • a Make scenario
  • a Python worker
  • a Zapier automation
  • maybe a background agent

...all hitting the same Google project, failures can look random unless you correlate them with project-wide traffic.

That means you should track at least:

  • requests per minute
  • tokens per minute
  • requests per day
  • spend spikes
  • retry volume
  • error rate by status code

If you only inspect one failing execution, you’ll miss the real cause.

SDK vs raw HTTP vs managed routing

This is one of those boring implementation details that decides whether your week stays calm.

Option What happens when Gemini gets flaky
Gemini API via official SDK Safer defaults. Built-in transient retry behavior. Less custom work.
Gemini API via direct REST or n8n HTTP Request You own retries, jitter, caps, and safe replay boundaries. Easier to get wrong.
Vertex AI pay-as-you-go Shared-capacity behavior means burst shape matters a lot.
Vertex AI Provisioned Throughput Better when you need more consistent service and retries alone aren’t enough.

My bias: if you’re doing direct HTTP in production, be honest that you’re taking on reliability work.

That’s fine.

Just don’t pretend it’s the same as using an SDK with sane defaults.

Terminal commands that are actually useful during incident response

If you’re testing Vertex AI auth manually:

gcloud auth print-access-token
Enter fullscreen mode Exit fullscreen mode

If you want to inspect whether your worker is replaying too aggressively, log attempt counts explicitly:

grep "gemini_attempt" app.log | tail -100
Enter fullscreen mode Exit fullscreen mode

And if you aren’t logging operation IDs yet, fix that first.

A lot of debugging pain disappears once you can answer this question quickly:

Did the model fail once, or did our workflow replay the same business event four times?

What changed for me after fixing the boundary

Once we stopped retrying the whole workflow, the incidents got much less dramatic.

We still saw transient model failures.

That part never fully goes away.

But the failures became contained:

  • one failed model step
  • one replayable unit
  • one error record
  • zero duplicate side effects

That’s a very different operational story.

Where Standard Compute fits if you’re tired of this game

A lot of teams end up here because per-token pricing makes them afraid to add the reliability layers they actually need.

They avoid extra retries.
They avoid fallback models.
They avoid always-on agents.
They avoid richer automation because every failure path has a billing consequence.

That’s exactly the problem Standard Compute is trying to remove.

Standard Compute gives you an OpenAI-compatible API with unlimited AI compute at a flat monthly price, so you can run agents, retries, batching, and automations without token anxiety.

If you’re building on n8n, Make, Zapier, OpenClaw, or custom workers, that matters more than people admit.

Predictable cost changes architecture decisions.

It’s a lot easier to build safe retry boundaries and fallback paths when every extra call doesn’t feel like a tiny financial penalty.

The takeaway I wish I had earlier

If your Google AI API 500 errors keep showing up in production, stop asking only:

how do we retry harder?

Ask better questions:

  • Is the model call isolated from side effects?
  • Can I replay only the Gemini step?
  • Am I using the official SDK or raw HTTP?
  • Am I tracking project-wide traffic instead of staring at one failed request?
  • Do I need better throughput guarantees instead of more retry code?

When an LLM stops responding, the winning move usually isn’t prompt magic.

It’s boring architecture:

  • idempotency keys
  • isolated model steps
  • capped retries
  • jitter
  • sub-workflows
  • dead-letter queues
  • clean error lanes

Less exciting than blaming Gemini.

Much more effective.

Top comments (1)

Collapse
 
devanti profile image
DEV ANTI •

wооow

‍‍‌‌