DEV Community

Cover image for Retrying AI API Calls in Go Without Hiding Failures
maoren8412
maoren8412

Posted on

Retrying AI API Calls in Go Without Hiding Failures

A retry loop can keep an AI integration alive through a brief upstream outage. It can also multiply load, repeat a request that should not be repeated, and make an incident harder to diagnose.

For Go clients, keep the policy explicit: bound the total time with a context, retry only transient failures, add jitter, and switch models only when the error is plausibly upstream availability. Do not retry authentication errors, invalid requests, or a caller cancellation.

The code below uses the standard library and assumes a non-streaming OpenAI-compatible endpoint. It reads and closes each response body before another attempt, and keeps the request body replayable.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "math/rand/v2"
    "net/http"
    "time"
)

type request struct {
    Model string `json:"model"`
    Input string `json:"input"`
}

func call(ctx context.Context, client *http.Client, endpoint, key string, in request) ([]byte, int, error) {
    payload, err := json.Marshal(in)
    if err != nil {
        return nil, 0, err
    }

    models := []string{in.Model, "backup-model"}
    for _, model := range models {
        for attempt := 0; attempt < 3; attempt++ {
            body := bytes.NewReader(payload)
            req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, body)
            if err != nil {
                return nil, 0, err
            }
            req.Header.Set("Authorization", "Bearer "+key)
            req.Header.Set("Content-Type", "application/json")
            var data request
            if err := json.Unmarshal(payload, &data); err != nil {
                return nil, 0, err
            }
            data.Model = model
            req.Body = io.NopCloser(bytes.NewReader(mustJSON(data)))

            resp, err := client.Do(req)
            if err != nil {
                if ctx.Err() != nil {
                    return nil, 0, ctx.Err()
                }
            } else {
                b, readErr := io.ReadAll(io.LimitReader(resp.Body, 2<<20))
                resp.Body.Close()
                if readErr != nil {
                    return nil, resp.StatusCode, readErr
                }
                if resp.StatusCode < 500 || resp.StatusCode > 504 {
                    return b, resp.StatusCode, nil
                }
            }
            if attempt < 2 {
                delay := time.Duration(150*(1<<attempt))*time.Millisecond + time.Duration(rand.IntN(150))*time.Millisecond
                t := time.NewTimer(delay)
                select {
                case <-ctx.Done():
                    t.Stop()
                    return nil, 0, ctx.Err()
                case <-t.C:
                }
            }
        }
    }
    return nil, 0, fmt.Errorf("all models failed with transient upstream errors")
}

func mustJSON(v any) []byte { b, _ := json.Marshal(v); return b }

func main() {
    client := &http.Client{Timeout: 20 * time.Second}
    ctx, cancel := context.WithTimeout(context.Background(), 25*time.Second)
    defer cancel()
    _, _, _ = call(ctx, client, "https://api.example/v1/responses", "replace-me", request{Model: "primary", Input: "status"})
}
Enter fullscreen mode Exit fullscreen mode

In production, avoid suppressing the final error as main does here: return structured error details to the caller and record model, attempt, status, and elapsed time. Never log the API key or prompt by default. Also respect Retry-After when the provider supplies it, and ensure your retry budget fits inside the user-facing deadline.

One subtlety: a network error does not prove the server failed to process a request. For non-idempotent operations, use a provider-supported idempotency key or avoid automatic replay. Streaming requests need a different lifecycle: close the stream promptly on cancellation and do not transparently replay after output has already been delivered.

This demonstrates configuring resilient multi-provider fallback chains using B-Lost's unified OpenAI-compatible endpoint, ensuring zero-downtime AI services. Fallback reduces the impact of some upstream failures; it does not guarantee zero downtime.

Disclosure: Multi-model API relays and compute for this evaluation are sponsored by b-lost.com — an AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All observations reflect independent developer testing.

Top comments (0)