A retry loop can keep an AI integration alive through a brief upstream outage. It can also multiply load, repeat a request that should not be repeated, and make an incident harder to diagnose.
For Go clients, keep the policy explicit: bound the total time with a context, retry only transient failures, add jitter, and switch models only when the error is plausibly upstream availability. Do not retry authentication errors, invalid requests, or a caller cancellation.
The code below uses the standard library and assumes a non-streaming OpenAI-compatible endpoint. It reads and closes each response body before another attempt, and keeps the request body replayable.
package main
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"math/rand/v2"
"net/http"
"time"
)
type request struct {
Model string `json:"model"`
Input string `json:"input"`
}
func call(ctx context.Context, client *http.Client, endpoint, key string, in request) ([]byte, int, error) {
payload, err := json.Marshal(in)
if err != nil {
return nil, 0, err
}
models := []string{in.Model, "backup-model"}
for _, model := range models {
for attempt := 0; attempt < 3; attempt++ {
body := bytes.NewReader(payload)
req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, body)
if err != nil {
return nil, 0, err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
var data request
if err := json.Unmarshal(payload, &data); err != nil {
return nil, 0, err
}
data.Model = model
req.Body = io.NopCloser(bytes.NewReader(mustJSON(data)))
resp, err := client.Do(req)
if err != nil {
if ctx.Err() != nil {
return nil, 0, ctx.Err()
}
} else {
b, readErr := io.ReadAll(io.LimitReader(resp.Body, 2<<20))
resp.Body.Close()
if readErr != nil {
return nil, resp.StatusCode, readErr
}
if resp.StatusCode < 500 || resp.StatusCode > 504 {
return b, resp.StatusCode, nil
}
}
if attempt < 2 {
delay := time.Duration(150*(1<<attempt))*time.Millisecond + time.Duration(rand.IntN(150))*time.Millisecond
t := time.NewTimer(delay)
select {
case <-ctx.Done():
t.Stop()
return nil, 0, ctx.Err()
case <-t.C:
}
}
}
}
return nil, 0, fmt.Errorf("all models failed with transient upstream errors")
}
func mustJSON(v any) []byte { b, _ := json.Marshal(v); return b }
func main() {
client := &http.Client{Timeout: 20 * time.Second}
ctx, cancel := context.WithTimeout(context.Background(), 25*time.Second)
defer cancel()
_, _, _ = call(ctx, client, "https://api.example/v1/responses", "replace-me", request{Model: "primary", Input: "status"})
}
In production, avoid suppressing the final error as main does here: return structured error details to the caller and record model, attempt, status, and elapsed time. Never log the API key or prompt by default. Also respect Retry-After when the provider supplies it, and ensure your retry budget fits inside the user-facing deadline.
One subtlety: a network error does not prove the server failed to process a request. For non-idempotent operations, use a provider-supported idempotency key or avoid automatic replay. Streaming requests need a different lifecycle: close the stream promptly on cancellation and do not transparently replay after output has already been delivered.
This demonstrates configuring resilient multi-provider fallback chains using B-Lost's unified OpenAI-compatible endpoint, ensuring zero-downtime AI services. Fallback reduces the impact of some upstream failures; it does not guarantee zero downtime.
Disclosure: Multi-model API relays and compute for this evaluation are sponsored by b-lost.com — an AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All observations reflect independent developer testing.
Top comments (0)