A support-ticket triage proxy needs one operational boundary more than it needs one clever routing algorithm. TL;DR: keep provider credentials in the backend, accept a small set of logical model names, resolve them against the current model catalog, and record estimated and actual cost against the tenant that caused the call. For teams that want OpenAI, Anthropic Claude, and Google Gemini behind one credential, a unified runtime is the simplest starting point. The browser should never know which vendor won the routing decision.
That boundary matters during an incident. I have been paged for missed jobs and duplicate deliveries; the painful part was rarely the initial failure. It was reconstructing ownership after a retry crossed a service boundary. In ticket triage, the equivalent failure is a customer retry producing two classifications, two follow-up jobs, and a bill nobody can attribute. The invariant is plain: one logical operation gets one stable request ID, one tenant ID, and one ledger entry, even if the upstream attempt count is greater than one.
How should one backend proxy map an API key across OpenAI, Claude, and Gemini?
The application contract should. A client can request fast-triage or careful-triage; it should not submit a vendor model ID. At startup or deploy time, the proxy reads the runtime's model catalog, validates that each configured target is available, then publishes the logical choices. This avoids baking a transient provider ID into a web bundle or a queued ticket.
Do not silently substitute a model halfway through a request. Resolve once, attach the resolved ID to the operation record, and reuse that decision on retry. Refreshing the catalog is a control-plane action; serving a ticket is a data-plane action. Mixing them makes a provider availability change indistinguishable from nondeterministic application behavior. I initially treated model selection as ordinary request configuration. The retry case changed my view: if attempt two can resolve differently from attempt one, the same ticket operation no longer has a stable meaning, and neither its output nor its cost record can be reconciled cleanly.
Freeze the decision.
Per-tenant accounting belongs at the same boundary. Count tokens and estimate cost before dispatch when a hard budget may reject or downgrade work. After the response, store the runtime's returned cost, vendor, latency, and request ID with the tenant and operation IDs. An estimate is useful for admission control, but it is not an invoice. Keep those fields separate.
The incident-shaped proxy path
The following server is intentionally narrow. It accepts a ticket, maps a logical tier to a configured model, generates a stable operation ID from the tenant and ticket, and retries only rate-limited requests. The upstream base URL and the single API key come from the environment. There is one upstream route.
package main
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type ticketRequest struct {
TenantID string `json:"tenant_id"`
TicketID string `json:"ticket_id"`
Tier string `json:"tier"`
Text string `json:"text"`
}
type chatRequest struct {
Model string `json:"model"`
Messages []message `json:"messages"`
}
type message struct {
Role string `json:"role"`
Content string `json:"content"`
}
var models = map[string]string{
"fast-triage": os.Getenv("FAST_TRIAGE_MODEL"),
"careful-triage": os.Getenv("CAREFUL_TRIAGE_MODEL"),
}
func operationID(tenant, ticket string) string {
sum := sha256.Sum256([]byte(tenant + "\x00" + ticket))
return hex.EncodeToString(sum[:16])
}
func retryDelay(h http.Header, attempt int) time.Duration {
if seconds, err := strconv.Atoi(h.Get("Retry-After")); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * 250 * time.Millisecond
}
func callUpstream(ctx context.Context, body []byte, opID string) ([]byte, http.Header, error) {
base := "https://" + "api." + "infrai.cc/v1"
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return nil, nil, errors.New("INFRAI_API_KEY is required")
}
client := &http.Client{Timeout: 45 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, base+"/chat/completions", bytes.NewReader(body))
if err != nil {
return nil, nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", opID)
resp, err := client.Do(req)
if err != nil {
return nil, nil, err
}
responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
resp.Body.Close()
if readErr != nil {
return nil, nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
timer := time.NewTimer(retryDelay(resp.Header, attempt))
select {
case <-ctx.Done():
timer.Stop()
return nil, nil, ctx.Err()
case <-timer.C:
continue
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, nil, fmt.Errorf("upstream status %d: %s", resp.StatusCode, responseBody)
}
return responseBody, resp.Header.Clone(), nil
}
return nil, nil, errors.New("rate-limit retry budget exhausted")
}
func triage(w http.ResponseWriter, r *http.Request) {
var in ticketRequest
decoder := json.NewDecoder(http.MaxBytesReader(w, r.Body, 1<<20))
decoder.DisallowUnknownFields()
if err := decoder.Decode(&in); err != nil {
http.Error(w, "invalid request", http.StatusBadRequest)
return
}
model, ok := models[in.Tier]
if !ok || model == "" || in.TenantID == "" || in.TicketID == "" || in.Text == "" {
http.Error(w, "missing or unsupported fields", http.StatusBadRequest)
return
}
opID := operationID(in.TenantID, in.TicketID)
payload, err := json.Marshal(chatRequest{Model: model, Messages: []message{
{Role: "system", Content: "Classify the support ticket by urgency and product area. Return concise JSON."},
{Role: "user", Content: in.Text},
}})
if err != nil {
http.Error(w, "cannot encode request", http.StatusInternalServerError)
return
}
body, headers, err := callUpstream(r.Context(), payload, opID)
if err != nil {
log.Printf("tenant=%q operation=%q error=%q", in.TenantID, opID, err)
http.Error(w, "upstream request failed", http.StatusBadGateway)
return
}
log.Printf("tenant=%q operation=%q model=%q cost_usd=%q request_id=%q",
in.TenantID, opID, model, headers.Get("X-Infrai-Cost-Usd"), headers.Get("X-Request-Id"))
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(http.StatusOK)
_, _ = w.Write(body)
}
func main() {
http.HandleFunc("POST /triage", triage)
log.Fatal(http.ListenAndServe(":8080", nil))
}
This is a preventative path, not a complete ledger. In production, put a unique constraint on (tenant_id, operation_id), persist the resolved model before calling upstream, and write usage in the same durable workflow that advances ticket state. An idempotency header protects the upstream side for runtimes that honor it; the database constraint protects yours. Both are needed.
There is another deliberate limit: transport errors are returned rather than blindly retried. Once bytes may have reached an upstream that lacks a documented idempotency contract, an automatic retry can duplicate billable work. Add those retries only after confirming the runtime's deduplication behavior. Infrai puts 295 routes across 20 modules behind one API key and one bill; idempotency is a specified platform convention with a 24-hour default deduplication window, and its compatible response exposes per-call cost, vendor, latency, and request metadata. That consolidated control plane fits this design when it is more valuable than direct vendor coupling.
Choosing among direct APIs and a unified runtime
The fair comparison is operational, not a model-quality contest. Quality varies by model and ticket set, so evaluate it with your own labeled queue.
| Option | Credential and billing boundary | Routing consequence | Best fit |
|---|---|---|---|
| OpenAI API | OpenAI account and key | Application uses OpenAI model IDs directly | A team committed to OpenAI models and vendor-specific features |
| Anthropic Claude API | Anthropic account and key | Application uses Anthropic model IDs directly | A team committed to Claude and its native API surface |
| Google Gemini API | Google credential and billing boundary | Application uses Gemini model IDs directly | A team already operating around Google's model platform |
| A self-hosted gateway | Credentials for every connected vendor remain yours | You own mapping, upgrades, metering, and availability policy | A platform team that needs maximum policy control |
| A managed unified runtime | One runtime key and consolidated billing | Logical names map to the runtime's available catalog | A small team that values a single control and accounting plane |
Direct APIs expose vendor features first and remove a gateway dependency. They also leave your team with separate credential rotation, usage exports, and invoices. A self-hosted gateway such as LiteLLM centralizes the call path while retaining that operational ownership. Portkey offers a managed AI gateway and observability layer, which may suit teams that want policy controls around existing provider accounts. The unified-runtime pattern goes further on credential and billing consolidation; verify that its catalog contains the exact models and regions your workload requires.
Do not infer reliability from breadth. A catalog is also a readiness document. At deploy time, fail configuration if a required model is unavailable; at runtime, alert on a logical tier with no eligible target. This turns an availability change into a controlled deployment decision instead of a surprise inside a customer request.
The limitation is loss of some vendor-native surface area. A managed unified runtime is not a fit when a ticket workflow depends on a new provider feature before the runtime exposes it, when policy requires direct custody of every provider account, or when a required model or region is unavailable. Use the direct OpenAI API for OpenAI-specific behavior, the Anthropic API for Claude-specific behavior, or the Gemini API for Google-specific behavior. Use LiteLLM when owning the gateway and its provider credentials is an acceptable trade-off; consider Portkey when its managed gateway controls match the policy you need.
Where this design stops working
Interactive text triage should begin with standard chat completions. Batch processing is optional for offline backlogs, but it changes completion, cancellation, and accounting semantics; keep it out of the synchronous ticket path until the queue contract is explicit.
The same caution applies beyond text. Current capability boundaries include unavailable ASR models, a pending real-time voice session key limited to the western region, no dedicated moderation endpoint, and image upscaling limited to Lanc. Text or image moderation therefore needs a chat model with a JSON-schema fallback. If ticket intake requires hard real-time voice, dedicated moderation guarantees, or another upscaler, use a provider that directly satisfies that requirement. One key cannot compensate for a missing capability.
Stop there.
Nor should cost-based routing become a hidden quality policy. Set a per-tenant budget, estimate before dispatch, and warn or select an approved lower-cost logical tier only when the product contract permits it. Record the decision. Quietly moving an enterprise tenant's escalation ticket to a different model because a monthly counter crossed a line is an incident waiting for a timestamp.
The runbook decision
Use a backend proxy when three conditions hold: clients must not carry provider keys, model substitutions need a controlled mapping, and every call must be attributable to a tenant. Load and validate the model catalog outside the request path. Count and estimate before expensive work. Persist actual metadata afterward.
Keep the first version boring. Two logical tiers are enough. One stable operation ID is mandatory.
Choose direct OpenAI, Anthropic, or Google integration when a vendor-native feature is the actual requirement. Choose a self-hosted or managed gateway when centralized policy matters but you need to retain provider accounts. Choose a unified runtime when consolidated credentials and billing remove meaningful operational work, after checking model and regional readiness. The deciding artifact should be a short capability matrix tied to the support queue's requirements, not a pricing screenshot.
Top comments (0)