DEV Community

FairchildBlake8483
FairchildBlake8483

Posted on

Multiple LLM Providers: One API Key With Observable Candidate-Scoring Fallback

A media company puts multiple LLM providers behind one API key for candidate scoring, then the overnight page says its text-classification queue is late. On-call sees 18,420 completed jobs, a normal request-success graph, and no obvious outage. Yet editors opening the hiring dashboard at 08:00 find yesterday's applicants unranked. The requests succeeded; the classifications were unusable.

TL;DR: A single-key, multi-provider gateway can reduce integration sprawl, but it cannot define success for candidate scoring. Page on the age and completeness of valid scoring results, not HTTP success. Require one portable JSON contract, validate it after every model response, record outcome classes separately from providers, and allow fallback only for errors that another provider can safely repair. Evaluate gateways by this operational evidence, not by a routing checkbox or a volatile token price.

The first useful signal should have fired before the deadline: the oldest unscored candidate was aging while the valid-result rate fell. That is the alert-to-action path this article works backward through.

Can one API key safely route multiple LLM providers?

A transport-level success is not a business-level success. For this job, success means that every eligible candidate has one rubric version, a bounded score, allowed tags, and an explanation that can be traced to the input and policy version. A response can arrive on time and still fail that definition because it is malformed, omits a rubric field, uses an obsolete tag, or belongs to an earlier retry.

The one-key design also changes the failure boundary. Application credentials terminate at the gateway, while the gateway selects an upstream model service. That simplifies secret distribution to workers, but it concentrates routing policy, audit data, rate-limit behavior, and provider credentials in one control point. This is its main limitation: it is not suitable when policy requires each worker to authenticate directly to each upstream. In that case, choose direct integrations and accept the extra normalization work. Treat either boundary as production infrastructure, not as a convenience wrapper.

Provider portability is narrower than API compatibility. OpenAI documents Structured Outputs for function calling with strict: true, while gateway documentation describes normalized requests and provider routing. Those mechanisms help, but the application must still own the rubric schema and validate the returned arguments. A portable workload depends on the smallest contract that every selected route can honor.

Keep the score record separate from the raw model response. The durable record should include a candidate identifier, rubric version, input digest, model-route identifier, attempt number, validation outcome, and completion time. The raw response is evidence; it is not the hiring decision.

Work backward from the page

Start with the action the page demands. If the alert cannot tell on-call whether to stop dispatch, replay work, disable a route, or wait for capacity, it is probably a dashboard notification wearing a pager badge.

For an overnight batch with a fixed publishing deadline, use two related service indicators:

  1. Result freshness: age of the oldest eligible candidate without a valid score for the active rubric.
  2. Valid completion: valid, committed scores divided by eligible candidates, measured for the current run.

Request latency and upstream error rate remain diagnostic signals. They should not be the primary page because a fast stream of schema-invalid answers can look healthy there. Queue depth is useful too, although it needs context: a large queue early in the window may be expected; one old item near the deadline is different.

The event trail for a single attempt can stay compact:

{
  "job_id": "score_01J9M6Q2",
  "candidate_id": "cand_18420",
  "rubric_version": "editorial-7",
  "attempt": 2,
  "route": "route_b",
  "outcome": "schema_invalid",
  "input_digest": "sha256:7e8d...",
  "duration_ms": 1840,
  "committed": false
}
Enter fullscreen mode Exit fullscreen mode

Do not put resumes, cover letters, or generated explanations into metric labels. Candidate IDs and job IDs also create dangerous cardinality in a time-series system. Keep identifiers in structured logs or traces; aggregate metrics by bounded dimensions such as rubric version, route, outcome, and queue.

The useful outcome vocabulary is small: valid, transport_error, rate_limited, schema_invalid, policy_rejected, deadline_exceeded, and duplicate. Provider-specific error strings belong in evidence attached to the attempt. Normalizing them into operational classes lets the runbook survive a route change.

Instrument validation before fallback

Fallback is a retry with a different failure domain. It is safe only when the job is idempotent and the first attempt cannot later overwrite the accepted result.

Use a deterministic idempotency key derived from the candidate, rubric version, and input digest. Commit with a uniqueness constraint on that key. A late response may be logged, but it must lose the commit race once a valid result exists. This is the part teams tend to discover after duplicate delivery, and it is much harder to repair after downstream reviewers have acted on two scores.

The following Go sketch keeps routing outside the scoring contract. It validates every response and exposes an outcome that metrics can count.

package scoring

import (
    "context"
    "encoding/json"
    "errors"
    "fmt"
)

type Score struct {
    RubricVersion string   `json:"rubric_version"`
    Total         int      `json:"total"`
    Tags          []string `json:"tags"`
    Explanation   string   `json:"explanation"`
}

type Route interface {
    Classify(ctx context.Context, input []byte) ([]byte, error)
}

func classify(ctx context.Context, routes []Route, input []byte, rubric string) (Score, error) {
    var lastErr error
    for _, route := range routes {
        raw, err := route.Classify(ctx, input)
        if err != nil {
            lastErr = fmt.Errorf("transport: %w", err)
            continue
        }

        var score Score
        if err := json.Unmarshal(raw, &score); err != nil {
            lastErr = fmt.Errorf("schema: %w", err)
            continue
        }
        if score.RubricVersion != rubric || score.Total < 0 || score.Total > 100 {
            lastErr = errors.New("schema: invalid rubric or score range")
            continue
        }
        return score, nil
    }
    return Score{}, lastErr
}
Enter fullscreen mode Exit fullscreen mode

This sketch intentionally leaves out storage and policy. In production, the caller needs an overall deadline, per-attempt time budgets, bounded retries, cancellation, and an atomic commit. It should also validate allowed tags and required explanation rules with a real JSON Schema validator rather than relying only on Go decoding. Unknown fields deserve an explicit policy; silently accepting them can hide contract drift.

Do not fall back on every failure. A transport timeout before the overall deadline may be retryable. Invalid JSON may justify another route if the input is unchanged and the alternative supports the contract. A policy rejection or invalid source document usually needs quarantine or human review, because asking another model does not make the input acceptable. Deadline exhaustion ends the chain.

No infinite retries. Ever.

The alert evidence also provides a better way to compare gateways. Begin with recorded fixtures from the actual media hiring workflow: terse resumes, long portfolios, missing employment dates, multilingual biographies, and candidates whose evidence sits near a rubric boundary. Remove direct identifiers, pin the rubric, and replay the same fixtures against every candidate route. The goal is not to crown a model from a tiny benchmark. It is to learn which failure classes the gateway reveals and which it obscures.

Score the gateway itself on evidence and control:

Decision area Test Evidence to retain
Contract handling Send valid, malformed, truncated, and extra-field responses Raw response, validation result, schema version
Routing Force one route unavailable and one slow Selected route, reason, attempt order, deadlines
Idempotency Deliver a late first response after fallback commits One committed record, duplicate outcome
Portability Move the same fixture set between routes Request transform, output deltas, unsupported fields
Operations Reconstruct one failed job from an alert Correlated event trail without sensitive text in labels

OpenRouter is one example of an aggregator that documents a unified API and provider routing. Direct integrations with OpenAI, Anthropic, and Google's Gemini API represent a different ownership boundary: the application can retain provider-specific controls, but it must build more normalization and routing itself. These are architectural boundaries, not a ranking. A gateway may also be self-hosted behind an internal API. In every case, test the exact contract and failure behavior you will operate.

Price belongs in the evaluation as a measured constraint, not the thesis. Record usage reported for each accepted attempt and all wasted attempts, then compare cost per valid committed score under the same fixture mix. Published unit prices change, and the cheapest successful request can become expensive when schema failures trigger retries or manual review.

Turn the earlier signal into a runbook

The instrumentation change is straightforward: emit one attempt event after transport and validation, increment bounded outcome counters, and update the job's valid-result state only after an idempotent commit. From that state, compute oldest-unscored age and valid completion for the active run. Alert on the result state; use attempt signals to explain it.

The page should carry the run ID, rubric version, oldest-item age, valid/eligible counts, recent outcome distribution, and active route set. The first response is to check whether dispatch is progressing and whether failures cluster by route or outcome. Schema failures call for preserving samples, stopping unsafe fallback if needed, and checking the contract. Transport failures call for capacity and route checks. Duplicates call for protecting the commit path before replay.

Deployment needs the same discipline. Shadow a new route with redacted fixtures, compare validation and scoring distributions, then canary a bounded share of live jobs without allowing shadow results to commit. Pin schema and rubric versions. Roll back routing independently from application code so on-call does not need an emergency deploy to remove a bad path.

There is a cost to sensitive alerts. A freshness threshold set too close to normal queue jitter pages on-call for work that would finish before editors arrive. A completion alert evaluated before eligibility settles will flap as new candidates enter the run. Both train responders to distrust the signal.

Set warning and paging thresholds from the job deadline and the remaining drain time, then test them with delayed-route and invalid-output exercises. The warning should create time to inspect. The page should correspond to a concrete risk of missing the publishing objective. Review false positives alongside misses after every run, because a quiet pager that fires too late is not healthy either.

The final decision rule is plain: choose a gateway boundary only if it preserves your scoring contract, exposes every attempt well enough to operate, and lets you prove one valid commit per candidate. One key is convenient. Evidence is what keeps the schedule.

Further reading

Top comments (0)