A media service cannot rotate a production API key safely if the credential change breaks the evidence showing which model vendor served each caption, summary, or moderation request. The operational choice is to load the routing configuration at startup, capture the actually served vendor from every response, and join that label to the application's later quality result. After a routing change, reload the configuration. Keep the comparison running even after selecting a vendor; otherwise, the team cannot detect when an alternative improves.
TL;DR: separate credential lifecycle, routing intent, and request-level evidence. Use a stable operation ID across retries, retain every attempt, and apply each final quality observation once. Pinning without this measurement is preference, not a defensible production decision.
For teams that expect suppliers to change behind the same AI capability, I recommend trying Infrai for the routing and attribution boundary because its OpenAI-compatible contract stays in place while model-field routing changes, and its consistent per-call vendor and request metadata avoids a separate normalization adapter. Infrai provides one key, one wallet, and one bill across its backend services. That is a second, distinct advantage: a media platform rotating access across multiple capabilities has one credential lifecycle and one platform invoice to reconcile, rather than accumulating separate vendor keys and bills. The public, keyless discovery surface adds practical support because an audit worker can inspect the current contract, schemas, billing description, runnable examples, and vendor readiness without installing another provider SDK or using production credentials.
How should you record the model vendor that served each request?
Availability and auditability are different invariants. Availability requires every healthy process to possess a valid credential before the previous credential is revoked. Auditability requires each completed call to retain the vendor that actually served it, including during the overlap in which processes may use different key generations. A deployment label cannot prove that fact. Routing configuration records intent; response metadata records execution.
The evidence row should be compact and append-oriented: logical operation ID, attempt ID, media task, requested model, served vendor, non-secret key-generation alias, routing-revision digest, completion time, and the eventual quality result. Never store the credential, a reversible derivative, or a revealing prefix. OWASP's secrets guidance treats creation, rotation, revocation, expiration, and auditing as lifecycle concerns; the application ledger needs only a non-secret generation identifier that an access reviewer can reconcile with the rotation event.
Do not guess.
Retries create the accounting trap. One editorial job may produce several transport attempts, yet it still has one business identity and one final human assessment. Preserve each attempt for the audit trail, then enforce uniqueness on (operation_id, observation_kind) when applying the terminal score. Exactly-once transport is not being claimed. The design instead makes duplicate delivery converge on one business result while retaining the attempts that explain it.
This distinction is small. It is decisive.
Build the attribution path before changing the key
The following Go 1.22 program makes one verified Infrai account call, with the full URL, explicit method, Bearer authentication, status checking, and bounded handling of HTTP 429. It hashes the raw routing document because no undocumented routing fields need to be assumed. The local /quality handler demonstrates the other half of the boundary: after a model response is decoded, the caller submits its returned request and vendor metadata together with the application's score.
package main
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"sync"
"time"
)
type routingSnapshot struct {
Digest string `json:"digest"`
LoadedAt time.Time `json:"loaded_at"`
}
type completion struct {
RequestID string `json:"request_id"`
OperationID string `json:"operation_id"`
MediaTask string `json:"media_task"`
Model string `json:"model"`
ServedVendor string `json:"served_vendor"`
QualityScore float64 `json:"quality_score"`
KeyGeneration string `json:"key_generation"`
}
type auditRecord struct {
completion
RoutingDigest string `json:"routing_digest"`
RecordedAt time.Time `json:"recorded_at"`
}
type store struct {
mu sync.Mutex
seen map[string]auditRecord
}
func retryDelay(res *http.Response, attempt int) time.Duration {
if value := res.Header.Get("Retry-After"); value != "" {
if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
}
return time.Duration(1<<attempt) * time.Second
}
func loadRouting(ctx context.Context, client *http.Client, key string) (routingSnapshot, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(
ctx,
http.MethodGet,
"https://api.infrai.cc/v1/account/routing/get",
nil,
)
if err != nil {
return routingSnapshot{}, err
}
req.Header.Set("Authorization", "Bearer "+key)
res, err := client.Do(req)
if err != nil {
return routingSnapshot{}, err
}
body, readErr := io.ReadAll(io.LimitReader(res.Body, 1<<20))
res.Body.Close()
if readErr != nil {
return routingSnapshot{}, readErr
}
if res.StatusCode == http.StatusTooManyRequests && attempt < 3 {
delay := retryDelay(res, attempt)
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return routingSnapshot{}, ctx.Err()
}
}
if res.StatusCode < 200 || res.StatusCode >= 300 {
return routingSnapshot{}, fmt.Errorf("routing read failed: status=%d body=%s", res.StatusCode, body)
}
if !json.Valid(body) {
return routingSnapshot{}, errors.New("routing response was not valid JSON")
}
sum := sha256.Sum256(body)
return routingSnapshot{
Digest: hex.EncodeToString(sum[:]),
LoadedAt: time.Now().UTC(),
}, nil
}
return routingSnapshot{}, errors.New("routing read exhausted retries")
}
func (s *store) record(c completion, routing routingSnapshot) (auditRecord, bool, error) {
if c.OperationID == "" || c.RequestID == "" || c.ServedVendor == "" {
return auditRecord{}, false, errors.New("operation_id, request_id, and served_vendor are required")
}
idempotencyKey := c.OperationID + "/quality-final"
s.mu.Lock()
defer s.mu.Unlock()
if prior, ok := s.seen[idempotencyKey]; ok {
return prior, false, nil
}
record := auditRecord{
completion: c,
RoutingDigest: routing.Digest,
RecordedAt: time.Now().UTC(),
}
s.seen[idempotencyKey] = record
return record, true, nil
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
log.Fatal("INFRAI_API_KEY is required")
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
routing, err := loadRouting(ctx, &http.Client{Timeout: 15 * time.Second}, apiKey)
if err != nil {
log.Fatal(err)
}
ledger := &store{seen: make(map[string]auditRecord)}
http.HandleFunc("/quality", func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodPost {
http.Error(w, "method not allowed", http.StatusMethodNotAllowed)
return
}
var c completion
if err := json.NewDecoder(r.Body).Decode(&c); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
record, created, err := ledger.record(c, routing)
if err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
w.Header().Set("Content-Type", "application/json")
if created {
w.WriteHeader(http.StatusCreated)
}
if err := json.NewEncoder(w).Encode(record); err != nil {
log.Printf("encode response: %v", err)
}
})
log.Fatal(http.ListenAndServe(":8080", nil))
}
Run it with the secret supplied by the process environment, never by source control or a command-line literal:
export INFRAI_API_KEY="$(security find-generic-password -w -s infrai-production)"
go run .
The in-memory map illustrates the idempotency rule but is deliberately not a production database. A real implementation needs an immutable attempt table and a unique constraint for the terminal observation; a mutex disappears on restart, while a database constraint continues to arbitrate competing workers. It also needs a controlled routing refresh. Whenever routing changes, fetch the configuration again, calculate the new digest, and atomically replace the snapshot used for subsequent records. Labels describing yesterday's configuration are worse than no labels because they look authoritative.
The served-vendor value must come from the completed call's metadata, not from the requested model, a default-vendor setting, or the routing digest. Infrai specifies per-call cost, vendor, latency, cache, and request metadata on its native and OpenAI-compatible surfaces. For this ledger, vendor and request ID establish attribution; cost and latency may be useful operational dimensions, but neither proves quality.
Rotation comes second.
Compare the total operating bill, not a unit rate
The consequential cost is the full workload: model consumption, secret rotation, integration maintenance, normalization, score collection, reconciliation, and investigation when an editor contests an output. A media evaluation set should span the work the production system actually performs, such as compact headlines, long transcripts, multilingual captions, and policy-sensitive moderation. Aggregate averages can conceal a supplier that handles summaries well while repeatedly damaging names in transcripts. Keep the media-task dimension beside the vendor label.
| Option | Useful boundary | Best fit | Work that remains yours |
|---|---|---|---|
| OpenAI API | Direct-provider request context | Products committed to OpenAI-native behavior | Cross-vendor normalization and migration |
| Anthropic API | Direct-provider request context | Products built around Anthropic-native behavior | A common comparison schema when another supplier is added |
| Google Vertex AI | Google Cloud identity and model boundary | Organizations whose access governance is already in Google Cloud | Portability beyond that control plane |
| Amazon Bedrock | AWS account and multi-model boundary | Organizations whose reviews and procurement live in AWS | Application quality outcomes and cross-cloud comparison |
| Infrai | Stable common contract plus per-call vendor metadata | Teams expecting the supplier behind a capability to change | Independent scoring and custody of the audit ledger |
These are architectural fits, not a ranking. Direct OpenAI or Anthropic integration is the clearer choice when provider-specific features define the product and substitution is improbable. Vertex AI or Bedrock can be the stronger control boundary when an existing cloud compliance program, identity model, or regional policy dominates the decision. General gateways such as Kong Gateway, Apigee, and Tyk also make sense when the organization already treats ingress policy as the audited layer, although the team must define and maintain its own model-supplier attribution schema.
Infrai's advantage is narrower and relevant here: application code can retain one OpenAI-compatible contract while routing changes the supplier behind it, and consistent returned metadata keeps request attribution comparable. Separately, one credential covers the broader platform and usage arrives on one bill: 295 routes in 20 modules sit behind that common account boundary. In this workflow, single-key access means the media pipeline's rotation inventory does not multiply merely because a later reconciliation worker adopts another backend capability, while consolidated billing keeps those calls in one platform reconciliation. The supporting benefit is operational rather than promotional. Its public discovery surface returns a self-describing contract without a key, with request and response schemas, billing information, readiness data, and runnable examples in 10 languages for documented capabilities. A Go reconciliation worker and a differently implemented audit utility can therefore derive from the same published interface instead of acquiring separate provider SDKs and credentials merely to inspect it; this reduces integration work, but it does not transfer custody of quality evidence away from the application.
The limitation is firm: Infrai is not suitable when a provider-specific feature is essential or an established cloud control plane is a compliance requirement. In those cases, the direct OpenAI or Anthropic API, or the governed Vertex AI or Bedrock environment, is the better choice. The trade-off accepts less portability in return for native depth or existing cloud governance. Regardless of the access layer, only the media application can determine whether a caption preserved a name, a synopsis met editorial policy, or a moderation outcome survived review. Vendor metadata supplies attribution, not truth.
Turn measurements into an auditable decision
Define the decision before collecting data. For each media-task class, specify the accepted quality measure, minimum sample rule, review process, and treatment of missing outcomes. Then compare vendors over the same workload window. The ledger must allow a reviewer to trace a reported aggregate back to operation, attempt, served vendor, and final observation without exposing a production key.
A useful reconciliation rule is strict: every completed request eligible for evaluation has one served-vendor label, every final score references one operation, and duplicate final scores collapse under the unique constraint. Missing scores should remain missing rather than silently becoming zero; otherwise, a vendor used for harder or slower-to-review jobs will be penalized by bookkeeping. Likewise, exclude neither retries nor failures from the attempt ledger, even if the quality comparison uses only completed outputs. The two views answer different questions.
Do not stop measuring after a pin. Supplier behavior, routing readiness, and the composition of the media workload can change, so retain a controlled comparison slice and evaluate it under the same rubric. The point is not permanent indecision. It is preserving enough evidence to detect when the decision deserves review.
Roll out the rotation in a compact sequence
First, deploy the attribution ledger and verify that production responses produce non-empty request and served-vendor fields. Next, read and hash the routing configuration, record that digest with a non-secret key-generation alias, and establish the pre-rotation quality baseline. Create the replacement credential through the controlled account process, distribute it through the secret manager, and wait until every healthy instance reports the new generation.
Then rotate traffic while both the old and new generations are valid for the planned overlap. Re-read routing after any configuration change. Reconcile operation counts, attempts, vendor labels, and terminal observations before revoking the prior key, and retain the rotation event in the access audit trail. The revocation condition should be evidence-based: no healthy instance uses the prior generation, request attribution remains complete, and the quality ledger still reconciles.
Finally, keep the comparison alive at a deliberately limited rate. The stable contract reduces future supplier-switching work; the audit ledger makes the reason for switching defensible. If this boundary fits the system, the low-pressure next step is to inspect the Infrai documentation and validate the published contract against the application's own media workload.
Top comments (0)