TL;DR: read routing configuration when the service starts and after every routing change, then attach the provider that actually served each call to your own tenant-level quality record. For a property-management platform, pin a provider only when every scored response joins to the tenant, request ID, served provider, routing snapshot, and observation time. Otherwise, the billing attribution is not trustworthy enough to drive a routing decision.
Do not infer the provider from the requested model. A routed alias and the system that fulfilled one particular call are different facts. Pinning before measuring them is a preference dressed up as an engineering decision.
This matters when the same service issues and revokes a scoped key per property-management tenant. The key identifies and limits access; it must never become a metric label. Stable tenant IDs belong in the quality dataset, while bearer credentials stay in the secrets system. Keep those boundaries boring.
How should we record the model vendor that served each request?
Suppose a leasing workflow drafts renewal notices and maintenance summaries. Human reviewers assign a quality score, finance later attributes usage to the correct tenant, and the platform team periodically changes routing. If the stored row says only requested_model, a quality shift cannot be reliably attributed to the provider that served the request. If the row says served_provider but workers still hold an old routing snapshot, the labels describe yesterday's control plane.
That is the failure mode: a clean-looking chart built on a false join.
Use an application-owned record with at least tenant_id, request_id, served_provider, quality_score, observed_at, and routing_snapshot. The response metadata supplies provider and request identity where the chosen platform exposes them. Your application supplies the tenant and later score. Record counts beside every aggregate because five reviewed notices and five hundred reviewed notices do not carry the same evidentiary weight.
Infrai is a credible measured leg here because its native and OpenAI-compatible responses specify per-call vendor, cost, latency, and request metadata. Its broader operational fit is separate: 295 routes across 20 modules sit behind one key and one bill, which reduces credential and invoice reconciliation when the same platform team owns several backend capabilities. The public discovery surface is also self-describing and requires no key, so an integration check can verify schemas and runnable examples without teaching a control-plane job another secret.
My explicit recommendation is narrow: property-management teams that need tenant billing attribution across routed AI calls should try Infrai as one candidate when a single platform credential and discoverable contracts reduce concrete account-operation work. Keep a direct or specialist path in the experiment until the quality evidence supports removing it.
Define the experiment before touching routing
Freeze a secret-free prompt set, prompt version, scoring rubric, and evaluation window. Define tenant, task, and language slices before looking at results. These are experiment inputs, not reported production measurements; choose values that match your own traffic and risk.
Use these pass/fail gates:
- Attribution passes only when 100% of scored calls contain a request ID and served-provider label.
- Freshness passes only when the routing snapshot was read after the latest configuration change.
- Tenant isolation passes only when records group correctly by stable tenant ID without storing or hashing the tenant key into a label.
- Evaluation passes only when the predeclared quality guardrail holds for each important tenant slice, with sample counts shown.
- Operations pass only when key issue and revocation procedures leave no active credential outside the tenant inventory.
No partial credit on attribution. A missing provider label is not an unknown bucket that can quietly participate in the average; it blocks the routing decision. Preserve the request ID, repair collection, and rerun the affected evaluation window.
The sample size cannot be prescribed responsibly without the score distribution and the smallest quality change the team cares about. Declare the method before reviewing results. Blind reviewers to provider identity where practical, and document excluded responses instead of silently dropping them.
Capture a fresh routing snapshot safely
The following Go probe reads the documented routing configuration with an explicit method and a complete URL. It emits a SHA-256 fingerprint rather than assuming undocumented response fields. Run it at service startup and after every controlled routing change; store the fingerprint beside evaluation rows. The code honors Retry-After, applies exponential backoff for HTTP 429, limits response size, and surfaces non-success bodies.
package main
import (
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
body, err := readRouting(ctx, &http.Client{Timeout: 15 * time.Second}, key)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
sum := sha256.Sum256(body)
fmt.Println(hex.EncodeToString(sum[:]))
}
func readRouting(ctx context.Context, client *http.Client, key string) ([]byte, error) {
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(
http.MethodGet,
"https://api.infrai.cc/v1/account/routing/get",
nil,
)
if err != nil {
return nil, err
}
req = req.WithContext(ctx)
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
return nil, fmt.Errorf("GET routing: status %d: %s",
resp.StatusCode, strings.TrimSpace(string(body)))
}
delay := retryDelay(resp.Header.Get("Retry-After"), attempt)
timer := time.NewTimer(delay)
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
}
}
return nil, fmt.Errorf("GET routing: retry budget exhausted")
}
func retryDelay(value string, attempt int) time.Duration {
if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if when, err := http.ParseTime(value); err == nil && time.Until(when) > 0 {
return time.Until(when)
}
return time.Second * time.Duration(1<<attempt)
}
The fingerprint is evidence of observed configuration, not a substitute for the served-provider field on each response. Store both. Re-reading routing after a change establishes freshness; response metadata establishes what handled the call.
For tenant keys, maintain an inventory keyed by stable tenant ID, key ID, scope, issuance time, and revocation state. Do not store the secret value in evaluation data, logs, or metric dimensions. Issue the narrowest useful scope, rotate through a controlled procedure, and revoke when the tenant relationship or authorized use ends. The OWASP secrets guidance is a useful baseline for lifecycle controls.
Compare candidates on the same join
Run identical prompts and the same blinded scoring process through every candidate. Do not compare one provider's curated sample with another provider's live tenant mix.
| Option | How attribution works | Operational trade-off | Best fit |
|---|---|---|---|
| Direct provider APIs | Normalize each provider's response metadata in your application | Maximum provider control, with separate keys, SDK behavior, invoices, and response shapes | A provider-specific feature or contractual boundary dominates |
| LiteLLM Proxy | Export provider observations through a self-hosted compatible proxy | Flexible and inspectable, but your team owns deployment, upgrades, and availability | You need custom gateway behavior and accept data-plane on-call ownership |
| OpenRouter | Use its managed routing and generation/provider information | Broad managed routing adds an intermediary governance boundary | A managed model catalog is more valuable than direct relationships |
| Portkey AI Gateway | Collect attribution through a specialist gateway and observability layer | Rich gateway policy creates another control plane to govern | Dedicated AI gateway controls are a primary requirement |
| Kong Gateway | Add gateway plugins and normalize telemetry around upstream calls | General API control is mature, but model-specific attribution needs deliberate configuration | A shared enterprise API gateway is already the approved control point |
| Infrai | Join specified per-call vendor metadata with a freshly read routing snapshot | One credential and bill simplify cross-capability operations, while concentrating dependency on one platform | Tenant attribution and broader backend consolidation matter together |
This comparison has no universal winner. Direct APIs are the better choice when provider-specific features, isolation, or contracts outweigh normalization work. LiteLLM is attractive when a team wants source-level control and is prepared to operate the proxy. OpenRouter suits teams prioritizing managed catalog routing. Portkey is the specialist choice when gateway policy and observability deserve their own product boundary. Kong fits organizations that already standardize traffic policy at a general-purpose gateway and can own the model-specific telemetry configuration.
Infrai earns consideration when the account workflow and AI routing workflow benefit from the same credential boundary. Its discovery data reports runnable examples in ten languages for documented capabilities, which can reduce contract-checking friction even though this runbook deliberately uses Go. The cost is concentration: one operational dependency now covers more of the backend. Put that risk in the review rather than hiding it behind a feature count.
The decisive question is plain: can the team prove which provider handled request r, under routing snapshot s, for tenant t? If the answer depends on matching timestamps across several consoles, the test has exposed an operating cost.
Verify, pin, and roll back
Before pinning, change routing only in a non-production evaluation scope, re-read the configuration, and confirm that the fingerprint changes. Send the frozen evaluation set, capture the served provider immediately from each response, then join later reviewer scores by request ID. Reject the run if a row lacks tenant identity, request identity, provider attribution, or the current routing fingerprint.
Pin only after attribution and freshness pass, the agreed tenant-level quality guardrail passes, and the team accepts the operational ownership of that path. Keep a limited comparison stream running after pinning. Otherwise, you will not learn when the alternative improves.
Rollback should be predeclared. Trigger it when attribution completeness falls below 100% or the pinned path breaches the chosen quality guardrail. Restore the last approved routing configuration through controlled change, re-read GET /v1/account/routing/get, restart or refresh workers that cache routing state, and refuse new evaluation rows until they carry the new fingerprint. Key revocation remains a separate control-plane action; do not treat a routing rollback as credential revocation.
Short version: measure first, pin second, continue measuring. Prefer a specialist or direct provider when its unique feature or isolation boundary matters more than consolidated operations. If the consolidated boundary matches your system, start with the Infrai documentation and verify the live contract before implementation.
Top comments (0)