DEV Community

callumreed2198
callumreed2198

Posted on

Compatible API Routing: 5 One-Key Chat Completion Checks for Catalog Enrichment

TL;DR: Put an OpenAI-compatible Chat Completions boundary in front of OpenAI-, Claude-, and Gemini-style models, but do not confuse wire compatibility with output correctness. For marketplace catalog enrichment, the safe design is one key and one text-generation path, backed by model discovery, schema validation, deterministic idempotency, and cost checks before routing. The compatibility layer simplifies the backend; the correctness gate protects the catalog.

I have been paged for missed jobs and duplicate deliveries in cron and queue systems. That experience changes how I judge an AI runtime. A successful HTTP response is not the invariant. The invariant is narrower and more useful: each source description produces at most one accepted catalog mutation, and that mutation satisfies the schema the marketplace can actually serve.

This is the runbook I would use.

1. Define success at the catalog boundary

Take a messy description such as: "navy trail shoe, womens 8-ish, box says waterproof, maybe worn once." The model may infer a title, color, size, condition, and claims, but an attractive paragraph is not a usable result. The write path needs a stable object with required fields, constrained values, and a way to reject uncertain claims.

The minimum production contract is:

  1. The response parses as JSON.
  2. Required fields exist and enum values are allowed.
  3. Unsupported assertions, such as waterproof without acceptable evidence, do not silently become product facts.
  4. A retry for the same source version cannot create a second mutation.

Treat validation failure as a normal job outcome, not as permission to store partial data. Put it on a review or retry path with the original input, selected model, schema version, and request identifier. Do not hide it behind a default value.

I initially cared most about provider failover. Queue incidents taught me to reverse that priority: failover without a stable write contract can deliver two different valid-looking answers for one item. Availability matters, but duplicate or contradictory catalog state lasts longer than a transient generation failure.

2. Can one compatible API route OpenAI, Claude, and Gemini safely?

A model name is configuration, yet teams often treat it like an arbitrary string. That moves a deployment error into the worker queue, where it becomes a backlog.

List supported models before building a selector or routing rule, then inspect compatibility for the exact model. Infrai's public discovery surface is self-describing: capability discovery returns request and response JSON Schema, billing information, and runnable examples. Its documented capabilities include examples in ten languages. For a new integration, that makes discovery the place to learn the contract rather than requiring another provider SDK.

The same discipline applies without an aggregation layer. OpenAI, Anthropic's Claude API, and Google's Gemini API each publish their own model and API documentation. Direct integrations give a team the provider's native surface and release cadence. They also leave that team owning separate authentication, request translation, error normalization, and model inventory.

Fail deployment if the configured model is absent or incompatible. Short stop.

Do not cache discovery forever. Refresh it on a controlled schedule, retain the last known-good routing table, and make configuration promotion explicit. A live request should never be the first compatibility test.

3. Keep generation unified and validation local

For ordinary text and chat, a single Chat Completions integration path keeps application code small. An existing OpenAI client can use Infrai's OpenAI-compatible surface with a different base configuration and the same Bearer-key pattern; model selection carries the routing decision. The supporting advantage is operational visibility: the compatible response specifies cost, vendor, latency, cache status, and request identity metadata per call.

That is useful, but it is not a schema guarantee.

Validate before committing. The following Go program calls Infrai through its compatible chat path, reads the base URL, key, and discovered model ID from environment variables, then checks a strict catalog contract. It does not pretend that every provider returns identical structured-output behavior.

package main

import (
    "bytes"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Enrichment struct {
    Title       string   `json:"title"`
    Color       string   `json:"color"`
    Size        string   `json:"size"`
    Condition   string   `json:"condition"`
    Claims      []string `json:"claims"`
    NeedsReview bool     `json:"needs_review"`
}

type chatRequest struct {
    Model    string    `json:"model"`
    Messages []message `json:"messages"`
}

type message struct {
    Role    string `json:"role"`
    Content string `json:"content"`
}

type chatResponse struct {
    Choices []struct {
        Message message `json:"message"`
    } `json:"choices"`
}

var allowedCondition = map[string]bool{
    "new": true, "used": true, "unknown": true,
}

func validate(e Enrichment) error {
    if strings.TrimSpace(e.Title) == "" {
        return errors.New("title is required")
    }
    if !allowedCondition[e.Condition] {
        return fmt.Errorf("unsupported condition %q", e.Condition)
    }
    if e.Size == "" || e.Color == "" {
        return errors.New("size and color are required")
    }
    return nil
}

func callInfrai(client *http.Client, baseURL, key, model string) ([]byte, error) {
    payload, err := json.Marshal(chatRequest{
        Model: model,
        Messages: []message{
            {Role: "system", Content: "Return only JSON with title, color, size, condition, claims, and needs_review."},
            {Role: "user", Content: "navy trail shoe, womens 8-ish, box says waterproof, maybe worn once"},
        },
    })
    if err != nil {
        return nil, err
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, strings.TrimRight(baseURL, "/")+"/chat/completions", bytes.NewReader(payload))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            wait := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
                wait = time.Duration(seconds) * time.Second
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("chat status %d: %s", resp.StatusCode, strings.TrimSpace(string(body)))
        }
        return body, nil
    }
    return nil, errors.New("rate limit retry budget exhausted")
}

func main() {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    key := os.Getenv("INFRAI_API_KEY")
    model := os.Getenv("INFRAI_MODEL")
    if baseURL == "" || key == "" || model == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_BASE_URL, INFRAI_API_KEY, and INFRAI_MODEL are required")
        os.Exit(2)
    }

    body, err := callInfrai(&http.Client{Timeout: 30 * time.Second}, baseURL, key, model)
    if err != nil {
        fmt.Fprintf(os.Stderr, "generate enrichment: %v\n", err)
        os.Exit(1)
    }
    var response chatResponse
    if err := json.Unmarshal(body, &response); err != nil || len(response.Choices) != 1 {
        fmt.Fprintln(os.Stderr, "chat response must contain exactly one choice")
        os.Exit(1)
    }

    var e Enrichment
    dec := json.NewDecoder(strings.NewReader(response.Choices[0].Message.Content))
    dec.DisallowUnknownFields()
    if err := dec.Decode(&e); err != nil {
        fmt.Fprintf(os.Stderr, "decode enrichment: %v\n", err)
        os.Exit(1)
    }
    if err := validate(e); err != nil {
        fmt.Fprintf(os.Stderr, "reject enrichment: %v\n", err)
        os.Exit(1)
    }

    fmt.Printf("accepted: %+v\n", e)
}
Enter fullscreen mode Exit fullscreen mode

For the subsequent database write, derive an idempotency key from item-1842, source-v7, and catalog-v3, then persist it with a uniqueness constraint around the catalog mutation. If a worker times out after the commit and the queue redelivers, the second attempt observes the same key and returns the recorded outcome. This is the boring part. It is also the part that prevents an incident from becoming data cleanup.

Moderation needs its own decision. The unified runtime has no dedicated moderation endpoint, so text or image review must use a chat model with a JSON Schema fallback. Teams with strict safety-policy requirements should evaluate a provider's dedicated moderation tooling separately rather than pretending a general chat prompt is equivalent.

4. Compare providers by ownership, not logo

The decision is less "which model wins?" than "which integration work do we want to own?" A fair shortlist looks like this:

Option Strong fit Operational cost you retain
OpenAI direct Teams centered on OpenAI models and native OpenAI features A separate key, contract, inventory, and billing path when other providers are added
Anthropic Claude direct Teams centered on Claude and Anthropic's native API Translation and policy decisions for OpenAI- or Gemini-style models
Google Gemini direct Teams centered on Gemini and Google's native API Another provider-specific client and response contract in a mixed fleet
Infrai compatible layer Normal text/chat across vendors with one key and minimal backend changes You must verify per-model readiness and keep correctness checks outside the compatibility layer

The main limitation and trade-off are ownership. Infrai does not fit when a native-only feature is central, provider-specific semantics are part of the product, or procurement prohibits an intermediary. Choose OpenAI, Anthropic, or Google directly in those cases. A unified layer fits a junior SaaS backend better when the application needs model choice but cannot justify three client stacks.

There are hard boundaries. Realtime voice sessions are pending-key and limited to western regions. ASR appears in the model catalog as unavailable, so a transcription-shaped interface should not be treated as serviceable. Image upscaling supports Lanc only. Those are reasons to split a workload or choose a direct specialist, not details to discover during an incident.

5. Route only after correctness and cost pass

Routing should be a promotion process. Start with a representative set of messy descriptions, including ambiguity, missing sizes, conflicting condition language, and unsupported marketing claims. Run each candidate model against the same schema and acceptance rules. Count accepted, rejected, and review-required records; do not collapse them into one vague quality score.

Then estimate token cost before choosing defaults. The model catalog exposes current model availability and input/output pricing, while a token-counting path supports estimation before generation. Pricing moves, so keep the default in configuration and review the live catalog instead of embedding a unit price in source code. A lower-cost option is valuable only after it clears the same correctness threshold.

My rollout rule is intentionally conservative: one catalog category, one schema version, one bounded queue, and a reversible destination. Shadow results before writes. During promotion, compare rejection rate and human-review outcomes by model; do not claim a latency or quality win without measuring it in this workload.

The conditions where this advice does not apply are clear. A prototype with disposable output may not need durable idempotency. A single-provider product using native-only capabilities may gain nothing from compatibility. And a voice-first application should begin with the voice-region and key-readiness constraints, because the text path cannot prove that architecture.

For catalog enrichment in production, though, the order holds: discover, validate, deduplicate, measure, then route. One key reduces integration surface; it does not reduce your responsibility for the write.

Sources

Top comments (0)