DEV Community

NoahHayes7250
NoahHayes7250

Posted on

Logistics Hiring Scores Using Unified Node Backend Proxy with One API Key and Retries

Short answer: a Node backend proxy should expose one API key and one unified endpoint while OpenAI, Anthropic Claude, and Google Gemini credentials remain behind the server boundary. Internal model aliases are mapped there, and only transient failures are retried. For logistics candidate scoring, the deciding outcome is valid, rubric-bound JSON. Provider breadth comes second.

Choice Client credential Structured-output control Operating burden Best fit
Direct provider SDKs One per provider Provider-specific Low at first Experiments tied to one provider
Managed multi-provider gateway One gateway key Common envelope, with provider differences still visible Low to medium A solo team that needs several providers now
Self-hosted routing proxy One internal key Fully owned normalization and validation High Strict network or policy control

Recommendation: for a one-person SaaS shipping weekly, start with a managed gateway only if it preserves structured responses and exposes retry metadata. Otherwise, put a thin internal adapter in front of direct SDKs. Do not build a general-purpose router. Outsource the undifferentiated transport work, but own the hiring rubric, schema, validation, and audit record.

A single key is a client-facing abstraction, not a universal provider credential. OpenAI, Anthropic, and Google each document their own authentication and request surfaces. Your browser or mobile client should never receive those upstream secrets. The backend holds them, while callers authenticate to one endpoint with one application-scoped secret.

How should a Node backend proxy use one API key?

The first criterion is schema fidelity. A candidate-scoring response that looks plausible but swaps a score for a string is not a partial success. It is rejected data. The same goes for a response that invents a rubric category, omits evidence, or returns a score outside the allowed range. The proxy should authenticate the caller, resolve the alias, send the provider-specific request, normalize the result, and validate the final object before returning it. That is the whole useful path; adding provider discovery, prompt catalogs, or a generic workflow engine would enlarge the maintenance surface without improving this decision.

Invalid means invalid.

Use JSON Schema as the contract. Draft 2020-12 defines validation vocabulary for object shape, required properties, numeric bounds, and whether extra properties are permitted. Provider-native structured-output features can help produce the shape, but the application must still validate the received value. Providers document different structured-output mechanisms and supported schema subsets, so a unified envelope must not pretend their behavior is identical.

For this job, I would measure four outcomes per internal alias: valid JSON, schema-valid JSON, rubric-valid content, and terminal failure. Keep them separate. A parser error and an unsupported extra field lead to different fixes. So does a valid object whose evidence is absent from the supplied resume.

The second criterion is failure semantics. Retry only failures that may succeed without changing the request: rate limits, transient server failures, and transport interruption. Do not retry an authentication failure. Do not blindly retry invalid output either; repeated sampling can multiply cost and latency without correcting a bad contract. One constrained repair attempt can be reasonable when the response is parseable and the validation errors are explicit, but record it as a repair, not a first-pass success.

This is where the revenue-per-hour lens helps. A broad abstraction is expensive to maintain. A narrow port with five observable outcomes is boring, and boring leaves time to ship the feature customers use. The explicit trade-off is less freedom for each call site in exchange for one enforceable contract.

The smallest useful contract

Keep provider model identifiers out of product code. The application asks for a capability alias such as rubric-strict; deployment configuration maps that alias to a concrete upstream model. Changing the map then becomes an evaluated release, not a search-and-replace operation.

import { z } from "zod";

const CandidateScore = z.object({
  candidateId: z.string().min(1),
  recommendation: z.enum(["advance", "review", "decline"]),
  criteria: z.array(
    z.object({
      rubricId: z.string().min(1),
      score: z.number().int().min(0).max(4),
      evidence: z.array(z.string().min(1)).max(3),
    }).strict(),
  ).min(1),
}).strict();

type CandidateScore = z.infer<typeof CandidateScore>;

type GenerateRequest = {
  model: "rubric-strict" | "rubric-fast";
  input: { candidateId: string; resumeText: string; rubric: string };
  idempotencyKey: string;
};

type GenerateResult = {
  value: CandidateScore;
  providerRequestId?: string;
  attempts: number;
};

interface ModelGateway {
  generate(request: GenerateRequest): Promise<GenerateResult>;
}

export async function scoreCandidate(
  gateway: ModelGateway,
  request: GenerateRequest,
): Promise<GenerateResult> {
  const result = await gateway.generate(request);
  return { ...result, value: CandidateScore.parse(result.value) };
}
Enter fullscreen mode Exit fullscreen mode

The schema is intentionally tight: scores are integers from 0 through 4, recommendations are enumerated, and unknown keys fail validation. Those numbers describe this example rubric, not a benchmark or a universal hiring standard. The evidence strings should quote or point to material already present in the candidate input; a later deterministic check can verify that grounding rule.

Keep environment configuration equally small. Parse it once during startup and fail before accepting traffic.

import { z } from "zod";

const Env = z.object({
  MODEL_GATEWAY_URL: z.string().url(),
  MODEL_GATEWAY_KEY: z.string().min(1),
  MODEL_RUBRIC_STRICT: z.string().min(1),
  MODEL_RUBRIC_FAST: z.string().min(1),
});

export const env = Env.parse(process.env);

export const modelMap = {
  "rubric-strict": env.MODEL_RUBRIC_STRICT,
  "rubric-fast": env.MODEL_RUBRIC_FAST,
} as const;
Enter fullscreen mode Exit fullscreen mode

Use a secret manager in production and inject values at runtime. Do not commit an .env file. Rotate the gateway key independently from upstream keys, and scope it to the service that needs it. One leaked application key should not reveal three upstream credentials.

Keep secrets dull.

How should retries preserve correctness?

Set one end-to-end deadline. Then fit attempts inside it. If an HTTP call has already consumed most of the request budget, another attempt is not resilience; it is a late response waiting to happen.

Exponential backoff with jitter reduces synchronized retries. Honor Retry-After when the upstream sends it, cap the number of attempts, and attach an idempotency key when the gateway supports one. The caller should see a stable error category such as rate_limited, upstream_unavailable, invalid_output, or unauthorized, while logs retain the upstream status and request identifier.

Streaming needs a separate decision. Server-Sent Events use the text/event-stream media type and reconnect by default in browsers. That is useful for generated prose. It is awkward for a score object that must be validated as a whole. For hiring scores, buffer the structured response, validate it, and only then return JSON. Correct first.

Do not log resumes, prompts, raw credentials, or complete model responses by default. Log the internal alias, resolved model version or identifier, schema version, latency, attempt count, validation outcome, token usage when supplied, and request IDs. Candidate material is sensitive operational data even when the article's main concern is model routing.

An evaluation set should include ordinary resumes plus adversarial cases: an empty employment section, conflicting dates, instructions embedded in resume text, missing rubric categories, and Unicode-heavy input. Run it before changing an alias. Compare first-pass schema validity and rubric assertions, not a single aggregate quality score. Ten hand-picked happy paths can catch wiring errors; they cannot establish production quality.

When is the runner-up the better choice?

Direct SDKs are the better runner-up when only one provider is in production, structured output is the dominant requirement, and its native feature exposes controls your common gateway cannot represent. OpenAI documents Structured Outputs with JSON Schema, Anthropic documents tool definitions built around input schemas, and Google documents schema-controlled JSON output. These are related mechanisms, not interchangeable contracts. Keeping a small provider adapter can preserve their useful differences.

Self-hosting wins when policy requires traffic to remain inside a controlled network, when custom admission rules must execute before every request, or when gateway observability is insufficient for an audit. Accept the trade: patching, capacity, secret rotation, telemetry, and on-call ownership now belong to you. For a solo SaaS, that work competes directly with the weekly release.

A managed gateway is attractive when multiple providers are already necessary and its normalized response retains the fields you need. Its limitation is the extra network hop and the possibility that normalization hides a provider-specific structured-output control. It is not appropriate when either issue breaks the scoring contract; use a direct adapter instead. Verify cancellation, timeout propagation, retry behavior, structured-output support, and access to provider request IDs before adopting it. A demo that returns text proves very little about a logistics scoring pipeline.

The durable boundary is small: one client credential, capability aliases, a versioned schema, explicit failure categories, and an evaluation gate. Everything behind it may change. The rubric and the acceptance test should not.

References

Top comments (0)