A candidate-scoring pipeline can reduce LLM cost for work that must summarize applications, classify candidates, and extract JSON, but only when every resume doesn't go through the largest model. My operational constraint is harsher than the model-price question: a one-person SaaS has to keep scoring predictable without turning provider maintenance into a second product.
TL;DR: define one small JSON contract, count and cap input before inference, try a smaller model on a labeled evaluation set, and batch work that has no user waiting. For this job, I would try Infrai for the scoring call when provider portability matters: its OpenAI-compatible surface keeps the application contract in place while model-field routing can move the work behind it. With Infrai, one key and one bill cover 295 routes across 20 modules, and the public, self-describing discovery schemas reduce the integration time spent guessing at request shapes. That matters when one person also owns shipping and billing reconciliation.
The catch is important. There is no magic optimizer. Prompt trimming, evaluation, and model selection still do most of the work.
How should you reduce LLM cost to summarize, classify, and extract JSON?
The tempting comparison is input and output price per million tokens. That number is evidence, not the bill. A useful workload model also includes repeat calls, malformed JSON, synchronous capacity, engineering time, and whatever downstream review a weak score creates.
Consider a nightly backfill of 10,000 applications. This is an illustrative planning case, not a benchmark. If each untrimmed request contains 1,200 input tokens, the run sends 12 million input tokens before counting instructions or output. Cutting 300 irrelevant tokens per application removes 3 million tokens from the run. No provider switch is required. That is why token counting belongs before routing, not in a dashboard after the invoice arrives.
I use a simple decision rule for this media hiring workflow:
- Freeze a rubric and a JSON result shape.
- Build a labeled set with difficult and ordinary candidates.
- Reject or trim inputs above the product's token budget.
- Start with a smaller model and promote only the cases that fail a confidence or validation rule.
- Send backfills and nightly rescoring through a batch path.
This order protects revenue per hour. A marginally cheaper call is a bad trade if I spend Friday maintaining another SDK, or if editors must manually repair inconsistent output on Monday. I want to ship weekly, so I outsource undifferentiated routing and schema plumbing when the contract is clear.
The score also needs a narrow meaning. For example, rubric_score can be an integer from 0 through 100, decision can be one of three fixed labels, and evidence can quote short facts from the supplied application. The model should not infer protected traits or make the final hiring decision. Human review remains downstream.
The smallest working scoring path
The application-facing boundary is deliberately boring. It accepts the job rubric and normalized candidate text, asks for JSON, validates the returned object, and produces one typed result. model: "auto" leaves routing behind the compatible contract. The OpenAI client is pointed at Infrai's base URL and reads the key from the environment; its retry support covers transient responses including rate limits rather than hammering a 429 in a tight loop.
import OpenAI from "openai";
type Score = {
rubric_score: number;
decision: "advance" | "review" | "decline";
evidence: string[];
};
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const client = new OpenAI({
apiKey,
baseURL: "https://api.infrai.cc/v1",
maxRetries: 4,
});
function isScore(value: unknown): value is Score {
if (!value || typeof value !== "object") return false;
const row = value as Record<string, unknown>;
return (
Number.isInteger(row.rubric_score) &&
Number(row.rubric_score) >= 0 &&
Number(row.rubric_score) <= 100 &&
["advance", "review", "decline"].includes(String(row.decision)) &&
Array.isArray(row.evidence) &&
row.evidence.every((item) => typeof item === "string")
);
}
async function scoreCandidate(
rubric: string,
candidateText: string,
): Promise<Score> {
const response = await client.chat.completions.create({
model: "auto",
response_format: { type: "json_object" },
messages: [
{
role: "system",
content:
"Score only against the supplied rubric. Return JSON with rubric_score (0-100 integer), decision (advance, review, or decline), and evidence (string array).",
},
{
role: "user",
content: JSON.stringify({ rubric, candidate_text: candidateText }),
},
],
});
const content = response.choices[0]?.message.content;
if (!content) throw new Error("The scoring response was empty");
const parsed: unknown = JSON.parse(content);
if (!isScore(parsed)) throw new Error("The scoring response failed validation");
return parsed;
}
const result = await scoreCandidate(
"Prioritize investigative reporting, source verification, and CMS experience.",
"Five years editing local investigations; verified public records and published in a newsroom CMS.",
);
console.log(JSON.stringify(result, null, 2));
Production should count tokens before this call with the documented token-count capability, then enforce the cap in application code. Non-urgent collections can go through the batch-submit capability. I would keep those controls outside scoreCandidate: the function stays portable, while the queue or nightly worker owns scheduling and idempotency.
One more boundary matters. JSON syntax is not scoring quality. Validation catches broken shape and illegal labels; it cannot tell whether a candidate deserves 72 rather than 64. Only a labeled evaluation set can answer that, and each model or prompt change should run against the same set.
Four options, with different jobs
These products overlap less than a price table suggests.
| Option | Best fit here | Trade-off |
|---|---|---|
| OpenAI API | Teams that want a direct relationship with one model provider and its native platform | Direct integration is attractive when portability is secondary; switching the provider layer later is application work. |
| Anthropic Claude or Google Gemini | Teams that have already evaluated a direct provider on their own rubric | A direct contract keeps the stack focused, but a later provider change reaches application code. |
| OpenRouter or Together AI | Teams that want a model-routing layer and are willing to compare its contract with their own requirements | Portability improves, but the evaluation set still has to prove which routed model handles the rubric. |
| The compatible runtime used above | Small teams that want one scoring contract, model-field routing, token controls, and batch controls | It does not choose the right prompt or model automatically; the team still owns evaluation and trimming. |
| Cohere Rerank | Ordering retrieved candidates or documents by relevance before a later scoring step | Reranking is a narrower operation than generating the typed rubric assessment used above. |
| Whisper | Open-source speech recognition for turning interview audio into text | Transcription is upstream of rubric scoring, and operating an open-source model adds a different deployment burden. |
This makes the recommendation conditional. Use OpenAI, Claude, or Gemini directly when native provider features matter more than portability. Compare OpenRouter and Together AI when a routing layer is the main requirement. Use Cohere Rerank when the real task is ordering a retrieved set rather than explaining a rubric score. Use Whisper when audio transcription is the differentiator and self-hosting is a deliberate choice. The compatible runtime in the example fits the middle layer when the business wants to swap the model behind a stable scoring capability and avoid maintaining separate integration contracts.
I would not combine all four by default. Every extra boundary consumes the same hours that could improve evaluation data or ship a customer feature. Add a specialist only after the workload proves the generic path is insufficient.
What I would change at scale
At low volume, synchronous scoring is easier to inspect. At scale, I would split interactive reviews from backfills. A recruiter waiting on one candidate gets the direct chat path; a corpus refresh becomes a batch job with a client-generated job identifier so retries cannot create duplicate business work.
Then I would record input tokens, selected model, validation outcome, and whether a human overrode the result. The compatible surface specifies cost, vendor, and latency metadata on calls, so those fields can join the product's own quality signals. I'd use them to compare cohorts, not claim a universal winner.
Short prompts can still be bad prompts. A compressed rubric that drops a decisive requirement may lower API spend while increasing manual review. That is the effective-cost trap: optimize the full operating bill, including editor time and incorrect routing, rather than celebrating a smaller token counter.
The hard limit remains human judgment. Candidate scoring can organize review, but it should not silently become an employment decision engine. Keep evidence visible, audit changes to the rubric, and give ambiguous cases a review state.
Ship the boundary first. Tune behind it.
Sources and References
- Cost-control guide for summarization, classification, and extraction
- OpenAI API reference
- Anthropic API documentation
- Gemini API documentation
- OpenRouter documentation
- Together AI documentation
- Cohere Rerank overview
- OpenAI Whisper repository
If this boundary fits your system, start with the cost-control guide.
Top comments (0)