TL;DR: For a fintech hiring workflow, start with a queued micro-batch that records rubric version, input and output tokens, queue delay, model duration, validation status, and retry count for every candidate summary. Compare processing modes by the latency the review team can tolerate and by measured rubric quality, not by a provider's lowest advertised token rate. A cheap call that produces an unsupported score, stalls a review queue, or gets charged twice after a blind retry is expensive in the way that matters.
| Processing mode | Pick this when | Quality control | Latency shape | Main operational risk |
|---|---|---|---|---|
| Synchronous request | A reviewer is waiting and the text is already bounded | Validate before returning | Immediate, with a hard deadline | Timeout ambiguity and duplicate work |
| Queued micro-batch | Reviews can wait minutes and traffic arrives unevenly | Validate each result, then release the group | Queue delay plus processing time | Backlog hides a slow quality regression |
| Scheduled bulk run | Inputs close at a known time and results are needed later | Evaluate a sample before publishing | Long, predictable window | One bad configuration affects a large set |
The table is the starting point. Instrumentation makes it a decision rather than a guess.
How should a startup compare text summarization API batch cost?
Treat quality and latency as separate budgets. Quality is whether the summary preserves evidence from the candidate material, follows the job rubric, and refuses to invent missing evidence. Latency has at least two parts: time waiting in the queue and time spent generating and validating the result. One combined duration conceals the difference.
For this workflow, the output should help a human review evidence. It should not turn a generated summary into an automatic employment decision. Keep the source material, rubric version, and generated artifact linked under the retention and access rules already used for recruiting data. Short-lived telemetry can use opaque identifiers; logs do not need candidate names, email addresses, or raw application text.
The least complex useful scorecard has four quality checks: valid schema, every rubric claim tied to supplied text, required rubric dimensions present, and no score when evidence is absent. Measure those checks on a fixed evaluation set before changing a prompt, model, or batching policy. Then measure them again after the change. Fast is irrelevant if the artifact cannot pass review.
Price comes later.
There is one important boundary. Reranking and vector search can help select relevant passages before summarization, but they do not validate the final rubric summary. Cohere documents reranking as ordering documents by relevance, while pgvector provides vector similarity search for Postgres. Those are retrieval components. The runtime still owns prompt versioning, output checks, retry behavior, and release criteria.
Pick synchronous processing for an active reviewer
Use a synchronous call when a person has opened one candidate record and expects a response during that interaction. Give the operation a deadline. If it expires, report an unresolved state instead of quietly starting overlapping attempts.
The useful trace has a small shape: review request to model call to schema validation to result storage. In words, that is the whole diagram. Record the same correlation ID at each boundary so an operator can distinguish a slow upstream request from local validation or storage time.
This path buys immediacy, but it narrows the retry budget. A request that may have completed remotely before the connection failed is ambiguous. An idempotency key derived from the candidate record, rubric version, prompt version, and input revision prevents a retry from becoming a second logical job. Do not derive that key from personal data.
Pick queued micro-batches for uneven arrivals
This is the practical middle ground when applications arrive throughout the day but reviewers do not need each summary in the same second. Grouping work smooths bursts and gives the system room to respect concurrency limits. It also creates a new user-visible metric: queue age.
Queue age changes the diagnosis.
Watch the oldest ready job, not only average queue time. A healthy average can coexist with one partition that has stopped moving. Pair that gauge with completion counts by status and a histogram for end-to-end duration. Alert on a sustained breach of the review team's service objective, not on one slow call.
Micro-batches make quality releases easier. A new prompt version can process a bounded slice, remain quarantined, and be compared with the current version on the same rubric checks. Promotion should require both an acceptable error rate and an acceptable quality result. This is an explicit trade: a few more minutes of delay buys a smaller failure radius.
Cost belongs in the same event stream. Store input tokens, output tokens, logical job ID, attempt number, and processing mode. Normalize later to cost per 1,000 tokens using the applicable contract, because rates and batch terms can change. Never infer cost from character count when the bill is based on tokens.
Pick scheduled bulk runs for a closed review window
Use a scheduled run when the candidate set closes at a known time and summaries are due for a later review session. Freeze the rubric, prompt, and evaluation threshold for that run. A bulk job should have a manifest that says exactly which input revisions belong to it.
Bulk processing provides the most scheduling freedom and the largest blast radius. Release in stages: process a small validation slice, inspect automated quality results, and only then allow the remaining manifest to proceed. If the slice fails, stop. Short and sharp.
No partial release.
Completion percentage alone is a weak dashboard. Show accepted, rejected, retryable, and permanently failed counts separately. Add the age of the oldest unfinished item and token totals by rubric version. Those fields answer the questions an on-call engineer actually gets: Is it moving? Is it correct? Will it finish inside the review window? Did a revision change the workload?
Instrument one recoverable Node.js worker
The worker below uses generic interfaces on purpose. Its contract is more durable than any one transport or model API. Every fence is TypeScript, and the event names are intentionally boring: boring fields are easy to query at 02:00.
type CandidateJob = {
jobId: string;
candidateRef: string; // Opaque internal reference, not a name or email.
rubricVersion: string;
promptVersion: string;
inputRevision: string;
text: string;
};
type Summary = {
evidence: Array<{ rubricItem: string; excerpt: string }>;
missingEvidence: string[];
};
type Usage = { inputTokens: number; outputTokens: number };
type ModelResult = { value: unknown; usage: Usage };
interface Runtime {
summarize(input: {
text: string;
rubricVersion: string;
promptVersion: string;
idempotencyKey: string;
}): Promise<ModelResult>;
}
interface Telemetry {
event(name: string, fields: Record<string, string | number | boolean>): void;
}
function validateSummary(value: unknown): Summary {
if (typeof value !== "object" || value === null) {
throw new Error("summary_not_object");
}
const candidate = value as Partial<Summary>;
if (!Array.isArray(candidate.evidence) || !Array.isArray(candidate.missingEvidence)) {
throw new Error("summary_schema_invalid");
}
return candidate as Summary;
}
export async function processJob(
job: CandidateJob,
runtime: Runtime,
telemetry: Telemetry,
attempt: number,
queuedAtMs: number,
): Promise<Summary> {
const startedAtMs = Date.now();
const idempotencyKey = [
job.candidateRef,
job.rubricVersion,
job.promptVersion,
job.inputRevision,
].join(":");
telemetry.event("summary.started", {
jobId: job.jobId,
rubricVersion: job.rubricVersion,
promptVersion: job.promptVersion,
attempt,
queueDelayMs: startedAtMs - queuedAtMs,
});
try {
const result = await runtime.summarize({
text: job.text,
rubricVersion: job.rubricVersion,
promptVersion: job.promptVersion,
idempotencyKey,
});
const summary = validateSummary(result.value);
telemetry.event("summary.accepted", {
jobId: job.jobId,
rubricVersion: job.rubricVersion,
promptVersion: job.promptVersion,
attempt,
modelDurationMs: Date.now() - startedAtMs,
inputTokens: result.usage.inputTokens,
outputTokens: result.usage.outputTokens,
evidenceCount: summary.evidence.length,
missingEvidenceCount: summary.missingEvidence.length,
});
return summary;
} catch (error) {
telemetry.event("summary.rejected", {
jobId: job.jobId,
rubricVersion: job.rubricVersion,
promptVersion: job.promptVersion,
attempt,
modelDurationMs: Date.now() - startedAtMs,
reason: error instanceof Error ? error.message : "unknown_error",
});
throw error;
}
}
Keep retry classification outside this function. A queue consumer can retry timeouts and transient transport failures with a capped policy, while schema or evidence failures go to evaluation rather than looping. That distinction prevents a quality defect from masquerading as an availability problem. The before-and-after dashboard should be crisp. Before a runtime change, capture accepted-result rate, quality-check pass rate, p50 and p95 queue delay, p50 and p95 processing duration, retry rate, and tokens per accepted summary. After the change, compare the same fields by rubric and prompt version. Do not mix versions in one aggregate and call it an improvement. For a concrete comparison, hold the candidate set and rubric fixed, send one bounded slice through the current configuration and another through the proposed configuration, and label every event with its configuration version. Review quality failures before reading the latency chart. Next, separate queue delay from model duration so a worker-capacity problem is not blamed on generation. Finally, compare tokens per accepted summary rather than tokens per attempt; rejected output and duplicate attempts still consume the startup's budget even though no reviewer receives a usable artifact. This sequence is slower than choosing the smallest rate in a pricing table. It also produces an answer the team can defend.
The release rule should name both axes: ship only when quality remains above the team's validated threshold and end-to-end latency remains inside the review window. Token-normalized cost can break a tie between configurations that already satisfy those constraints. It cannot rescue a configuration that fails them.
Limits to keep visible
The limitations are concrete. Synchronous processing is not suitable for a review window that can absorb queue delay when burst control matters more than immediate feedback. Scheduled bulk processing is not suitable when a reviewer needs one changed record now. Micro-batching trades some immediacy for isolation and smoother capacity, and it still needs backlog monitoring. Automated schema and evidence checks do not prove that a summary is fair, complete, or suitable for an employment decision. Human review remains part of this workflow. Evaluation data also ages as job rubrics and candidate materials change, so version it and refresh it deliberately.
Averages hide tails. Token counts do not express summary quality. Retrieval relevance does not prove generation accuracy. Keep those boundaries visible, and the runtime comparison stays honest even as providers, rates, and workload volume change.
Further reading
- Cohere, “Rerank overview”: https://docs.cohere.com/docs/rerank-overview
- pgvector, “Open-source vector similarity search for Postgres”: https://github.com/pgvector/pgvector
Top comments (0)