To reduce the LLM API bill for a healthtech SaaS app, preserve the quality and latency of human review first. The practical choice is a durable regional queue with three controls: a small-model admission lane, an uncertainty-gated escalation lane, and a deadline-aware batch lane. Keep the product-facing Node.js service thin; put retries, idempotency, and scheduling policy in a worker boundary.
TL;DR: route ordinary reports through the lowest-latency model that has passed a task-specific quality gate, escalate ambiguous or safety-sensitive results, and batch only work whose review deadline can absorb the wait. This can reduce paid inference without turning price into the scheduler. Quality, queue age, and duplicate suppression remain the controlling signals.
I run cron and queue infrastructure in production, and I have been paged by both missed jobs and duplicate deliveries. The lesson that transfers to this workload is blunt: a successful API response is not proof that a report completed exactly once. A worker can finish inference and lose its acknowledgment; a lease can expire while a slow request is still running; a retry can then classify the same report again. Billing is one symptom. Conflicting labels and delayed human review are worse.
The invariant I use is therefore: one durable report identity, many permitted attempts, one accepted classification version. Everything else in the design follows from that.
Duplicates happen.
How can a SaaS app reduce its LLM API bill safely?
A router that inspects a prompt and immediately chooses “small” or “large” is missing the operational state around the prompt. The same moderation report deserves a different execution path when its human-review deadline is ten seconds away, the escalation queue is saturated, or the report already has an accepted result. Model choice is one scheduling decision among several.
Price cannot answer that scheduling question.
For this system, each queue message should carry a report ID, tenant ID, region, policy version, enqueue time, review deadline, attempt number, and payload reference. Do not place raw health information in logs or metrics labels. The payload can remain in the system of record while the message carries the minimum reference needed by an authorized worker. Data minimization and purpose limitation matter in the EU, while US deployments may also need administrative, physical, and technical safeguards for electronic protected health information. Those are design inputs, not a reason to mix regions in one convenient queue.
I would start with a policy like this and tune it against held-out, adjudicated reports:
| Lane | Admission rule | Execution | Failure action |
|---|---|---|---|
| Immediate | Human-review deadline is near | Small model now | Escalate uncertainty; never wait for a batch |
| Standard | Normal deadline and queue age | Small model now | Escalate low confidence or policy-sensitive labels |
| Batch | Deadline has measured slack | Regional micro-batch | Drain to standard lane before slack is exhausted |
| Human-only | Policy forbids automation or input is malformed | No model call | Send to reviewer with a reason code |
“Near,” “low confidence,” and “measured slack” are local policy values. They are not universal constants. A team could begin with an internal target such as a 60-second classification deadline and reserve 15 seconds for escalation, but those numbers are an example configuration, not an industry benchmark. Derive the real values from the human-review service objective and load tests.
There is a correction worth making early. I first tend to suspect the slow provider call when queue age rises; production queue work teaches a less comfortable possibility: admission may exceed sustainable concurrency long before an individual call looks abnormal. Track both. Queue age reveals accumulated work, while call latency describes only the attempts that obtained capacity.
Make retries boring before optimizing inference
At-least-once delivery is the useful default assumption for a queue worker. Exactly-once business effect has to be built at the write boundary. The worker below claims a report-policy pair, runs the selected lane under a deadline, and commits only if another attempt has not already won. Its interfaces are deliberately generic; the same contract works with a self-hosted queue or a managed one.
package moderation
import (
"context"
"errors"
"time"
)
var ErrAlreadyAccepted = errors.New("classification already accepted")
type Job struct {
ReportID string
PolicyVersion string
Region string
Deadline time.Time
PayloadRef string
}
type Result struct {
Label string
Confidence float64
ModelClass string
}
type Store interface {
Claim(ctx context.Context, key string, lease time.Duration) (bool, error)
CommitIfAbsent(ctx context.Context, key string, result Result) error
Release(ctx context.Context, key string) error
}
type Classifier interface {
Classify(ctx context.Context, payloadRef, modelClass string) (Result, error)
}
func Handle(ctx context.Context, job Job, store Store, models Classifier, now time.Time) error {
key := job.ReportID + ":" + job.PolicyVersion
claimed, err := store.Claim(ctx, key, 30*time.Second)
if err != nil || !claimed {
return err
}
committed := false
defer func() {
if !committed {
_ = store.Release(context.Background(), key)
}
}()
ctx, cancel := context.WithDeadline(ctx, job.Deadline)
defer cancel()
result, err := models.Classify(ctx, job.PayloadRef, "small")
if err != nil {
return err
}
if shouldEscalate(result) {
result, err = models.Classify(ctx, job.PayloadRef, "large")
if err != nil {
return err
}
}
err = store.CommitIfAbsent(ctx, key, result)
if errors.Is(err, ErrAlreadyAccepted) {
return nil
}
if err == nil {
committed = true
}
return err
}
func shouldEscalate(result Result) bool {
return result.Confidence < 0.82 || result.Label == "needs_safety_review"
}
The 0.82 threshold illustrates where a calibrated policy value belongs; copying it would be a mistake. Select it from precision and recall requirements for each label, then version it with the prompt and model class. A single global threshold can hide poor calibration on a rare, consequential category.
The lease duration is also an example. In a real implementation, either renew the lease while useful work continues or size it from an observed upper bound and ensure the commit is compare-and-set. A lease alone does not prevent duplicate effects. The atomic commit does.
Retry only failures that might succeed on another attempt: timeouts, transient transport errors, and explicit capacity responses. Malformed input belongs in the human-only lane. Policy rejections should be recorded with a stable reason code. Unbounded retries are delayed drops wearing a friendlier name, so cap attempts and send exhausted jobs to a reviewable dead-letter workflow.
Spend quality and latency budgets separately
The small-first path should earn its place with task-level evidence. Build a versioned evaluation set from de-identified, adjudicated moderation reports; split it so prompt tuning does not leak into the final check. Report per-label precision and recall, escalation rate, abstention rate, and disagreement with reviewers. An average accuracy number can conceal the exact class that wakes the on-call engineer.
Then run the candidate policy in shadow mode. It may compute a proposed route without changing the reviewer-visible decision. Compare that proposal with the currently accepted path, and examine disagreements by label and region before increasing traffic. This is also the cleanest way to learn whether confidence scores are calibrated enough to drive escalation.
Quality and latency need separate budgets because they fail differently. A quality breach should stop or narrow automated acceptance even if the queue is empty. A latency breach should shed optional work, reduce batch dwell time, or route reports directly to humans; it must not quietly lower the quality threshold. Put differently: saturation is not permission to become less careful.
Batching is useful when the backend supports it or when a worker can efficiently group independent calls, but the scheduler needs a latest-start calculation. For each job, subtract the measured upper-bound execution time, escalation reserve, and commit reserve from its review deadline. The earliest latest-start time controls when a partial batch must flush. Full batches are an optimization. Deadlines win.
Keep US and EU execution pools separate when residency or organizational policy requires it. Region should be assigned before enqueueing and included in authorization, queue selection, payload storage, and telemetry routing. A fallback that silently crosses the Atlantic is not a fallback; it is a policy change. If a regional model pool is unavailable, the explicit alternatives are to wait within the deadline, use an approved in-region fallback, or hand the report to a human.
Operate the queue as the product boundary
The dashboard should answer whether reports are being completed correctly, not merely whether inference endpoints are responding. I would page on sustained deadline risk and loss of processing capacity, then use lower-urgency alerts for changes in routing behavior. The useful signals include oldest-job age by region and lane, completions versus admissions, attempt count, duplicate-commit rejections, escalation rate, batch dwell time, deadline misses, human-only rate, and classification disagreement sampled from reviewer outcomes.
Do not label metrics with report IDs, tenant IDs, prompts, or free-form error text. Those dimensions are both sensitive and unbounded. Put trace correlation in access-controlled logs, retain it according to policy, and keep metrics dimensions small enough to remain operable during an incident.
The runbook should be short enough to use while tired:
- Confirm whether age is rising in one region, one lane, or all lanes.
- Compare admission rate with completed and exhausted work; check worker capacity and dependency latency.
- Stop batch admission and drain deadline-near jobs to the standard lane.
- Preserve idempotency keys when replaying. Never manufacture new report identities to make a retry move.
- If quality telemetry breaches its gate, disable automated acceptance for the affected policy version and retain human review.
- Record the first bad policy version and the first affected enqueue time for the postmortem.
Cost telemetry still belongs here, just not as the only objective. Attribute input and output units, attempt count, and model class to a policy version using bounded labels or an analytics stream. Look for waste from duplicates, escalation loops, overlong context, and batches that time out and replay. The most reliable reduction often comes from deleting repeated work before negotiating a cheaper unit of work.
Rollout deserves the same discipline. Deploy a policy version to a small, representative slice, keep the previous version available for new jobs, and never reinterpret an already-enqueued job under an unrecorded prompt. Roll back on quality or deadline gates. Review cost after those gates remain healthy.
When should you skip small-first routing?
Skip it when nearly every report escalates, when the first pass consumes enough deadline that the second pass misses review objectives, or when the small model cannot meet the required per-label quality gate. Two calls are then slower and may consume more inference than one appropriately selected call. Direct routing by validated report type can be cleaner.
It also does not apply to categories where policy requires a human decision or forbids automated processing. Send those reports straight to the review queue. No clever confidence threshold overrides governance.
The final design rule is narrow: use the cheapest eligible execution path only after eligibility has been established by quality, region, deadline, and idempotency state. That ordering keeps a cost-control project from becoming an incident generator. The Node.js application can enqueue and poll status; the durable worker system owns completion.
Top comments (0)