Choose adapter-owned decision events for a portable Node.js moderation layer, and keep provider details out of the telemetry contract. The deciding constraint is not how quickly one classifier can return JSON. It is whether a B2B SaaS team can change the upstream model without breaking safety policy, code-review findings, or the observability budget.
TL;DR: one ingress API key is an access boundary, not a moderation architecture. Put a small classifier contract behind it, validate every response, and emit one bounded event per decision. Sample diagnostic traces separately. This gives code-review services enough evidence to audit decisions while preventing model names, rule text, repository identifiers, and free-form findings from turning into permanent high-cardinality labels.
The choice has a boundary. Adapter events are the better default for synchronous moderation on user-submitted code changes. Inline provider metrics remain reasonable for a narrow, single-provider service whose labels are fixed and whose migration risk is explicitly accepted.
Can one API key keep moderation portable across providers?
This architecture decision starts with invariants. A request enters through one service credential, but the credential does not imply that OpenAI, Claude, and Gemini expose identical request fields or identical structured-output behavior. The Node.js boundary must therefore own the stable parts: the input envelope, the allowed verdicts, schema validation, timeouts, and the fail-open or fail-closed policy. Provider-specific translation belongs behind that boundary.
For a code-review product, the useful output is deliberately small. A decision can be allow, block, or review; a reason code comes from a controlled vocabulary; and findings are structured records rather than label values. The service may retain a short text explanation for authorized review, but it should never copy that text into a metric dimension.
Three failure boundaries matter. A transport failure means no classifier result arrived. A contract failure means bytes arrived but did not satisfy the expected schema. A policy result means the validated classifier intentionally returned block or review. Combining all three into moderation_failed=true saves a line of code and destroys the operational distinction needed during an incident.
The request path also needs an explicit deadline. Once that deadline expires, application policy decides what happens next; the adapter should not silently invent a verdict. Retries must be bounded because a retry changes both latency and usage. For the same reason, log the attempt count as a small integer field, not a unique request label.
The invariants are compact:
- The public decision schema does not contain a provider-specific finish reason or model identifier.
- Every accepted result passes local schema validation before application code reads it.
-
provider,outcome,reason_code, andschema_versionuse enumerated values. - Request bodies, code patches, explanations, and repository names stay out of metric labels.
- Missing, late, malformed, and policy-denied results remain separate outcomes.
These rules are stricter than a shared chat-completions-shaped endpoint. They have to be. Surface similarity can reduce integration work, but it cannot make upstream semantics interchangeable.
Portability is conditional.
The two telemetry designs
Inline metrics let each provider branch increment counters and attach whatever context is available there. Adapter events instead emit a normalized record after validation, then derive metrics, logs, and sampled traces downstream. Both designs can count requests. They differ sharply once provider portability becomes a real requirement.
| Decision factor | Inline provider metrics | Adapter-owned decision events |
|---|---|---|
| Schema ownership | Repeated inside each branch | One versioned application contract |
| Provider switch | Dashboards and alerts may need label changes | Stable fields survive adapter replacement |
| Cardinality control | Depends on every call site | Enforced once at the event boundary |
| Debug detail | Easy to add, easy to retain accidentally | Stored separately under explicit sampling and retention |
| Failure taxonomy | Often coupled to an SDK exception shape | Normalized before aggregation |
| Best fit | One provider, small fixed label set | Multiple providers or planned portability |
I choose adapter events because the migration unit should be an adapter, not every dashboard and alert. This is an operational choice, not a claim that one upstream classifier is safer than another. OpenAI, Claude, and Gemini are provider families behind the boundary; none should define the application's durable event schema.
There is a cost argument, but it is not a price comparison. Suppose the bounded dimensions have 3 providers, 4 outcomes, 12 reason codes, and 2 schema versions. The full Cartesian ceiling is 288 series before deployment and region dimensions. Add 20 deployments and 4 regions and the theoretical ceiling becomes 23,040. Most combinations may never appear, yet the multiplication exposes the budget before production traffic does.
Now consider a repository_id label with 50,000 possible values. That single choice multiplies the theoretical space to 1,152,000,000 series. The arithmetic is a warning, not a forecast: actual series creation depends on traffic and the monitoring system. Still, no sampling percentage repairs an unbounded label design. Keep tenant and repository identifiers in access-controlled event storage only when an investigation requires them, and give that store its own retention policy.
Short labels win.
Measure first.
The critical path in one request
The external route below is an application-owned example, not a claim about any provider's native API. Its contract is intentionally narrow. The gateway chooses an adapter from server-side configuration, so callers present one service key and do not encode a provider into business logic.
response=$(curl --fail-with-body --silent --show-error \
--max-time 8 \
--request POST \
--url https://moderation.example.invalid \
--header "Authorization: Bearer ${MODERATION_SERVICE_KEY}" \
--header "Content-Type: application/json" \
--header "Idempotency-Key: review-8f3d2" \
--data '{
"schema_version": "1",
"task": "code_change_review",
"content": {
"language": "javascript",
"patch": "const query = userInput;"
},
"response_schema": {
"verdict": ["allow", "block", "review"],
"reason_code": "string",
"findings": "array"
}
}')
printf '%s\n' "${response}"
The service validates the JSON response against its local schema before returning it. It then emits one decision event with fields such as provider_family, outcome, reason_code, schema_version, latency_bucket, and attempt_count. The code patch and returned findings travel through the request path but are not copied into telemetry by default. A separate audit record can hold the minimum content required by the product's review workflow, protected by access controls and a stated deletion schedule.
This split matters for structured findings. A finding may include a path, line number, category, and explanation. Path and explanation are useful to a reviewer, but both are hostile metric labels: paths expand with every repository, while explanations are effectively unbounded strings. Count finding_category only if its vocabulary is controlled. Count the number of findings as a histogram or bounded bucket rather than turning each count into a new label.
The synchronous path should record no more detail than operators can justify retaining. If a malformed response appears, preserve a redacted sample under a low, explicit diagnostic sampling rate and a shorter retention window. Do not sample the aggregate decision counter; sample the bulky evidence. A 1% trace sample and a 100% outcome counter answer different questions, so placing both behind one sampling switch is a category error.
How much evidence is enough?
Retention begins with a question, not a default. Which investigation must this record support, and for how long? Aggregate counters may need a longer horizon for regressions. Raw user content usually deserves the shortest defensible window because every additional day increases stored bytes and exposure. Regulatory obligations, contractual commitments, deletion requests, and internal incident-response needs can change that answer; 45 CFR Part 164 is one example of why a health-data system cannot borrow a generic SaaS retention setting without legal and security review.
A practical budget can be expressed without vendor prices. Estimate daily event bytes as requests × average encoded event size, apply the retention days, then add index and replication overhead measured in the chosen storage system. For example, 8,000,000 decisions per day at a measured 420 bytes per normalized event produce 3.36 GB of raw events per day and 100.8 GB across 30 days before overhead. Those are illustrative inputs, not benchmark results. Measure the serialized event from the actual implementation.
Cardinality needs a separate budget because byte volume and series count fail differently. Set a maximum allowed value count for every metric dimension. Reject or map unknown reason codes to other; do not pass arbitrary upstream strings through. Alert when the known vocabulary changes. A new provider adapter should ship with contract tests that prove its errors map into the existing taxonomy and that no payload field becomes a label.
For evaluation, keep a versioned corpus of representative code changes and expected policy outcomes. Run it before changing an adapter, classifier, prompt, or schema. Measure false accepts, false blocks, review rates, malformed-output rates, and deadline misses independently. Provider portability without behavioral evaluation merely makes replacement faster; it does not make replacement correct.
Asynchronous classification has a valid place outside the critical path. Large backfills, historical reclassification, and offline evaluation do not need to hold an interactive request open. Batch facilities can suit that work, but batch results should pass through the same schema validator and event taxonomy before entering reports. The cited Batch API guide documents one concrete batch mechanism; it does not turn a synchronous moderation decision into a batch workload.
Why reject inline metrics, and when are they valid?
I reject inline provider metrics for this B2B SaaS review path because each branch becomes an observability schema author. One branch may label a timeout by exception class, another by status family, and a third by raw message. The dashboard looks unified until the first switch, at which point comparisons become a data-cleaning project. The more damaging failure is quieter: a helpful engineer adds repo, model, or error as a label and the series count expands without a design review.
The rejected option is still valid under a narrow condition: one provider is an accepted architectural dependency, its emitted dimensions are fixed and reviewed, and the service has no portability objective. A small internal tool can reasonably choose that simplicity. It should document the dependency instead of claiming a portable abstraction it does not possess.
Events are not free.
The main limitation of the adapter design is its extra control plane. It adds a schema, translation code, contract tests, and responsibility for taxonomy changes. Normalization can hide provider-specific evidence, so the raw diagnostic channel must remain available under controlled access and retention. This trade-off is justified when portability is a requirement, but it is needless machinery for a small internal service with one fixed provider and no shared reporting surface. In that case, inline metrics are the better choice, provided their dimensions remain bounded.
Deployment should be incremental. Shadow the new adapter on an approved evaluation corpus, compare structured outcomes, then canary a bounded traffic slice while watching outcome distribution, schema failures, latency buckets, and diagnostic sample volume. Rollback selects the previous adapter configuration; it does not require reverting dashboards. No provider name belongs in the conclusion because the decision rule is independent of the current roster.
The resulting architecture is modest: one credential at ingress, a versioned moderation contract, adapters at the unstable edge, bounded decision events, and deliberately sampled evidence. For synchronous code-change review, choose adapter events. Choose inline metrics only when provider coupling is explicit, small, and expected to remain so.
Top comments (0)