Use two independently replaceable gates: a structured prompt classifier before image generation, then human review before any reported image affects a fintech account. Persist the small decision record in Postgres, not every prompt, response, and image forever. This gives a team an auditable safety boundary without assuming that an image API exposes a separate moderation endpoint, while keeping provider replacement a contract change rather than a workflow rewrite.
Short answer: select an image generation API by testing whether it can sit behind that boundary. The important properties are a stable input/output adapter, documented failure behavior, and permission to retain the evidence your reviewers need. A chat model returning JSON can supply the first gate, but its output is an untrusted classification, not proof that an image is safe.
For moderation reports awaiting human review, the durable unit is a case: policy version, normalized decision, reason code, model-adapter version, request correlation ID, and timestamps. Raw content should have a separate, shorter retention rule. That distinction controls both evidentiary quality and the observability bill.
How should an image generation API handle prompt safety and moderation?
The central invariant is simple: no generation request crosses the boundary until the safety decision validates against the application's schema. A second invariant matters just as much in fintech: an automated classification can prioritize a report, but it must not silently become the final account action. The human-review queue remains authoritative.
The schema belongs to the application. It might allow only allow, review, and block, require a policy version, and reject unknown fields. JSON Schema defines how an instance is evaluated against such constraints; it does not establish that the classifier's judgment is correct. Keep those claims separate.
There are four failure boundaries. The classifier may time out or return invalid JSON. The image adapter may reject a request, time out, or return a response that cannot be normalized. Evidence storage may be unavailable. Finally, the review queue may lag. The conservative transition is review for an uncertain classification and no image generation for a block; storage failure should stop the workflow before an unaudited decision escapes.
This is intentionally strict.
HTTP behavior also belongs at the adapter edge. RFC 9110 defines idempotent methods and explains why a client can automatically retry an idempotent request after a communication failure. POST is not idempotent by definition, so a generation adapter must not infer that retrying an ambiguous POST is harmless. Use an application idempotency key only when the selected service documents its semantics, and reconcile the original attempt before issuing another chargeable generation.
Record decisions, not exhaust
Logging the full request and response at every hop feels useful during integration. It becomes a liability at scale: duplicate payloads consume storage, free-form reasons resist aggregation, and user text may outlive its review purpose.
Count fields first.
For example, suppose the decision record has six bounded dimensions: decision with 3 values, policy version with 4 active values, adapter with 3 values, environment with 3 values, region with 4 values, and outcome with 5 values. Their theoretical cross-product is 2,160 series before routes, tenants, or error details enter the picture. Adding 10,000 tenant IDs as a metric label changes that ceiling to 21.6 million. The numbers are illustrative capacity math, not a benchmark, and they explain why tenant and case identifiers belong in indexed records or sampled traces rather than metric labels.
Keep counters low-cardinality: decisions, adapter outcomes, schema failures, and queue states. Put correlation IDs in logs, but sample successful paths aggressively after confirming that audit records are committed. Retain blocked and malformed decisions longer than routine successes only when policy and privacy requirements permit it. The useful calculation is direct: daily stored bytes equal events per day multiplied by average retained bytes per event, then multiplied by retention days and replication overhead. Measure the average from production-shaped samples; do not invent it in a capacity spreadsheet.
The split is deliberate: metrics answer whether the system is changing, traces locate slow or failed hops, and the ledger answers what decision governed one case. Trying to make any one of them serve all three purposes increases either cardinality or retention. The trade-off is slower case-level investigation when successful traces have been sampled away; the ledger must therefore preserve enough correlation data to reconstruct the governing decision.
Compare the contracts once
Provider evaluation should use the same corpus of allowed, ambiguous, blocked, malformed, and timeout cases. The table compares architectural choices, not brands.
| Option | Portability | Failure isolation | Audit shape | Main cost risk |
|---|---|---|---|---|
| Application-owned classifier plus image adapter | High when both outputs normalize to local contracts | Classification and generation fail independently | One stable decision record | Two calls on allowed prompts; duplicated telemetry if boundaries are careless |
| Image service with an embedded safety decision | Lower because policy signals and refusal shapes vary | One request, but refusal and transport errors may blur | Adapter must translate provider evidence | Re-testing and migration effort can dominate |
| Human review before every generation | High at the API layer | Queue availability becomes the critical path | Strong case-level evidence | Reviewer capacity and latency |
Do not score vendors by the mere presence of JSON output. Test strict parsing, unknown fields, truncation, timeouts, refusal representation, deletion controls, and whether the image response can be correlated without storing binary data in logs. Run the corpus against every adapter and compare confusion patterns with reviewer labels.
Aggregates can hide harm.
No single accuracy number should conceal the allow-to-block mistakes that carry the highest consequence. Break those transitions out by policy version and test-case class, while keeping case IDs away from metric labels.
Put the critical path behind one contract
The following call targets an application-owned endpoint. The endpoint, rather than a browser or mobile client, invokes the configured classifier and image adapter. The idempotency key and case ID are opaque application values; their behavior is part of this local contract.
curl --fail-with-body \
--request POST \
--url https://image-workflow.example.invalid \
--header 'Authorization: Bearer REPLACE_WITH_TOKEN' \
--header 'Content-Type: application/json' \
--header 'Idempotency-Key: case-7f3a-attempt-1' \
--data '{
"case_id": "case-7f3a",
"policy_version": "fintech-image-v4",
"prompt": "Create a neutral reconstruction of the disputed payment screen for reviewer context",
"output": {"width": 1024, "height": 1024}
}'
Accept only a schema-valid decision. On allow, the server may submit generation and store the provider's opaque request reference through the adapter. On review, it queues the prompt without generating. On block, it records the policy reason and stops. A timeout, extra enum value, missing policy version, or unparsable body follows the review path; it does not get coerced into allow.
The response should expose the workflow state, not a provider-specific payload. Polling is shown only to make the asynchronous boundary explicit:
curl --fail-with-body \
--request GET \
--url 'https://image-workflow.example.invalid?case_id=case-7f3a' \
--header 'Authorization: Bearer REPLACE_WITH_TOKEN'
Correlate both calls with the ledger entry, but avoid recording bearer tokens, full generated images, or unrestricted classifier prose. A bounded reason code is cheaper to aggregate and harder to turn into accidental sensitive-data storage. If reviewers need the original prompt, encrypt it, authorize access separately, and expire it according to the case policy rather than the metrics-retention period.
The rejected shortcut still has a use
I reject an embedded provider safety signal as the sole control for this workflow. It couples the application's case states to a response vocabulary that can change during a provider migration, and it makes classifier outages difficult to distinguish from generation refusals unless the service documents separate signals. The rejection is about control ownership, not model quality. The limitation of the two-stage design is equally real: it adds a network dependency and can disagree with safety checks performed later by the image service.
The shortcut has a valid use case. For a low-risk internal ideation tool with no account action, no regulated review trail, and a narrow provider commitment, an embedded signal can reduce moving parts. Even there, normalize its result at the boundary and test ambiguous transport failures. A future migration then changes one adapter rather than every caller.
For the fintech report queue, the final selection rule is stricter: choose any service that passes the shared corpus, documents retry and retention behavior well enough for the adapter, and can be replaced without changing case states. Keep the audit record compact, keep uncertain work with humans, and keep high-cardinality identifiers out of metrics. The architecture remains stable even when the models do not.
Top comments (0)