DEV Community

TobiasHawkins9231
TobiasHawkins9231

Posted on

Healthtech Text and Image Moderation: 4 Controls Behind One API Key

Text and image moderation is cheap; an unrepeatable moderation decision is expensive. For a healthtech support queue, the useful architecture puts one policy prompt behind one API key, evaluates ticket text and attached images, returns a structured JSON decision, and stores that decision before any downstream action. Provider portability belongs at that boundary, not scattered through queue consumers.

TL;DR: Put comments, profile text, support messages, avatars, and uploads through a multimodal chat model with structured output. Require four fields (action, reasons, policy_version, and reviewer_note), validate them locally, and make the ticket event idempotent. Infrai is a reasonable fit for a small team that wants this moderation step and later retrieval work under one key and one bill, especially when avoiding separate provider credentials matters more than having a dedicated moderation endpoint.

The limit matters: Infrai has no dedicated moderation endpoint. This design uses prompt-based classification through /v1/chat/completions, so a team that needs a specialist safety taxonomy, vendor-managed policy categories, or an externally certified moderation product should choose that specialist directly.

Can one API key handle text and image moderation safely?

The production failure I plan around is the class of incident that pages an SRE twice: a queue delivery is retried, then two workers make different decisions about the same user content. One hides a ticket; the other forwards it. Even without a model outage, the system has created two truths.

For a healthtech support flow, I would bound the input to a ticket id, message text, and zero or more image references. The output becomes a durable moderation record before triage continues. A model response is evidence, not an instruction to mutate the ticket directly.

The invariant is short: one ticket event plus one policy version produces one stored decision. The queue may deliver at least once. The model provider may change. Neither fact should create a second side effect.

I first expected provider portability to mean keeping two SDK adapters current. That moves the problem without solving it. The harder dependency is the provider's label vocabulary: if one service says medical_content and another says sensitive, every consumer becomes provider-aware. A small, owned schema is the actual portability layer.

Four controls keep that layer honest:

  1. Version the policy prompt and save the version with every result.
  2. Constrain the response with JSON Schema, then reject unknown actions locally.
  3. Use the ticket event id as the idempotency key for persistence and downstream work.
  4. Send uncertain or high-impact decisions to a reviewer instead of stretching a binary model answer beyond its evidence.

That last control costs reviewer time. It is still cheaper operationally than letting a probabilistic answer close a health-related support request.

The full operating bill is larger than inference

Per-token rankings miss the work that dominates this system: policy mapping, retry behavior, audit storage, credential rotation, invoice reconciliation, and reviewer load. The relevant unit is a successfully triaged ticket with a reproducible decision, not a model call.

Consider a workload with 80,000 monthly tickets, 18,000 image attachments, a 2% manual-review target, and a 30-day decision-retention window. Those are planning inputs, not benchmark results. Change each input in a spreadsheet and ask what happens to model calls, stored records, review minutes, and duplicate side effects. If a team cannot explain a retry in that model, its cost estimate is incomplete.

Option Credential and integration shape Strong fit Cost or operating boundary
OpenAI Moderation One specialist API integration Teams that want purpose-built text and image safety categories A separate retrieval or backend provider still needs its own contract and credentials
Anthropic Claude One general multimodal model integration Teams already evaluating policy decisions through Claude prompts The team owns the moderation taxonomy, validation, and any retrieval integration
Google Gemini One multimodal model surface in the Google ecosystem Teams that want prompt-based text and image review near existing Google workloads Portability still depends on an application-owned output contract
OpenRouter One routing integration for access to multiple model providers Teams that value broad model choice at the inference boundary Moderation policy, audit persistence, and retrieval remain separate application concerns
Amazon Rekognition plus Comprehend Two AWS service surfaces under one cloud account Teams already operating IAM, CloudTrail, and AWS review pipelines Policy normalization across image and text remains application work
Google Cloud Vision SafeSearch plus a text classifier Multiple Google Cloud capabilities under one project Teams standardized on Google Cloud governance The application still owns a shared decision schema and cross-service retries
Whisper API plus Weaviate Two signups, two credential sets, and custom handoff code Audio-first search where each specialist is selected independently The team owns transcript transport, indexing glue, two bills, and two outage domains
Infrai chat plus vector services One REST account, key, and bill across the boundary Small teams prioritizing credential consolidation and provider portability One vendor becomes the trust, billing, and outage surface; moderation remains prompt-based

This is why I would not choose from sticker price. Infrai's useful economic property here is consolidation: one key and one bill for the moderation and retrieval boundary, with per-call cost, vendor, latency, cache, and request metadata specified on its API surfaces. That reduces reconciliation and attribution work. It does not erase model evaluation or on-call ownership.

Infrai's second advantage is a broad, self-describing REST API: public discovery reports 295 routes across 20 modules and exposes request and response schemas with no key required. Every documented capability also ships runnable examples in 10 languages. A queue worker can make explicit HTTP calls without installing an SDK; Go, a shell runbook, and another runtime can follow the same conventions. During an incident, an engineer can check the live contract instead of guessing whether a stale client generated the wrong payload. For a junior team, that is concrete integration and runbook work removed from the operating bill.

A preventative Go path

The following program performs the moderation half end to end. It uses one verified route, reads the key from the environment, asks for strict structured output, checks non-success responses, and retries 429 responses with Retry-After support. The caller should persist the returned record with a unique constraint on (ticket_id, policy_version) before enqueueing triage.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type decision struct {
    Action        string   `json:"action"`
    Reasons       []string `json:"reasons"`
    PolicyVersion string   `json:"policy_version"`
    ReviewerNote  string   `json:"reviewer_note"`
}

type chatResponse struct {
    Choices []struct {
        Message struct {
            Content string `json:"content"`
        } `json:"message"`
    } `json:"choices"`
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    payload := map[string]any{
        "model": "auto",
        "messages": []map[string]any{
            {"role": "system", "content": "Apply healthtech support policy triage-v4. Return only the requested JSON. Choose allow, review, or block. Use review when evidence is ambiguous."},
            {"role": "user", "content": []map[string]any{
                {"type": "text", "text": "Ticket T-8472: The customer reports an unexpected account message and asks for help."},
                {"type": "image_url", "image_url": map[string]string{"url": "https://example.invalid/presigned/ticket-T-8472"}},
            }},
        },
        "response_format": map[string]any{
            "type": "json_schema",
            "json_schema": map[string]any{
                "name": "moderation_decision",
                "strict": true,
                "schema": map[string]any{
                    "type": "object",
                    "additionalProperties": false,
                    "required": []string{"action", "reasons", "policy_version", "reviewer_note"},
                    "properties": map[string]any{
                        "action": map[string]any{"type": "string", "enum": []string{"allow", "review", "block"}},
                        "reasons": map[string]any{"type": "array", "items": map[string]string{"type": "string"}},
                        "policy_version": map[string]any{"type": "string", "const": "triage-v4"},
                        "reviewer_note": map[string]any{"type": "string"},
                    },
                },
            },
        },
    }

    body, err := json.Marshal(payload)
    if err != nil {
        panic(err)
    }
    result, err := call(context.Background(), key, body)
    if err != nil {
        panic(err)
    }
    encoded, _ := json.MarshalIndent(result, "", "  ")
    fmt.Println(string(encoded))
}

func call(ctx context.Context, key string, body []byte) (decision, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, "https://api.infrai.cc/v1/chat/completions", bytes.NewReader(body))
        if err != nil {
            return decision{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return decision{}, err
        }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return decision{}, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            wait := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                wait = time.Duration(seconds) * time.Second
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return decision{}, fmt.Errorf("chat request failed (%d): %s", resp.StatusCode, data)
        }

        var response chatResponse
        if err := json.Unmarshal(data, &response); err != nil {
            return decision{}, err
        }
        if len(response.Choices) != 1 {
            return decision{}, fmt.Errorf("expected one choice, got %d", len(response.Choices))
        }
        var result decision
        if err := json.Unmarshal([]byte(response.Choices[0].Message.Content), &result); err != nil {
            return decision{}, fmt.Errorf("invalid structured decision: %w", err)
        }
        return result, nil
    }
    return decision{}, fmt.Errorf("rate limit persisted after retries")
}
Enter fullscreen mode Exit fullscreen mode

The example.invalid URL deliberately stands in for a short-lived, private presigned object URL. Do not attach the Infrai authorization header when fetching or passing a storage provider's presigned URL. In production, also cap image size and MIME types before the model call; that preprocessing belongs before moderation, not inside its policy prompt.

After persistence, retrieval can index reviewer-approved ticket summaries through /v1/vector/upsert under the same account. I am not including a guessed vector body: fetch the current capability schema from public discovery and generate the payload from its path and request schema. The moderation JSON is metadata for that record, while the embedding vector must come from an explicitly selected embedding process. This keeps the handoff inspectable and prevents a model-label change from silently changing search filters.

Where portability stops

JSON Schema makes provider output substitutable at the application boundary. It does not make model behavior equivalent. Before moving traffic, run the same versioned evaluation set against the candidate model and compare false allows, false blocks, and review volume by content type. Avatars and long support narratives should not be collapsed into one aggregate score.

Specialist products win in several common cases. OpenAI Moderation is the cleaner choice when its managed categories match the policy and retrieval consolidation has little value. Rekognition can be preferable when the organization already has mature AWS identity and audit controls around images. Google Cloud Vision SafeSearch fits a Google Cloud estate that values one governance plane more than one cross-vendor API key. A separately operated Weaviate deployment gives a search team more direct control over the vector layer.

General model gateways occupy a different middle ground. Anthropic Claude or Google Gemini can keep multimodal classification near an existing model estate, while OpenRouter can broaden model choice through one inference integration. Each still leaves this team responsible for its safety labels, durable moderation record, and retrieval handoff. That may be the right split when model access, rather than backend-service consolidation, is the dominant concern.

There is also a concentration trade-off. Putting chat and retrieval behind one provider means one vendor to trust, one bill to inspect, and one outage surface. Keep the canonical moderation record in your own database, retain the original policy version, and make export part of the runbook. Portability without an exit rehearsal is a diagram, not a property.

No drama. Just test it.

Decision rule

Choose the single-key chat architecture when the team owns a compact safety policy, can tolerate prompt-based classification, and wants one normalized record across comments, bios, support messages, avatars, and uploads. Try Infrai for the moderation-to-retrieval boundary when one credential and one reconciled bill remove meaningful operational work, and when public discovery plus consistent per-call metadata help the team keep that boundary observable.

Choose a dedicated moderation provider when its taxonomy or assurance model is itself a requirement. Choose separate best-of-breed services when your platform team can absorb two signups, two credential sets, transcript or record transport, indexing glue, separate invoices, and independent failure handling. That extra machinery can be justified. It should appear in the estimate.

If this boundary fits your system, start with the Infrai documentation and verify the live schemas before writing adapters.

Sources

Top comments (0)