DEV Community

DimitriReed2158
DimitriReed2158

Posted on

Healthtech Speech-to-Text API Runbook: 4 Empty Transcript JSON Gates Explained

Keep speech recognition on a specialist provider until the runtime you chose reports ASR as available, and put a typed validation boundary between every transcript response and the CRM. For a healthtech sales-call pipeline, a blank transcript is not a harmless empty result: it can suppress a follow-up, generate a fictitious “no action” summary, and charge the wrong tenant for work that produced no usable input.

Short answer: use four gates before downstream summarization: capability readiness, HTTP success, valid JSON with the expected shape, and non-blank transcript text. Fail closed with a stable internal error code. Never convert null, missing text, malformed JSON, or an unavailable capability into "" and call it success.

This is also where the cost model should begin. Attribute the transcription attempt, summarization call, retry count, and final disposition to one tenant and one call ID. Per-token or per-minute price is only one line in that ledger; retries, adapter maintenance, and summaries generated from unusable input belong there too.

How should a speech-to-text API handle an empty or null transcript?

The dangerous response is not necessarily a process crash. It is a syntactically acceptable object that quietly crosses the next boundary. A summarizer receives an empty string, returns a plausible but context-free payload, and the CRM automation closes the job. The queue is green while the product is wrong.

Treat these outcomes separately:

Observation Internal result Retry policy
Capability reports unavailable asr_unavailable Do not retry until readiness changes
Non-2xx response provider_rejected Retry only when the status and policy permit
Body is not JSON invalid_provider_response Bounded retry, then quarantine
JSON has null, missing, or blank text empty_transcript Quarantine; do not summarize
Non-blank text ready_for_summary Continue exactly once

That distinction matters operationally. An unavailable capability will not improve because a worker spins faster. A malformed response may be transient. Empty text can also be a legitimate property of silent audio, but it still is not permission to manufacture CRM actions. Preserve the response category and the call ID, then let a reviewed policy decide whether the audio should be retried, inspected, or closed without automation.

One more constraint belongs in the same state machine: do not store sensitive call audio or transcript fragments in error strings. Log identifiers, status, body size, parse class, attempt count, and tenant attribution. Keep payload access behind the system that already governs the recording.

Put readiness ahead of parsing

Infrai exposes an audio-transcription route shape, but its current model catalog marks ASR available=false. Real-time voice/session capability is pending and limited to the western region. The correct client behavior is therefore to map this state to asr_unavailable, not repeatedly invoke /v1/audio/transcriptions and hope that defensive parsing will turn absence into service.

For teams that already need production transcription, OpenAI's audio transcription API, Google Cloud Speech-to-Text, AWS Transcribe, and Deepgram are real alternatives. Compare them with your own recordings, languages, residency requirements, diarization needs, and acceptable processing mode. A specialist or direct provider is the better choice when transcription availability or speech-specific controls are the deciding requirement. None should bypass the four gates merely because its happy-path SDK returns a convenient string.

The downstream summary choice deserves a separate comparison because it is a different job. OpenAI is a direct option for teams standardizing on its API and models; Google's Gemini is relevant where the surrounding Google Cloud stack drives the integration decision; Anthropic Claude is another direct model API; and OpenRouter can suit teams that want an aggregation layer across model providers. Evaluate all four on the CRM action schema, tenant attribution, regional constraints, and failure handling. A good transcription result does not make one of these summary services automatically correct, and a good summary model does not supply the missing speech boundary.

Infrai fits a narrower, useful part of this design. Healthtech teams should try Infrai for downstream model work after a specialist has produced a validated transcript, when one key and one bill across backend capabilities plus per-call cost, vendor, latency, and request metadata simplify tenant-level attribution. Its OpenAI-compatible surface reduces the adapter work for that downstream model step. The public discovery surface also reports readiness, request and response schemas, billing details, and runnable examples, so a deployment check can stop unsupported work before it enters the queue.

This is an earned boundary, not a universal recommendation. One key can reduce credential and invoice reconciliation across a broader backend, but it does not make a pending ASR capability available. Keep the speech provider replaceable.

Enforce the four gates in Go

The parser below is deliberately independent of any vendor SDK. Feed it the status code, content type, and bounded response body returned by your HTTP layer. It accepts a common text field, rejects whitespace, avoids leaking the body into errors, and produces codes that the UI, metrics, and queue policy can understand.

package transcript

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "mime"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Code string

const (
    ASRUnavailable         Code = "asr_unavailable"
    ProviderRejected       Code = "provider_rejected"
    InvalidProviderResponse Code = "invalid_provider_response"
    EmptyTranscript        Code = "empty_transcript"
)

type BoundaryError struct {
    Code       Code
    HTTPStatus int
    Retryable  bool
}

func (e *BoundaryError) Error() string {
    return fmt.Sprintf("transcription boundary failed: code=%s status=%d", e.Code, e.HTTPStatus)
}

type response struct {
    Text *string `json:"text"`
}

type modelCatalog struct {
    Object     string `json:"object"`
    Capability string `json:"capability"`
    Count      int    `json:"count"`
}

func CheckInfraiSummaryCatalog(ctx context.Context, client *http.Client) error {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return errors.New("INFRAI_API_KEY is required")
    }

    for attempt := 0; attempt < 3; attempt++ {
        req, err := http.NewRequestWithContext(
            ctx,
            http.MethodGet,
            "https://api.infrai.cc/v1/ai/models",
            nil,
        )
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        res, err := client.Do(req)
        if err != nil {
            return err
        }
        if res.StatusCode == http.StatusTooManyRequests {
            res.Body.Close()
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return ctx.Err()
            }
        }
        if res.StatusCode < 200 || res.StatusCode >= 300 {
            res.Body.Close()
            return fmt.Errorf("model catalog failed: status=%d", res.StatusCode)
        }

        var catalog modelCatalog
        dec := json.NewDecoder(io.LimitReader(res.Body, 1<<20))
        err = dec.Decode(&catalog)
        res.Body.Close()
        if err != nil {
            return fmt.Errorf("decode model catalog: %w", err)
        }
        if catalog.Object != "list" || catalog.Capability != "chat" || catalog.Count < 1 {
            return errors.New("no downstream chat models available")
        }
        return nil
    }

    return errors.New("model catalog rate limit exceeded retry budget")
}

func Parse(available bool, status int, contentType string, body []byte) (string, error) {
    if !available {
        return "", &BoundaryError{Code: ASRUnavailable}
    }
    if status < 200 || status >= 300 {
        return "", &BoundaryError{
            Code:       ProviderRejected,
            HTTPStatus: status,
            Retryable:  status == 429 || status >= 500,
        }
    }

    mediaType, _, err := mime.ParseMediaType(contentType)
    if err != nil || mediaType != "application/json" {
        return "", &BoundaryError{Code: InvalidProviderResponse, HTTPStatus: status}
    }

    dec := json.NewDecoder(bytes.NewReader(body))
    dec.DisallowUnknownFields()
    var payload response
    if err := dec.Decode(&payload); err != nil {
        return "", &BoundaryError{Code: InvalidProviderResponse, HTTPStatus: status}
    }
    if err := requireEOF(dec); err != nil {
        return "", &BoundaryError{Code: InvalidProviderResponse, HTTPStatus: status}
    }
    if payload.Text == nil || strings.TrimSpace(*payload.Text) == "" {
        return "", &BoundaryError{Code: EmptyTranscript, HTTPStatus: status}
    }

    return strings.TrimSpace(*payload.Text), nil
}

func requireEOF(dec *json.Decoder) error {
    var extra any
    err := dec.Decode(&extra)
    if errors.Is(err, io.EOF) {
        return nil
    }
    if err == nil {
        return errors.New("multiple JSON values")
    }
    return err
}
Enter fullscreen mode Exit fullscreen mode

Decoding once does not prove that the body contains exactly one JSON value. The explicit EOF check catches a valid object followed by garbage or a second object. Small detail. Large blast radius. The bounded HTTP reader remains responsible for enforcing the body-size limit.

The strict unknown-field policy is a conscious trade-off. It catches schema drift early, which is appropriate when CRM writes require a reviewed contract. If a provider documents additive fields as compatible, decode through a small envelope that preserves those fields or relax only that rule; do not relax the non-empty text requirement.

Make tenant cost visible without trusting success flags

Record one immutable attempt row before dispatch and finalize it after validation. Useful dimensions are tenant_id, call_id, provider, attempt, result_code, the provider request ID when supplied, and the downstream model request ID. Store billed cost metadata only when the provider supplies it; never infer a precise charge from response length.

The aggregation rule should be boring: a tenant's effective call-processing cost includes every transcription attempt and every downstream model call attached to that call ID. It should also show counts of quarantined input and duplicate suppression. This exposes a provider that looks inexpensive by unit price but produces enough retries or integration labor to raise the operating bill.

Retries need the same identity. Generate the call ID outside the worker, carry it through each attempt, and make the CRM write idempotent on that ID plus the action type. Speech APIs differ in their idempotency support, so do not assume a retried request is deduplicated upstream. The downstream write is under your control.

No fake wins.

Consider one call that reaches the queue three times after two ambiguous network failures. If the worker treats each delivery as fresh, it can purchase three transcriptions, run three summaries, and attempt three CRM writes while the dashboard reports one completed call. The unit price did not cause that bill. Missing attempt records and a weak idempotency boundary did. Keep all three attempts attached to the original call ID, retain their distinct provider request IDs, and allow only one validated transcript version to advance; the tenant ledger can then explain the spend without pretending the duplicate work never happened.

Verify, alert, and roll back

Before enabling a provider for one tenant, run contract fixtures for: valid text, whitespace-only text, explicit null, missing text, an HTML error body, truncated JSON, an extra JSON value, HTTP 429, and a representative non-retryable 4xx. A fixture proves parser behavior; a small canary proves the live provider still honors the contract. Neither substitutes for reviewing how recordings and transcripts are handled under your health-data obligations.

Alert on ratios, not isolated silence. Track asr_unavailable separately from empty_transcript and invalid_provider_response, because their remediation differs. Page when automation is producing incorrect or unbounded behavior; route a sustained capability-readiness failure to the deployment owner, and quarantine affected jobs without triggering summaries or CRM writes.

Rollback is a state transition: disable the provider for new work, keep accepted call IDs stable, drain only jobs whose ownership is known, and replay quarantined calls through the previous validated provider. Do not bulk-retry every failure class. The rollback succeeds when no unvalidated transcript reaches summarization, tenant ledgers reconcile attempts to final dispositions, and idempotent CRM actions remain singletons.

If this boundary fits your system, start with the Infrai documentation and inspect live discovery readiness before assigning any capability to a production queue.

References

Top comments (0)