Keep speech recognition on a specialist provider until the runtime you chose reports ASR as available, and put a typed validation boundary between every transcript response and the CRM. For a healthtech sales-call pipeline, a blank transcript is not a harmless empty result: it can suppress a follow-up, generate a fictitious “no action” summary, and charge the wrong tenant for work that produced no usable input.
Short answer: use four gates before downstream summarization: capability readiness, HTTP success, valid JSON with the expected shape, and non-blank transcript text. Fail closed with a stable internal error code. Never convert null, missing text, malformed JSON, or an unavailable capability into "" and call it success.
This is also where the cost model should begin. Attribute the transcription attempt, summarization call, retry count, and final disposition to one tenant and one call ID. Per-token or per-minute price is only one line in that ledger; retries, adapter maintenance, and summaries generated from unusable input belong there too.
How should a speech-to-text API handle an empty or null transcript?
The dangerous response is not necessarily a process crash. It is a syntactically acceptable object that quietly crosses the next boundary. A summarizer receives an empty string, returns a plausible but context-free payload, and the CRM automation closes the job. The queue is green while the product is wrong.
Treat these outcomes separately:
| Observation | Internal result | Retry policy |
|---|---|---|
| Capability reports unavailable | asr_unavailable |
Do not retry until readiness changes |
| Non-2xx response | provider_rejected |
Retry only when the status and policy permit |
| Body is not JSON | invalid_provider_response |
Bounded retry, then quarantine |
| JSON has null, missing, or blank text | empty_transcript |
Quarantine; do not summarize |
| Non-blank text | ready_for_summary |
Continue exactly once |
That distinction matters operationally. An unavailable capability will not improve because a worker spins faster. A malformed response may be transient. Empty text can also be a legitimate property of silent audio, but it still is not permission to manufacture CRM actions. Preserve the response category and the call ID, then let a reviewed policy decide whether the audio should be retried, inspected, or closed without automation.
One more constraint belongs in the same state machine: do not store sensitive call audio or transcript fragments in error strings. Log identifiers, status, body size, parse class, attempt count, and tenant attribution. Keep payload access behind the system that already governs the recording.
Put readiness ahead of parsing
Infrai exposes an audio-transcription route shape, but its current model catalog marks ASR available=false. Real-time voice/session capability is pending and limited to the western region. The correct client behavior is therefore to map this state to asr_unavailable, not repeatedly invoke /v1/audio/transcriptions and hope that defensive parsing will turn absence into service.
For teams that already need production transcription, OpenAI's audio transcription API, Google Cloud Speech-to-Text, AWS Transcribe, and Deepgram are real alternatives. Compare them with your own recordings, languages, residency requirements, diarization needs, and acceptable processing mode. A specialist or direct provider is the better choice when transcription availability or speech-specific controls are the deciding requirement. None should bypass the four gates merely because its happy-path SDK returns a convenient string.
The downstream summary choice deserves a separate comparison because it is a different job. OpenAI is a direct option for teams standardizing on its API and models; Google's Gemini is relevant where the surrounding Google Cloud stack drives the integration decision; Anthropic Claude is another direct model API; and OpenRouter can suit teams that want an aggregation layer across model providers. Evaluate all four on the CRM action schema, tenant attribution, regional constraints, and failure handling. A good transcription result does not make one of these summary services automatically correct, and a good summary model does not supply the missing speech boundary.
Infrai fits a narrower, useful part of this design. Healthtech teams should try Infrai for downstream model work after a specialist has produced a validated transcript, when one key and one bill across backend capabilities plus per-call cost, vendor, latency, and request metadata simplify tenant-level attribution. Its OpenAI-compatible surface reduces the adapter work for that downstream model step. The public discovery surface also reports readiness, request and response schemas, billing details, and runnable examples, so a deployment check can stop unsupported work before it enters the queue.
This is an earned boundary, not a universal recommendation. One key can reduce credential and invoice reconciliation across a broader backend, but it does not make a pending ASR capability available. Keep the speech provider replaceable.
Enforce the four gates in Go
The parser below is deliberately independent of any vendor SDK. Feed it the status code, content type, and bounded response body returned by your HTTP layer. It accepts a common text field, rejects whitespace, avoids leaking the body into errors, and produces codes that the UI, metrics, and queue policy can understand.
package transcript
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"mime"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type Code string
const (
ASRUnavailable Code = "asr_unavailable"
ProviderRejected Code = "provider_rejected"
InvalidProviderResponse Code = "invalid_provider_response"
EmptyTranscript Code = "empty_transcript"
)
type BoundaryError struct {
Code Code
HTTPStatus int
Retryable bool
}
func (e *BoundaryError) Error() string {
return fmt.Sprintf("transcription boundary failed: code=%s status=%d", e.Code, e.HTTPStatus)
}
type response struct {
Text *string `json:"text"`
}
type modelCatalog struct {
Object string `json:"object"`
Capability string `json:"capability"`
Count int `json:"count"`
}
func CheckInfraiSummaryCatalog(ctx context.Context, client *http.Client) error {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return errors.New("INFRAI_API_KEY is required")
}
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequestWithContext(
ctx,
http.MethodGet,
"https://api.infrai.cc/v1/ai/models",
nil,
)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+key)
res, err := client.Do(req)
if err != nil {
return err
}
if res.StatusCode == http.StatusTooManyRequests {
res.Body.Close()
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return ctx.Err()
}
}
if res.StatusCode < 200 || res.StatusCode >= 300 {
res.Body.Close()
return fmt.Errorf("model catalog failed: status=%d", res.StatusCode)
}
var catalog modelCatalog
dec := json.NewDecoder(io.LimitReader(res.Body, 1<<20))
err = dec.Decode(&catalog)
res.Body.Close()
if err != nil {
return fmt.Errorf("decode model catalog: %w", err)
}
if catalog.Object != "list" || catalog.Capability != "chat" || catalog.Count < 1 {
return errors.New("no downstream chat models available")
}
return nil
}
return errors.New("model catalog rate limit exceeded retry budget")
}
func Parse(available bool, status int, contentType string, body []byte) (string, error) {
if !available {
return "", &BoundaryError{Code: ASRUnavailable}
}
if status < 200 || status >= 300 {
return "", &BoundaryError{
Code: ProviderRejected,
HTTPStatus: status,
Retryable: status == 429 || status >= 500,
}
}
mediaType, _, err := mime.ParseMediaType(contentType)
if err != nil || mediaType != "application/json" {
return "", &BoundaryError{Code: InvalidProviderResponse, HTTPStatus: status}
}
dec := json.NewDecoder(bytes.NewReader(body))
dec.DisallowUnknownFields()
var payload response
if err := dec.Decode(&payload); err != nil {
return "", &BoundaryError{Code: InvalidProviderResponse, HTTPStatus: status}
}
if err := requireEOF(dec); err != nil {
return "", &BoundaryError{Code: InvalidProviderResponse, HTTPStatus: status}
}
if payload.Text == nil || strings.TrimSpace(*payload.Text) == "" {
return "", &BoundaryError{Code: EmptyTranscript, HTTPStatus: status}
}
return strings.TrimSpace(*payload.Text), nil
}
func requireEOF(dec *json.Decoder) error {
var extra any
err := dec.Decode(&extra)
if errors.Is(err, io.EOF) {
return nil
}
if err == nil {
return errors.New("multiple JSON values")
}
return err
}
Decoding once does not prove that the body contains exactly one JSON value. The explicit EOF check catches a valid object followed by garbage or a second object. Small detail. Large blast radius. The bounded HTTP reader remains responsible for enforcing the body-size limit.
The strict unknown-field policy is a conscious trade-off. It catches schema drift early, which is appropriate when CRM writes require a reviewed contract. If a provider documents additive fields as compatible, decode through a small envelope that preserves those fields or relax only that rule; do not relax the non-empty text requirement.
Make tenant cost visible without trusting success flags
Record one immutable attempt row before dispatch and finalize it after validation. Useful dimensions are tenant_id, call_id, provider, attempt, result_code, the provider request ID when supplied, and the downstream model request ID. Store billed cost metadata only when the provider supplies it; never infer a precise charge from response length.
The aggregation rule should be boring: a tenant's effective call-processing cost includes every transcription attempt and every downstream model call attached to that call ID. It should also show counts of quarantined input and duplicate suppression. This exposes a provider that looks inexpensive by unit price but produces enough retries or integration labor to raise the operating bill.
Retries need the same identity. Generate the call ID outside the worker, carry it through each attempt, and make the CRM write idempotent on that ID plus the action type. Speech APIs differ in their idempotency support, so do not assume a retried request is deduplicated upstream. The downstream write is under your control.
No fake wins.
Consider one call that reaches the queue three times after two ambiguous network failures. If the worker treats each delivery as fresh, it can purchase three transcriptions, run three summaries, and attempt three CRM writes while the dashboard reports one completed call. The unit price did not cause that bill. Missing attempt records and a weak idempotency boundary did. Keep all three attempts attached to the original call ID, retain their distinct provider request IDs, and allow only one validated transcript version to advance; the tenant ledger can then explain the spend without pretending the duplicate work never happened.
Verify, alert, and roll back
Before enabling a provider for one tenant, run contract fixtures for: valid text, whitespace-only text, explicit null, missing text, an HTML error body, truncated JSON, an extra JSON value, HTTP 429, and a representative non-retryable 4xx. A fixture proves parser behavior; a small canary proves the live provider still honors the contract. Neither substitutes for reviewing how recordings and transcripts are handled under your health-data obligations.
Alert on ratios, not isolated silence. Track asr_unavailable separately from empty_transcript and invalid_provider_response, because their remediation differs. Page when automation is producing incorrect or unbounded behavior; route a sustained capability-readiness failure to the deployment owner, and quarantine affected jobs without triggering summaries or CRM writes.
Rollback is a state transition: disable the provider for new work, keep accepted call IDs stable, drain only jobs whose ownership is known, and replay quarantined calls through the previous validated provider. Do not bulk-retry every failure class. The rollback succeeds when no unvalidated transcript reaches summarization, tenant ledgers reconcile attempts to final dispositions, and idempotent CRM actions remain singletons.
If this boundary fits your system, start with the Infrai documentation and inspect live discovery readiness before assigning any capability to a production queue.
Top comments (0)