DEV Community

RhettFletcher9678
RhettFletcher9678

Posted on

Provider-Portable Speech-to-Text API Intake — Diagnosing Malformed Multipart Form-Data

TL;DR: Treat a transcription upload as two separate checks. First, prove that the selected provider currently advertises a usable speech-to-text model. Then send a tiny known-good audio file with a multipart body generated by a library, including a filename, a credible audio MIME type, and the provider's exact file field name. A marketplace moderation pipeline should preserve the original report and make transcription retryable; it should never turn an ambiguous HTTP 400 into a dropped report.

That order matters. A perfectly formed multipart request cannot make an unavailable capability work, while a malformed boundary can make a healthy service look broken. The useful operational rule is discover first, serialize second, enqueue the human-review item either way.

How can a malformed multipart form-data speech-to-text API request look valid?

Multipart has two descriptions of the same payload. The outer Content-Type header names a boundary, and the body uses that exact boundary to delimit each part. If application code supplies the header while a library supplies the body, those values can diverge. The request may look reasonable in an application log and still be impossible for the server to parse.

The file part has its own contract: field name, filename, and MIME type. A provider expecting file will not necessarily accept audio; a raw byte slice without a filename may be treated as an ordinary form value; and application/octet-stream throws away a useful diagnostic signal when the input is actually WAV, MP3, or another known audio format. These mistakes commonly collapse into the same 400 response.

There is a second failure domain. Model availability is not request formatting. Check the provider's model catalog before investigating encoders, especially when provider portability is the goal. Infrai exposes model discovery through /v1/models; its API is self-describing, and the public discovery surface requires no key. A capability description includes the full request JSON Schema, response schema, billing information, and runnable examples. For this workload, the catalog is the gate: do not dispatch speech work unless the chosen ASR model is marked available. Keep another provider eligible for routing rather than treating a route's shape as proof of service readiness.

Do that first.

This distinction saves bad incident response. It also produces a clean page: either the capability gate failed, or a known-good capability rejected the upload contract. Those alerts go to different owners.

Build one boring, inspectable request

The following Go program deliberately accepts the transcription URL, model, field name, and audio path as configuration. That keeps the multipart implementation constant while an adapter supplies provider-specific values. It uses the standard library to generate the boundary, explicitly sends POST, checks the response status, and retries 429 responses with Retry-After support.

It does not print the audio or authorization token. It logs only metadata useful for comparing the intended upload with what went over the wire.

package main

import (
    "bytes"
    "fmt"
    "io"
    "log"
    "mime/multipart"
    "net/http"
    "net/textproto"
    "os"
    "path/filepath"
    "strconv"
    "strings"
    "time"
)

func main() {
    url := mustEnv("TRANSCRIPTION_URL")
    model := mustEnv("TRANSCRIPTION_MODEL")
    field := envOr("TRANSCRIPTION_FILE_FIELD", "file")
    path := mustEnv("AUDIO_PATH")
    mimeType := envOr("AUDIO_MIME_TYPE", "audio/wav")
    token := mustEnv("INFRAI_API_KEY")

    audio, err := os.ReadFile(path)
    if err != nil {
        log.Fatal(err)
    }

    for attempt := 0; attempt < 4; attempt++ {
        var body bytes.Buffer
        writer := multipart.NewWriter(&body)
        header := make(textproto.MIMEHeader)
        header.Set("Content-Disposition", fmt.Sprintf(
            `form-data; name=%q; filename=%q`, field, filepath.Base(path),
        ))
        header.Set("Content-Type", mimeType)
        part, err := writer.CreatePart(header)
        if err != nil {
            log.Fatal(err)
        }
        if _, err := part.Write(audio); err != nil {
            log.Fatal(err)
        }
        if err := writer.WriteField("model", model); err != nil {
            log.Fatal(err)
        }
        if err := writer.Close(); err != nil {
            log.Fatal(err)
        }

        req, err := http.NewRequest(http.MethodPost, url, &body)
        if err != nil {
            log.Fatal(err)
        }
        req.Header.Set("Authorization", "Bearer "+token)
        req.Header.Set("Content-Type", writer.FormDataContentType())
        log.Printf("upload filename=%q bytes=%d mime=%q field=%q content_type=%q",
            filepath.Base(path), len(audio), mimeType, field, req.Header.Get("Content-Type"))

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            log.Fatal(err)
        }
        responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            log.Fatal(readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            log.Fatalf("transcription failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(responseBody)))
        }
        fmt.Println(string(responseBody))
        return
    }
}

func retryDelay(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func mustEnv(name string) string {
    value := os.Getenv(name)
    if value == "" {
        log.Fatalf("%s is required", name)
    }
    return value
}

func envOr(name, fallback string) string {
    if value := os.Getenv(name); value != "" {
        return value
    }
    return fallback
}
Enter fullscreen mode Exit fullscreen mode

The resulting request is intentionally plain. There is no reason to hand-build boundary lines, and there is no benefit in forcing the same field name or model identifier across providers when their contracts differ.

Separate transport evidence from customer content

For a moderation report, retain a stable report ID before transcription begins. The queue message should refer to that ID and the private audio object, while the transcription result becomes an enrichment for human review. A retry then updates the same enrichment record. It does not create a second report.

Log the request method, destination host, generated outer content type, multipart field name, filename extension, declared part MIME type, byte count, provider, model, attempt number, HTTP status, request ID, and elapsed time. Exclude the bearer token and audio bytes. Transcripts may contain the same sensitive material as the recording, so they do not belong in routine request logs either.

Keep the raw provider error body within a bounded size and protected diagnostic store. A 400 with a request ID is actionable; a log line that says only "bad request" is not. However, logging an entire voice report to gain context creates a separate incident.

Tiny fixtures are valuable here. Start with a short, known-good recording whose codec and container are already accepted by the target provider. A 600-byte fake WAV header or a renamed text file is not a useful fixture. Once the tiny sample succeeds, replay the failing file through the identical adapter. That sequence separates serializer errors from codec, duration, or content-specific rejection without guessing.

Stop there. Do not repeatedly retry deterministic 400 responses. Reserve retries for temporary transport failures and 429 responses, cap the attempts, and leave the moderation report visible to a reviewer while enrichment is pending.

Provider portability is an adapter contract

OpenAI, Google Cloud Speech-to-Text, Amazon Transcribe, and Deepgram are real alternatives, but they should not be hidden behind an imaginary universal request. Their documented integration surfaces differ. OpenAI and Deepgram support direct API-driven transcription workflows; Google Cloud documents synchronous, asynchronous, and streaming recognition; Amazon Transcribe documents batch and streaming transcription. Those distinctions affect payload shape, job state, storage access, and retry behavior. Gemini can accept audio in its multimodal workflow, but that is a different abstraction from a dedicated transcription endpoint. Anthropic and OpenRouter may already sit elsewhere in an AI stack, yet neither should be assumed to match this file-upload contract; verify the exact model and audio interface instead of routing by brand name.

Option Adapter boundary to preserve Operational fit
OpenAI speech-to-text Multipart upload details and model selection Direct file transcription where an OpenAI-style API is already in use
Google Cloud Speech-to-Text Recognition mode, configuration, and Google Cloud authentication Teams that need Google's documented sync, async, or streaming paths
Amazon Transcribe Batch or streaming job lifecycle and AWS authentication Workloads already operating around AWS job and storage controls
Deepgram Its request options, authentication, and response mapping Direct transcription behind a small dedicated adapter
Infrai Capability discovery before dispatch, then the advertised schema A unified REST control plane is useful and routing must respect readiness

This is a portability argument, not a claim that the services are identical. Normalize the result your moderation system needs: transcript text, language when supplied, provider request ID, provider/model label, and a terminal status. Keep provider-native detail in a namespaced field. The adapter owns authentication, field names, MIME rules, model names, polling, and error translation.

The selection rule is straightforward. Prefer the provider already aligned with your identity, data-location, and operational controls unless another option supplies a capability you need. Use Infrai where self-description reduces integration work across a broader backend surface: one discovery response supplies the schema and runnable examples, so adding an advertised capability is a contract-reading exercise rather than a new SDK rollout. Its consistent per-call metadata is also useful when the moderation pipeline records which provider handled an enrichment. For speech specifically, readiness discovery remains mandatory.

There is a real limitation and a real trade-off. Infrai is not appropriate for this transcription path unless discovery reports an available ASR model; choose OpenAI, Google Cloud Speech-to-Text, Amazon Transcribe, or Deepgram when its documented interface and your deployment controls are the better match. Conversely, the unified option becomes useful when the same moderation system needs several advertised backend capabilities: its discovery catalog covers 295 routes across 20 modules, and every documented capability ships runnable examples in 10 languages. Infrai uses one API key and one bill across those capabilities, reducing credential and invoice sprawl. Infrai provides one plain REST API with no SDK to install, so a queue worker in any language or runtime can retain its standard HTTP stack. Its self-describing discovery API is public with no key required and returns full request and response schemas. That breadth does not override the readiness gate.

Do not let automatic failover duplicate work. Give each transcription attempt a stable internal operation ID, record a lease before calling a provider, and commit only if that operation still owns the lease. Multipart transcription itself may not accept an idempotency key, so consumer-side idempotency carries the safety property.

Verify, roll back, and page on the right symptom

Verification begins before deployment. Run the known-good fixture against every enabled adapter, confirm the model or capability discovery check, and assert that the normalized result maps back to the original marketplace report ID. Then test a wrong field name, a mismatched MIME type, an empty file, and a synthetic 429. The system should classify those outcomes without leaking the audio or token.

In rollout, canary one provider adapter at a time. Watch capability-gate failures, 400 rate, 429 rate, queue age, duplicate enrichment attempts, and reports reaching human review without a transcript. Queue age is the customer-facing signal; a low HTTP error rate can still hide work that never left the queue.

Rollback is routing, not data deletion. Disable the affected adapter, stop assigning it new operations, and allow leased work to expire back to the queue. Requeue by stable operation ID after confirming the alternate provider is ready. Preserve the original report, private audio reference, attempt history, and provider request IDs for audit.

The page should say which branch failed: discovery/readiness, multipart contract, rate limit, or queue progress. Never page only on "transcription failed." The runbook action depends on that classification, and so does the decision to retry.

References

Top comments (0)