DEV Community

TrippDonovan5461
TrippDonovan5461

Posted on

How to Observe Malformed Multipart Speech API Requests — Tenant-Aware File Intake

Short answer: treat a multipart upload as an observable intake transaction, not a mysterious prelude to transcription. Preserve the browser-generated boundary, validate one agreed file field at the first server hop, and record bytes, media type, outcome, and tenant-scoped usage without logging audio or transcript content. That turns a vague malformed-request response into a specific stage failure. It also lets a B2B SaaS team explain which tenant created the work before a moderation report reaches human review.

The before/after mental model is small. Before: browser -> framework -> speech endpoint -> generic 400. After: browser -> intake receipt -> validated audio -> transcription -> classification -> human review, with one correlation ID and one tenant ID joining the signals. Each arrow has an outcome. Each expensive stage has a usage record.

How did a multipart form-data speech API request become malformed?

Start at the first parser. A multipart body is coupled to the boundary parameter in its Content-Type header. If client code creates a FormData body, let the runtime create that header; manually writing multipart/form-data can omit the generated boundary that describes the body. The field name is a separate contract. If the receiver expects audio and the sender appends file, the transport can be valid while application validation still rejects it.

That distinction matters operationally. boundary_missing, file_missing, media_type_rejected, and upstream_rejected are different outcomes. Collapsing all four into bad_request makes the dashboard tidy and the diagnosis useless.

Use a narrow client contract:

export async function submitAudioReport(input: {
  endpoint: string;
  audio: Blob;
  tenantId: string;
}): Promise<Response> {
  const form = new FormData();
  form.append("audio", input.audio, "moderation-report.webm");

  return fetch(input.endpoint, {
    method: "POST",
    headers: { "x-tenant-id": input.tenantId },
    body: form,
  });
}
Enter fullscreen mode Exit fullscreen mode

The missing Content-Type assignment is intentional. Tenant identity, however, must come from verified authentication context; an arbitrary request header is not sufficient attribution.

For retries, the client should reuse an idempotency key for the same logical report and apply bounded backoff when a downstream service returns 429. Never retry a malformed body unchanged. Fixing transport validation comes first.

Build one copyable intake receipt

The intake layer should emit a receipt after parsing and validation. Keep it boring. The receipt is structured metadata, not the uploaded report.

type IntakeOutcome =
  | "accepted"
  | "boundary_missing"
  | "file_missing"
  | "media_type_rejected"
  | "size_rejected"
  | "upstream_rejected";

type IntakeReceipt = {
  requestId: string;
  tenantId: string;
  outcome: IntakeOutcome;
  bytesReceived: number;
  mediaType?: string;
  durationMs: number;
};

export function hasMultipartBoundary(value: string | null): boolean {
  if (!value) return false;
  return /^multipart\/form-data(?:;|$)/i.test(value) &&
    /(?:^|;)\s*boundary=(?:"[^"]+"|[^;\s]+)/i.test(value);
}

export function writeReceipt(receipt: IntakeReceipt): void {
  process.stdout.write(`${JSON.stringify({
    event: "audio_report_intake",
    ...receipt,
  })}\n`);
}
Enter fullscreen mode Exit fullscreen mode

This helper does not parse multipart data. A maintained parser should do that job and enforce limits while reading the stream. The helper answers a narrower question early: does the declared media type include the parameter needed to interpret this body? After parsing, validate that exactly one audio part exists, that it is a file rather than a text field, and that its declared and inspected characteristics meet policy.

Do not log the boundary value. Do not log filenames, raw audio, transcripts, authorization headers, or free-form parser errors either. Map failures to a small outcome vocabulary, retain the request ID for controlled investigation, and keep sensitive moderation material out of routine telemetry.

Make tenant cost visible without turning logs into invoices

Per-tenant cost visibility starts with attribution, but application logs should not pretend to be a billing ledger. Record stable usage units at the stage that observes them: received bytes at intake, audio duration after media inspection, and model usage returned by later processing when such usage exists. Join those records with tenantId, requestId, and an operation name. Apply prices or internal rates in a separate, versioned calculation.

The separation is useful because rates can change while historical usage does not. It also prevents a malformed upload from being charged as completed transcription merely because it reached the route.

type UsageEvent = {
  requestId: string;
  tenantId: string;
  operation: "audio_intake" | "transcription" | "classification";
  status: "accepted" | "completed" | "rejected";
  audioMilliseconds?: number;
  inputBytes?: number;
  occurredAt: string;
};

export function recordUsage(event: UsageEvent): void {
  process.stdout.write(`${JSON.stringify({ event: "usage_observed", ...event })}\n`);
}
Enter fullscreen mode Exit fullscreen mode

Keep cardinality under control. Tenant IDs belong in logs and usage records where per-tenant aggregation is required, but putting an unbounded tenant ID on every time-series metric can make the metrics system itself expensive. Use low-cardinality metric labels such as operation, outcome, and deployment region; calculate tenant views from structured events or a purpose-built usage store. This is a deliberate trade-off: metrics stay fast for alerting, while tenant attribution remains queryable.

Alert on service symptoms, not individual customers. A rise in boundary_missing after a frontend deployment suggests request construction changed. A rise in file_missing with stable boundary health points toward a field-contract mismatch. A rise in upstream_rejected after intake acceptance moves the investigation one stage later.

Crisp signals. Faster ownership.

What about server-side forwarding?

A common trap appears when an Express or Next.js handler parses an upload and then forwards it. The original multipart stream has already been consumed. Forwarding the original Content-Type header with a newly constructed body reuses a boundary that does not describe the new body. Build fresh FormData, append the validated file under the downstream field name, and let the runtime generate a fresh header.

export async function forwardForTranscription(input: {
  endpoint: string;
  token: string;
  audio: Uint8Array;
  mediaType: string;
  requestId: string;
}): Promise<Response> {
  const body = new FormData();
  body.append("audio", new Blob([input.audio], { type: input.mediaType }), "report-audio");

  return fetch(input.endpoint, {
    method: "POST",
    headers: {
      authorization: `Bearer ${input.token}`,
      "x-request-id": input.requestId,
    },
    body,
  });
}
Enter fullscreen mode Exit fullscreen mode

Why buffer Uint8Array at all? For large or frequent audio, do not. Stream through a parser and downstream client that support backpressure, enforce byte and time limits during ingestion, and avoid duplicating the full payload in memory. The buffered code keeps the contract visible; it is not a performance prescription.

Test the boundary between layers with fixtures, not with a live model call. Send one valid file, one request without a boundary parameter, one wrong field name, one text field named audio, one disallowed media type, and one body beyond the configured limit. Assert both the HTTP response and the emitted outcome. Then test the forwarding adapter separately with a local receiver that inspects parsed parts.

Ship the signals before the classifier

Deploy the intake vocabulary first. Watch rejection ratios and latency by operation and outcome. Then enable transcription and classification for a small tenant cohort while checking that each accepted request produces the expected stage sequence and that rejected uploads stop before downstream usage is recorded.

Human review stays the final decision point. The classifier output should carry the same request and tenant context into the review queue, but routine observability should contain identifiers and outcomes rather than report content. The team can locate malformed transport, application contract errors, and later processing failures without reading a customer's audio.

The result is more than a fixed upload. You get an auditable path from tenant submission to human review, plus usage records that support cost allocation without coupling diagnostics to a vendor or a mutable price sheet.

Further reading

Top comments (0)