Most OpenAPI tutorials stop at application/json, yet a large share of real endpoints hand back a file: an invoice PDF, a CSV export, a ZIP of assets, a resized image, a generated spreadsheet. Teams either leave these endpoints undocumented or describe them as type: string with no format, so a generated client tries to decode a PDF as UTF-8 text and corrupts it. Non-JSON responses need explicit media types, a binary schema, and the headers that tell the client what to do with the bytes.
The schema for a binary body
A file response uses type: string with format: binary (OpenAPI 3.x), not an object and not an untyped string. The media type says what the bytes are:
paths:
/invoices/{invoiceId}/pdf:
get:
summary: Download an invoice as PDF
operationId: downloadInvoicePdf
parameters:
- name: invoiceId
in: path
required: true
schema: { type: string }
responses:
'200':
description: The invoice PDF.
headers:
Content-Disposition:
schema: { type: string }
description: 'attachment; filename="invoice-1042.pdf"'
content:
application/pdf:
schema:
type: string
format: binary
'404':
$ref: '#/components/responses/NotFound'
format: binary tells generators to treat the body as raw bytes (a Blob, ReadableStream, byte[], or Buffer depending on language) rather than parsing it. Use format: byte only for actual base64-encoded content embedded in JSON, which is rare for downloads; a normal file response is binary.
Common download media types
| Artifact | Media type | Notes |
|---|---|---|
application/pdf |
Invoices, reports, contracts | |
| CSV | text/csv |
Tabular exports; declare charset if needed |
| Excel | application/vnd.openxmlformats-officedocument.spreadsheetml.sheet |
.xlsx |
| ZIP | application/zip |
Bundled assets, bulk export |
| PNG/JPEG/WebP |
image/png, image/jpeg, image/webp
|
Generated or resized images |
| Plain text | text/plain |
Logs, .env templates |
| Unknown / any | application/octet-stream |
Fallback; prefer a specific type |
For CSV, state the dialect in the description when it matters (comma vs semicolon, header row, UTF-8 with BOM for Excel, date formatting). Clients parsing the file cannot infer these from the media type alone.
inline versus attachment
Content-Disposition decides whether the browser renders the file or downloads it:
-
inline; filename="invoice.pdf"lets the browser display it in a tab (common for PDF and images). -
attachment; filename="invoice-1042.pdf"forces a Save dialog and suggests the file name.
Document the header as part of the response and, where the client can choose, accept a query parameter:
parameters:
- name: download
in: query
schema: { type: boolean, default: false }
description: When true, sets Content-Disposition to attachment to force a download.
Always provide a stable filename, ASCII-safe or RFC 5987 encoded for non-ASCII names, so saved files are not named after the endpoint.
Multiple formats from one endpoint
An export that can be CSV, XLSX, or JSON is a content-negotiation decision. Use either an Accept header or an explicit format parameter and list each success representation:
responses:
'200':
description: The export in the requested format.
content:
text/csv:
schema: { type: string, format: binary }
application/vnd.openxmlformats-officedocument.spreadsheetml.sheet:
schema: { type: string, format: binary }
application/json:
schema:
type: array
items: { $ref: '#/components/schemas/Order' }
'406':
description: Requested format is not supported.
Listing each media type lets a generator expose per-format methods and lets documentation render the right example. Do not collapse them into application/octet-stream, which hides the real contract.
Errors are JSON, not the file
A download endpoint can still fail with a structured body. Give error responses their own JSON media type so clients do not attempt to parse an HTML proxy error page or a JSON error as the file:
'401':
description: Authentication required.
content:
application/json:
schema: { $ref: '#/components/schemas/ProblemDetail' }
'403':
description: The file exists but this account may not access it.
content:
application/json:
schema: { $ref: '#/components/schemas/ProblemDetail' }
Clients must check the response status and Content-Type before treating bytes as the expected file; a 200 with application/json after an auth redirect is a common source of a saved "PDF" that is actually a login page.
Resumable and ranged downloads
Large files should support range requests and conditional GET. Document the headers so download managers and clients can resume and cache:
'200':
headers:
Accept-Ranges:
schema: { type: string, example: bytes }
ETag:
schema: { type: string }
Last-Modified:
schema: { type: string, format: http-date }
content:
application/zip:
schema: { type: string, format: binary }
'206':
description: Partial content for a Range request.
headers:
Content-Range:
schema: { type: string, example: bytes 0-1048575/8388608 }
content:
application/zip:
schema: { type: string, format: binary }
'304':
description: Not modified; use the cached copy.
If-None-Match/ETag and If-Modified-Since avoid re-downloading unchanged files; Range/206 Partial Content enables resume and parallel fetch. Signed URLs that expire should return a documented 403 or 410 when stale, with a documented way to fetch a fresh link.
Async-generated files
A file that takes time to generate should not block. Kick off a job that returns 202 and a job resource; when the job succeeds, its result is a short-lived signed download URL modeled exactly as above. This composes the async-job pattern with the download pattern instead of holding the request open for minutes.
What codegen, mocks, and AI callers need
-
format: binaryis the signal that makes a generator return aBlob/byte[]; without it, JavaScript clients corrupt binary by decoding as text. - Multiple
contententries let typed clients request a specific representation and validate406. - A spec-driven mock should return a small valid fixture (a real one-page PDF or a few CSV lines) with the correct headers, plus a JSON
404, so the client's save-to-disk and error paths are both testable. - An AI agent wiring up a download needs the media type, the
Content-Dispositionbehavior, and the fact that errors are JSON; without that it will guess the extension and mishandle failures.
Checklist
- Model file bodies as
type: string, format: binarywith the precise media type; never as an object or plain string. - Document
Content-Disposition, including a safe filename and inline versus attachment. - List each supported format as a separate
contententry and add a406. - Give errors a JSON Problem Detail body and tell clients to verify status and Content-Type before saving.
- Add
Accept-Ranges,ETag,Last-Modified,206, and304for large or cacheable files. - Document signed-URL expiry and how to obtain a fresh link.
- Generate the client and confirm downloads arrive as bytes, and mock both a valid file and a JSON error.
Get these right and invoices export without corruption, large bundles resume cleanly, and a failed download shows a real error instead of a PDF full of HTML.
You can describe binary responses, generate byte-accurate clients, and mock a real file plus a JSON error in one local-first workspace, right in your browser. For the upload side of the same contract, see file uploads and multipart/form-data in OpenAPI.
Top comments (0)