"What should the type be when the value is missing?" is the question that most quietly corrupts an API contract. Teams answer it with one word, nullable, when there are actually three independent questions hiding inside: must the key be present, may the value be JSON null, and what does absence mean? Conflate them and generated clients crash on a null the spec said could not happen, or branch endlessly on keys that are never actually omitted.
Three orthogonal facts
| Fact | Expressed by | Question it answers |
|---|---|---|
| Presence | the required array |
Must the key exist in the JSON? |
| Nullability | including 'null' in type (3.1) |
May the value be JSON null? |
| Default | default |
What does the server assume when absent? |
These combine independently. A field can be required and nullable, optional but non-null when present, or optional with a server default. Do not use one keyword to answer two questions.
The four request-body combinations
For a field such as middleName, each combination is a distinct contract.
| required? | nullable? | Accepted JSON | Meaning |
|---|---|---|---|
| yes | no | "middleName": "Lee" |
Must send a real string |
| yes | yes | "middleName": null |
Must send the key; null explicitly means "no middle name" |
| no | no | key omitted, or a string | You may leave it out, but never send null |
| no | yes | omitted, a string, or null | Full three-state field |
In OpenAPI 3.1 the nullable combination is written with a type array:
middleName:
type: [string, 'null']
description: Explicit null means no middle name; key may be omitted to keep the current value.
The OpenAPI 3.0 keyword nullable: true is removed in 3.1. If a field is both optional and nullable, list 'null' in type and leave it out of required; do not also add the old keyword.
Responses: absent, null, and empty are not synonyms
Response modeling has its own trap. Four different payloads carry different meaning, and the spec should say which one the server actually returns.
{ "discount": null }
{}
{ "tags": [] }
{ "tags": null }
- Key present with
null: the concept exists and is explicitly empty for this resource. - Key omitted: the concept does not apply, or the server uses sparse fieldsets.
- Empty array
[]: the collection exists and currently has zero items. - Array with
null: usually a modeling mistake; prefer[]for "no items" and reserve null for "the collection itself is unknown or not loaded."
Clients iterate [] safely; they must guard null. If your success response always initializes collections, model the array as non-nullable and return []. If an error or partial-response branch can leave it unknown, model [array, 'null'] on that response, not on the happy path.
Do not tighten the type to make a diff pass
There is a recurring temptation in tooling: a comparison shows a nullable array in one document and a non-null one in another, so someone flips a single flag to make them match. That is how contracts lie. A Go handler that returns a nil slice on one branch and an initialized slice on another serializes them differently (null versus []) depending on the encoder and the path. The correct type is decided by the real success and error branches, not by which edit clears a warning. Forcing every array to non-null, or every one to nullable, hides genuine behavior.
| Evidence | Honest schema |
|---|---|
| Every success path initializes the collection |
type: array, non-nullable |
| One reachable success path returns null | type: [array, 'null'] |
| Only an error response omits it | Keep it required on the 200, absent on the error schema |
| Cannot be proven from source | Mark the field unresolved; do not guess |
One non-null sample proves nothing
A single happy-path request that returns a populated object is not evidence that a field is never null. Nullability is decided by branches: a not-found join, an optional relation, a failed enrichment, a soft-deleted record, or a feature flag can all produce null on a different input. When you reverse-engineer a spec from code, trace the handler's success and error branches and its cross-package helpers. If a computed mapper can return null and you cannot prove otherwise, the field is nullable or it is an explicit gap; a confident non-null claim backed by one sample is the most common and most dangerous false positive a scanner can emit.
The same discipline works the other way: a field absent from one response because the client requested sparse fields is not optional on the full representation. Match the schema to the specific response and status code it describes.
What codegen does with each choice
A TypeScript generator turns the distinctions into the type system:
// required, non-null
middleName: string;
// required, nullable
middleName: string | null;
// optional, non-null when present
middleName?: string;
// optional and nullable, the full three-state field
middleName?: string | null;
Collapse them and the generated client either forces callers to supply values that should be optional, or fails to guard a null that the server really sends. Server-side generators and validators use the same flags to decide whether a missing key is a 422 or a legitimate default.
Mocks and AI should generate all three states
A mock that only returns the populated happy path never exposes the null-handling bugs. Drive the mock from the schema and deliberately generate four variants: the full object, the object with the nullable field set to null, the object with the optional key omitted, and the empty collection. Then confirm the client renders each one instead of throwing. AI agents generating test data behave the same way: a schema that distinguishes the states gives the model permission to emit them; a vague string type yields only the populated case, so null handling ships untested.
A decision table for the next field you model
- Can the key be absent? No: add it to
required. Yes: leave it out and document why. - Is explicit
nulla meaningful value? Yes: add'null'totype. No: keep the type single. - Does absence imply a value the server fills? Add
defaultand say whether the response echoes it. - Is it a collection? Return
[]for "empty and known," null only for "unknown." - Which status codes return which shape? Model each response separately; never let an error branch widen the success schema.
- What is the evidence? A traced branch or a request on the specific path; otherwise mark unresolved.
Checklist
-
required, nullability, anddefaultare set independently and each is intentional. - 3.1 documents use
type: [T, 'null']; no leftover 3.0nullable. - Empty collections are
[]; null arrays mean "unknown," and only when a branch proves it. - Different status codes have different response schemas where the shapes actually differ.
- Nullability claims are backed by branch analysis or requests, never by one populated sample.
- Mocks and AI-generated tests cover full, null, omitted, and empty variants.
A contract that is honest about missing data is worth more than a complete-looking one that crashes on the first null. When you cannot prove a field, label it unresolved; that is a feature, not a gap in the report.
Model the three states, generate the precise client types, and have a mock return null, omitted, and empty variants, all in the browser app. Keeping unknown fields honestly marked instead of guessing is the core principle behind a reliable code scan; the split between deterministic analysis and reviewable AI gap-filling is explained in the AST vs AI for reverse-engineering OpenAPI guide.
Top comments (0)