DEV Community

Jeff
Jeff

Posted on Originally published at powerduck.com

Optional, nullable, or omitted? Modeling missing data in OpenAPI 3.1 without lying to clients

"What should the type be when the value is missing?" is the question that most quietly corrupts an API contract. Teams answer it with one word, nullable, when there are actually three independent questions hiding inside: must the key be present, may the value be JSON null, and what does absence mean? Conflate them and generated clients crash on a null the spec said could not happen, or branch endlessly on keys that are never actually omitted.

Three orthogonal facts

Fact Expressed by Question it answers
Presence the required array Must the key exist in the JSON?
Nullability including 'null' in type (3.1) May the value be JSON null?
Default default What does the server assume when absent?

These combine independently. A field can be required and nullable, optional but non-null when present, or optional with a server default. Do not use one keyword to answer two questions.

The four request-body combinations

For a field such as middleName, each combination is a distinct contract.

required? nullable? Accepted JSON Meaning
yes no "middleName": "Lee" Must send a real string
yes yes "middleName": null Must send the key; null explicitly means "no middle name"
no no key omitted, or a string You may leave it out, but never send null
no yes omitted, a string, or null Full three-state field

In OpenAPI 3.1 the nullable combination is written with a type array:

middleName:
  type: [string, 'null']
  description: Explicit null means no middle name; key may be omitted to keep the current value.
Enter fullscreen mode Exit fullscreen mode

The OpenAPI 3.0 keyword nullable: true is removed in 3.1. If a field is both optional and nullable, list 'null' in type and leave it out of required; do not also add the old keyword.

Responses: absent, null, and empty are not synonyms

Response modeling has its own trap. Four different payloads carry different meaning, and the spec should say which one the server actually returns.

{ "discount": null }
Enter fullscreen mode Exit fullscreen mode
{}
Enter fullscreen mode Exit fullscreen mode
{ "tags": [] }
Enter fullscreen mode Exit fullscreen mode
{ "tags": null }
Enter fullscreen mode Exit fullscreen mode
  • Key present with null: the concept exists and is explicitly empty for this resource.
  • Key omitted: the concept does not apply, or the server uses sparse fieldsets.
  • Empty array []: the collection exists and currently has zero items.
  • Array with null: usually a modeling mistake; prefer [] for "no items" and reserve null for "the collection itself is unknown or not loaded."

Clients iterate [] safely; they must guard null. If your success response always initializes collections, model the array as non-nullable and return []. If an error or partial-response branch can leave it unknown, model [array, 'null'] on that response, not on the happy path.

Do not tighten the type to make a diff pass

There is a recurring temptation in tooling: a comparison shows a nullable array in one document and a non-null one in another, so someone flips a single flag to make them match. That is how contracts lie. A Go handler that returns a nil slice on one branch and an initialized slice on another serializes them differently (null versus []) depending on the encoder and the path. The correct type is decided by the real success and error branches, not by which edit clears a warning. Forcing every array to non-null, or every one to nullable, hides genuine behavior.

Evidence Honest schema
Every success path initializes the collection type: array, non-nullable
One reachable success path returns null type: [array, 'null']
Only an error response omits it Keep it required on the 200, absent on the error schema
Cannot be proven from source Mark the field unresolved; do not guess

One non-null sample proves nothing

A single happy-path request that returns a populated object is not evidence that a field is never null. Nullability is decided by branches: a not-found join, an optional relation, a failed enrichment, a soft-deleted record, or a feature flag can all produce null on a different input. When you reverse-engineer a spec from code, trace the handler's success and error branches and its cross-package helpers. If a computed mapper can return null and you cannot prove otherwise, the field is nullable or it is an explicit gap; a confident non-null claim backed by one sample is the most common and most dangerous false positive a scanner can emit.

The same discipline works the other way: a field absent from one response because the client requested sparse fields is not optional on the full representation. Match the schema to the specific response and status code it describes.

What codegen does with each choice

A TypeScript generator turns the distinctions into the type system:

// required, non-null
middleName: string;
// required, nullable
middleName: string | null;
// optional, non-null when present
middleName?: string;
// optional and nullable, the full three-state field
middleName?: string | null;
Enter fullscreen mode Exit fullscreen mode

Collapse them and the generated client either forces callers to supply values that should be optional, or fails to guard a null that the server really sends. Server-side generators and validators use the same flags to decide whether a missing key is a 422 or a legitimate default.

Mocks and AI should generate all three states

A mock that only returns the populated happy path never exposes the null-handling bugs. Drive the mock from the schema and deliberately generate four variants: the full object, the object with the nullable field set to null, the object with the optional key omitted, and the empty collection. Then confirm the client renders each one instead of throwing. AI agents generating test data behave the same way: a schema that distinguishes the states gives the model permission to emit them; a vague string type yields only the populated case, so null handling ships untested.

A decision table for the next field you model

  1. Can the key be absent? No: add it to required. Yes: leave it out and document why.
  2. Is explicit null a meaningful value? Yes: add 'null' to type. No: keep the type single.
  3. Does absence imply a value the server fills? Add default and say whether the response echoes it.
  4. Is it a collection? Return [] for "empty and known," null only for "unknown."
  5. Which status codes return which shape? Model each response separately; never let an error branch widen the success schema.
  6. What is the evidence? A traced branch or a request on the specific path; otherwise mark unresolved.

Checklist

  • required, nullability, and default are set independently and each is intentional.
  • 3.1 documents use type: [T, 'null']; no leftover 3.0 nullable.
  • Empty collections are []; null arrays mean "unknown," and only when a branch proves it.
  • Different status codes have different response schemas where the shapes actually differ.
  • Nullability claims are backed by branch analysis or requests, never by one populated sample.
  • Mocks and AI-generated tests cover full, null, omitted, and empty variants.

A contract that is honest about missing data is worth more than a complete-looking one that crashes on the first null. When you cannot prove a field, label it unresolved; that is a feature, not a gap in the report.

Model the three states, generate the precise client types, and have a mock return null, omitted, and empty variants, all in the browser app. Keeping unknown fields honestly marked instead of guessing is the core principle behind a reliable code scan; the split between deterministic analysis and reviewable AI gap-filling is explained in the AST vs AI for reverse-engineering OpenAPI guide.

Top comments (0)