Here is the most common API test in the world. Send a request, assert status 200, call it green. It proves a route exists and the server did not crash. It does not prove that a user can complete a single task in your product, and it will happily stay green while the integration is completely broken.
Day 4 of the the series is about scenario tests: chained requests that walk a business flow end to end, carry data between steps, and assert against the contract from Day 2.
What a single request cannot see
A user never does one request. They authenticate, create something, read it back, act on it, and expect the terminal state to match. Every interesting bug lives between the requests:
- Create returns 201; the immediate read returns 404 because of a replication delay nobody modeled, and the client has no retry.
- Cancel accepts an illegal state transition and returns 200, leaving the record half-cancelled.
- The list endpoint silently caps
limitat 50 while the UI paginates assuming 100, so records vanish without an error. - The token refresh works in isolation but expires halfway through the flow, and step six sends a stale bearer.
- Money arrives as a string in one endpoint and a number in the next.
None of these are caught by per-route green checks. All of them are caught by a test that performs the actual sequence.
What a scenario is made of
A scenario is an ordered list of steps. Each step is one documented request plus optional pre-request work and assertions; steps can extract values from responses and hand them to later steps:
| Step | Request | Extraction / assertion |
|---|---|---|
| 1 | POST /auth/token |
Extract access_token into the environment |
| 2 |
POST /holds with Idempotency-Key
|
Assert 201, extract id
|
| 3 | Replay step 2 with the same key | Assert the same id comes back |
| 4 | GET /holds/{id} |
Assert 200, schema matches Hold, status: held
|
| 5 | POST /holds/{id}/confirm |
Assert 200, status: confirmed
|
| 6 | Subscribe to the SSE stream | Assert the confirmation event was emitted |
| 7 |
POST /holds/{id}/confirm again |
Assert 409 with the documented error envelope |
Three mechanisms do the heavy lifting:
- Extraction. JSONPath or property references pull ids and tokens out of responses into variables. No hard-coded ids means the test is repeatable against a fresh environment.
- Pre-request scripts. Signing, HMAC, timestamped tokens, nonces — anything the gateway expects gets computed per step instead of pasted.
- Schema assertions. Beyond status codes, assert the response validates against the schema in the spec. That is the part that catches the string-versus-number money bug and the missing field the mobile team is already rendering.
Run the same chain everywhere
Environments are just named variable sets: localhost, the mock from Day 3, staging, production (read-only scenarios only). The scenario file never changes; the environment pointer does. That gives you a progression the whole team can reason about:
- Against the mock, the scenario is the executable definition of done before the backend exists.
- Against a developer's localhost, it is the first integration check after a route is implemented.
- Against staging on every pull request, it is contract verification in CI, with an exported report attached to the ticket.
- Against production with read-only steps, it is a synthetic uptime check that walks a real customer path.
Most workspaces that support this also generate a run-host snippet — a one-command runner you can drop into CI so the exact chain a developer ran locally is the chain the pipeline runs.
Let AI write the first draft, then make it prove itself
This is one of the places AI assistance is genuinely worth it, with a caveat. Given the spec and a plain-language scenario ("reserve, idempotently replay, confirm, observe the event, reject double-confirm"), the assistant can draft the step list, extraction paths, and assertions quickly. The caveat is the same one from Day 2: the output is a proposal. Review three things specifically:
- Does every assertion come from the contract, or did the model invent fields that "look reasonable"?
- Does the scenario include the failure cases, or only the happy path it prefers?
- Are extracted values actually used downstream, or is the chain secretly independent steps in a trench coat?
Then the loop closes in the other direction. When an implementation fails a scenario, do not describe the failure to the coding agent in prose. Hand it the evidence: the request that was sent, the expected schema, the actual response. Agents fix evidence far more reliably than they fix vibes. A failed scenario is a bug report with a reproduction attached.
A practical starter set
If you do nothing else, write four scenarios per service:
- Happy path — the primary business flow, end to end, with extraction.
- Idempotency and retries — replay mutating requests, assert no duplicates.
- Permission and validation failures — bad token, missing required field, illegal state transition; assert the shared error envelope every time.
- Pagination and consistency — create N resources, page through them, read one back, assert nothing silently disappears.
Those four cover the majority of the "it passed in Postman but broke in production" tickets I have ever seen.
The tooling is commodity — any runner that supports chained requests, environments, and schema assertions works; I run scenarios inside Powerduck because they derive directly from the same local spec as the mock and the MCP server, and the reports and CI runner are built in. The argument for treating the business flow, rather than the route, as the unit of testing is the part that matters, and I made it at greater length here.
Tomorrow the direction reverses. Days 2 through 4 assumed no code existed. Day 5 covers the messier, more common reality: a couple hundred routes already running in production, no spec anywhere, and a deadline.
Top comments (0)