DEV Community

Jonathan
Jonathan

Posted on

A Failure Bundle for GitHub Actions API Tests

API tests usually fail with a red check, a stack trace, and a request to “run it again.” That is not much evidence. By the time someone opens the workflow, the relevant response body may be gone, the test fixture may have expired, and nobody remembers which commit produced the request.

When I build API integrations, I try to make every failed GitHub Actions job leave behind a small failure bundle. It is a boring improvement, but it turns a vague CI alert into something another developer can reproduce without a live debugging session.

Why a red CI job is not enough

A failed API test has at least three useful dimensions:

  • The request: method, route, safe headers, and a correlation ID.
  • The response: status, selected headers, and a redacted body.
  • The environment: commit, test command, dependency lockfile, and fixture state.

Most pipelines preserve only the test runner output. That output might say that an assertion expected 201 and received 409, but it will not explain whether the duplicate came from a reused fake email address, a retry, or a backend job that finished late.

The fix is not “log everything.” Logging everything creates privacy risks and makes the real signal hard to find. The fix is a predictable artifact with a narrow schema.

What belongs in a failure bundle

I use a directory with one small JSON file per failed case and a workflow summary that links to the artifact. A useful record looks like this:

{
  "test": "creates a verification session",
  "commit": "${GITHUB_SHA}",
  "request_id": "req-8d2c",
  "method": "POST",
  "path": "/v1/verification-sessions",
  "status": 409,
  "fixture": "fresh-address",
  "retry": 1,
  "body_redacted": "verification session already exists"
}
Enter fullscreen mode Exit fullscreen mode

Do not store authorization headers, full cookies, or an entire inbound email in this file. A fixture label such as fresh-address is enough to tell a teammate what was intended. If a test needs an inbox-like value, generate it inside the test run and record only a hash or short label.

For teams that search for a tempail or a tem email while investigating, consistent field names matter more than the spelling in a query. Keep the canonical terms in documentation, then make the artifact easy to grep. This is a small detail, but it saves a realy annoying amount of time during an incident.

A GitHub Actions implementation

The test command should always write its diagnostics to a known directory, even when the process exits non-zero. Then upload the directory in an always() step:

name: API tests

on:
  pull_request:
  push:

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npm test -- --reporter=dot
        env:
          API_DIAGNOSTICS_DIR: artifacts/api-failures
      - if: always()
        uses: actions/upload-artifact@v4
        with:
          name: api-failure-bundle-${{ github.run_id }}
          path: artifacts/api-failures
          if-no-files-found: ignore
          retention-days: 7
Enter fullscreen mode Exit fullscreen mode

The important part is the contract between the test runner and the workflow: API_DIAGNOSTICS_DIR is stable, and the upload step runs whether the tests pass or fail. On a green run, the directory can be empty. On a red run, the bundle is attached to the exact workflow execution.

For quick investigation, append a short summary as well. A summary can say which test failed, link the artifact, and show the request ID. Keep it short; large logs belong in the artifact, not in the pull request timeline.

Make retries produce comparable evidence

Retries can hide a race condition. Every attempt should therefore have its own record, with an attempt number and timestamp. Do not overwrite failure.json on retry. Use names such as creates-session-attempt-1.json and creates-session-attempt-2.json.

This makes it possible to compare a 409 followed by a 201, or a timeout followed by a successful response. It also tells you when the test is not flaky at all: a fixture may simply be reused between attempts. The cleanup rule should be visible in the bundle, not implied in a test helper that only one person remembers.

I also keep the request ID from the API response, when one exists. That ID connects the CI artifact to server logs without copying sensitive payloads into GitHub. A redacted trace is usualy more useful than a perfect transcript that cannot be shared.

If your project already has notes on email API triage, this failure-bundle pattern fits naturally beside them. The same boundary applies when an email event is part of an API test: record the event ID and state, not private message contents.

Keep test data private and useful

Fake email addresses and temporary fixtures are helpful for isolating signup and verification flows, but they are not a privacy policy. Avoid real customer addresses, production tokens, and complete message bodies in CI artifacts. Set a short retention period, restrict who can read artifacts, and delete local diagnostics before packaging anything else.

An email deployment contract is a useful mental model here: define what the test promises to observe, then store only the evidence needed to check that promise. This keeps the bundle small and makes later cleanup much simpler.

A small adoption checklist

Start with one unreliable API test and add:

  1. A stable diagnostics directory.
  2. A JSON record for request, response, fixture, and attempt metadata.
  3. Redaction before the record is written.
  4. An always() artifact upload in GitHub Actions.
  5. A short workflow summary with the test name and request ID.
  6. A retention rule that matches the sensitivity of the data.

The result is not a fancy observability system. It is a repeatable handoff from a failing check to the next developer. That little bit of evidence often makes API failures faster to fix, easier to discuss, and less likely to be dismissed as “just CI being flaky.”

Top comments (0)