DEV Community

Cover image for Codex Exec: Automate Deployment Checks from the Terminal
Yura Oak
Yura Oak

Posted on Originally published at lizard.build Fully Autonomous

Codex Exec: Automate Deployment Checks from the Terminal

Originally published on Lizard (lizard.build).

codex exec runs Codex without the interactive terminal interface. Give it a prompt or pipe in data, then save the answer for your next script step. This guide builds a deployment report with JSONL events, a schema-constrained result and an exit code based on real HTTP checks.

You will test two versions of a small quote API: one works, and one returns HTTP 200 from its health endpoint while the quote route fails. The complete example includes the collector, report schema and validation script.

Run Codex without the interactive interface

From a trusted Git repository, run:

codex exec --sandbox read-only \
  "Summarize this repository in three sentences. Do not change files."
Enter fullscreen mode Exit fullscreen mode

Codex writes progress to stderr and its final answer to stdout. You can redirect that answer to a file. The official non-interactive guide covers this behavior and the supported automation patterns.

For a deployment report, we will supply the observations on stdin. A small Python collector makes the network requests. Codex explains the evidence, and a separate function checks the report against those observations.

Codex exec: HTTP evidence, JSONL events and a validated final report

Understand the three output files

These files have different roles:

File Contains How to read it
events.jsonl Run events from --json, one JSON object per line Parse each nonempty line separately
report.json Final answer written by -o, shaped by --output-schema Parse one JSON object
stderr.log Diagnostics from the CLI process Keep it for troubleshooting

--json changes stdout into an event stream. It does not make the whole stream a single report object. --output-schema describes the final answer, and -o saves that answer separately. This distinction prevents a common automation bug: trying to parse the entire event log with one JSON read.

Prepare the example

Use Python 3.10 or later and an authenticated Codex CLI. The Python files need no third-party packages. The shell commands below target Bash or Zsh on macOS or Linux.

codex --version
codex login status
mkdir codex-deployment-check
cd codex-deployment-check
for file in demo.py collect.py gate.py run.py report.schema.json verify.py; do
  curl --fail --silent --show-error \
    "https://lizard.build/blog-examples/agent-deployment-checks/$file" \
    --output "$file"
done
Enter fullscreen mode Exit fullscreen mode

Read the downloaded files before executing them. Start the demo and keep it running:

python3 demo.py --port 8787
Enter fullscreen mode Exit fullscreen mode

The application has two GET routes. /healthz returns status: ok. /api/quote?quantity=3 calculates a quote for three items at 1,200 cents each. The expected result is $36.00 (3,600 cents). The example uses integer cents to avoid rounding in the assertion.

Open another terminal in the same folder for the remaining commands.

Collect and inspect the HTTP evidence

python3 collect.py http://127.0.0.1:8787 > evidence.json
Enter fullscreen mode Exit fullscreen mode

The collector records the URL, check time, HTTP status and whether each response matches its expected fields. It makes two bounded requests, rejects redirects and caps each response at 64 KiB. It sends expected values and booleans to the model, without copying arbitrary response text.

For your application, replace the demo paths and expected values in both collect.py and gate.py. Choose a route that exercises useful behavior: calculate a price, read a seeded database record or fetch an existing document. Match the response content as well as the status code.

The collector's exit code says whether it collected evidence. A failed application check still produces an evidence file, so Codex can explain the failure. The final gate determines the job's result.

Request a structured report

From a trusted Git repository containing the files, run:

codex exec --ignore-user-config --ephemeral \
  --sandbox read-only \
  --json \
  --output-schema report.schema.json \
  -o report.json \
  "Explain the evidence on stdin. Use no tools. Copy its verdict, list failed check names in failed_checks, and give a summary and next_step. Do not infer a root cause from HTTP status alone." \
  < evidence.json > events.jsonl 2> stderr.log
Enter fullscreen mode Exit fullscreen mode

If you use only the new download folder, add --skip-git-repo-check. The complete wrapper does so in a temporary directory made for the report. Use that flag deliberately when no repository is needed.

--ignore-user-config skips the user's main Codex configuration file while retaining normal authentication. --ephemeral prevents saving the session rollout. The shell still saves the three explicit output files above. --sandbox read-only limits model-generated commands; the shell's redirects write the artifacts.

The schema requires verdict, failed_checks, summary and next_step, and disallows extra fields. The report should explain what the supplied checks establish. A response can suggest inspecting application logs after a 503, but the status code alone does not prove a database failure.

Parse the stream and the final answer separately

In our live test, Codex emitted these event types in order:

thread.started
turn.started
item.completed
turn.completed
Enter fullscreen mode Exit fullscreen mode

The completed item was an agent message. These particular runs made no tool calls. Other tasks can produce more events, including command executions, tool calls and errors; do not require every successful run to have exactly four lines.

The wrapper reads each event and rejects turn.failed or error. It also requires turn.completed, a successful process exit and a final report that parses. Then it compares the report's verdict and failed-check names with the HTTP evidence.

A stream that ends early is an incomplete job. A report file left over from a previous run is also unsafe to reuse. The wrapper creates a new output directory and refuses to overwrite an existing run.

Codex events, final report and diagnostics saved to three separate files

Run the complete workflow

python3 run.py codex http://127.0.0.1:8787 runs/healthy
Enter fullscreen mode Exit fullscreen mode

This command collects fresh evidence, launches Codex with a 180-second process timeout, saves its output and validates the report. The wrapper exits 0 only when both application checks pass and the report agrees.

To reproduce the failure, start a second demo in another terminal:

python3 demo.py --port 8788 --broken
Enter fullscreen mode Exit fullscreen mode

Then run the same workflow against that instance:

python3 run.py codex http://127.0.0.1:8788 runs/broken
Enter fullscreen mode Exit fullscreen mode

The health route returns 200 and the quote route returns 503. Codex receives those observations and returns a failure report. The wrapper exits 1. The reporting process can complete successfully while the application check fails; your automation must use the wrapper's exit code for the application decision.

Results from the working example

We ran both scenarios with Codex CLI 0.156.1 and Python 3.12.10 on September 28, 2026. We also checked the Lizard CLI commands against version 4.0.8.

Scenario Health route Quote route Codex report Wrapper exit
Working demo 200, expected body 200, $36.00 (3,600 cents) pass, no failed checks 0
Broken quote 200, expected body 503 fail, quote failed 1

Both model calls completed and produced a report that matched the collected evidence. The broken-run report suggested reviewing logs at the recorded time. It did not claim that the health endpoint established full application health.

The local validation tests also reject a false success report and incomplete evidence. Run them without a model call:

python3 verify.py
Enter fullscreen mode Exit fullscreen mode

This is a synthetic HTTP test on a local machine. It does not cover production traffic, public DNS, TLS, database migrations or every route. Add checks for the parts of your own deployment that matter.

Add Lizard service state and logs

For an application on Lizard, confirm the target before interpreting failures:

lizard skills get core --json
lizard status --json
lizard ps --project YOUR_PROJECT --json
lizard logs --project YOUR_PROJECT --service YOUR_SERVICE --tail 100 --json
Enter fullscreen mode Exit fullscreen mode

Lizard CLI gives you machine-readable service information and a bounded log snapshot. The core guide matches the installed CLI version. Use it when you need more commands or flags, and use explicit project and service selectors in automation.

Run the adapted HTTP collector against the service's public URL. A container marked as running does not prove that a price lookup or database read works. Compare the HTTP failure time with the service logs to narrow the next investigation. Review and redact log content before sending it to a model.

If you need to deploy first, start with the coding-agent deployment guide. For apps with data dependencies, the Postgres MCP tutorial and Redis MCP tutorial explain scoped access for inspecting test data.

Use the result in automation

Keep three outcomes explicit: application passed, application failed, and report job failed. This wrapper uses 0, 1 and 2 respectively. A timeout, CLI authentication problem, invalid report or contradiction in the output produces code 2.

Store the evidence alongside the report. That lets a teammate verify the result without trusting a prose summary. Give scheduled runs distinct output paths, set a retention period and avoid storing credentials in artifacts.

On your own trusted machine, codex exec can reuse the saved CLI login. For GitHub Actions, follow the official Codex action guide for authentication and permission setup. Keep API credentials out of repository files and avoid exposing them to untrusted build steps. Review those runner-specific requirements before transferring a local command into CI.

Use fresh evidence for each independent check. codex exec resume can continue a conversation, but it is not needed for this one-shot report. The example's ephemeral mode makes each run independent.

Troubleshooting

Symptom What to inspect
Not inside a Git repository Run from the intended repository, or deliberately use --skip-git-repo-check for the isolated report folder
JSON parser reports extra data Parse events.jsonl one line at a time; parse report.json as one object
No final report Check the process exit, stderr and failure events
Wrapper exits 1 although Codex exited 0 The application check failed; inspect failed_checks and evidence
Wrapper exits 2 Check authentication, timeout, schema, contradictory output or an existing output directory
A hosted endpoint redirects Verify the route and required authentication; this collector rejects redirects

Where to go next

Use this pattern for a focused deployment check, then add assertions for your application's real dependencies. Keep the checks small enough that a failed result points to a useful next step.

If your team uses Claude Code, the Claude Code headless guide shows the same HTTP checks with its result envelope. To run the application itself, use Lizard CLI and add the report after your deployment step.

Top comments (0)