I work as a software engineer at a startup, where small changes go out to production all the time. One of them broke a feature in prod. The code looked fine. We went through the commit history line by line looking for what we had missed, while 500s kept piling up.
The cause was an env var. The change read a new variable that existed on the developer's machine but had never been added to .env.example or to production. One missing line, and the feature went down.
Why nothing caught it
This bug slips past every usual safety net:
-
Tests pass, because they mock the config or run with a local
.env. -
Review misses it, because a new
process.env.SOMETHINGis one short line in a long diff, and nobody checks it against.env.example. -
The build passes, because a missing env var isn't a compile error. It's just
undefinedat runtime. - Production finds it, after the deploy, through users or error logs.
The fix takes ten seconds. Finding it takes hours.
What you can do today, without any tool
-
Validate env vars at startup.
t3-env,envalidor a smallzodschema make the app fail on boot with a clear message, instead of on the first request with a mystery 500. -
Treat
.env.exampleas a contract. Every variable the code reads belongs there. Put it on your PR checklist. -
Check the diff before merging. Search it for
process.env.oros.environand compare each name against.env.example.
All three work. The third one only works if someone remembers to do it, and that's the part I wanted to automate.
What I built: deployhealth
It started as envcheck, a small scanner that finds every env var your code reads and compares them with your env files. Pull request checks and uptime monitoring grew around it, and it became deployhealth.
1. A pull request check (GitHub App). When a PR starts reading a new env var, the App comments with the name and the exact file and line, and warns if it isn't declared in .env.example. It also flags a committed .env file and anything on the added lines that looks like a leaked key.
2. A scan on every push (CLI). The CLI checks the whole repo the same way and reports, per deploy:
- variables the code reads but nothing declares,
- variables declared but never used,
-
.envand.env.exampledrifting apart.
Try it on your own repo. In dry-run mode it sends nothing:
npx -y deployhealth-scan --dry-run
3. Uptime checks that know about deploys. If an endpoint starts failing soon after a deploy, the alert names that deploy and the env vars it introduced. From the live demo:
Acme API started failing 4m after deploy b52952e, which introduced 2 missing env vars: REDIS_URL, STRIPE_KEY
It only ever handles variable names, file paths and line numbers, never values.
What it doesn't do yet
- It reads JS/TS, Python, Go and Ruby. On other stacks, the PR check says it can't check the repo yet instead of showing a green pass.
- It can't see what's actually set on Vercel or Railway, only what's in your repo.
- It's regex-based, so dynamic reads like
process.env[name]are missed.
How I built it
I built deployhealth in close collaboration with Claude Code, and the process mattered as much as the code.
Every phase started with a written plan from me: what to build, what must never happen, and how to test it. Before any code was written, the plan came back as a list of numbered decisions, and I approved, changed or rejected each one. Then came small commits, the full test suite, and a summary I reviewed before anything was merged. I read the risky parts myself: authentication, the webhook, anything public, anything that touches someone else's repo.
Some of the most useful moments came from that back-and-forth:
- A security review before launch found that one crafted pull request could freeze the worker that runs the checks. The check now runs in an isolated thread with a time limit, and every parser runs in linear time, with tests that fail on the slow version.
- A friend tried it on a Java/Spring repo and got a green check that had read nothing. That became a whole phase: every check now says what it actually read, and a repo it can't read says so.
- The docs and the code once disagreed about how jobs were locked. Since then I check every claim against the code, not just the summary.
The infrastructure, the releases and the final call on every decision were mine. The result is something I understand end to end and can stand behind.
Try it
- Live demo, no login: https://deployhealth.dev/demo
- GitHub App: https://github.com/apps/deployhealth
- Source: https://github.com/Utsav-mishra-25/deployhealth (the scanner and CLI are MIT, the rest is FSL and free to self-host)
The hosted version is free while in beta. It's built for small teams and freelancers who look after several apps.
Two questions for you: has a missing env var ever taken down your prod, and how long did it take to find? And if you run the dry run on your repo, does it catch anything, or get something wrong?
Top comments (0)