It was a Tuesday afternoon deploy, the kind that's supposed to be boring. CI went green, the deploy job ran, Slack posted its usual "🚀 deployed api to production" message, and everyone went back to what they were doing.
Twenty minutes later, the pager went off. Not with errors — with silence. Our api service's task count in ECS was dropping, and nothing was replacing the terminated tasks. Within ten minutes we had zero healthy instances and a load balancer returning 503s to every request.
The first three guesses, all wrong
Guess one: bad credentials. Someone had rotated an IAM role recently, so the obvious theory was that ECS couldn't pull from our registry anymore. We checked — the pull permissions were fine, confirmed by aws ecr get-login-password working from a bastion host seconds later.
Guess two: a bad health check. Maybe the new code introduced a slow startup path that was failing our /healthz probe. We pulled the CloudWatch logs for a replacement task. There were no logs. Not failing logs — no logs at all. The container wasn't starting.
Guess three: resource limits. Maybe the new image was bigger and hit a memory reservation ceiling on the cluster. Still wrong — the scheduler wasn't even getting that far. The task definition update had gone through, ECS was trying to launch tasks from it, and every single launch attempt was failing at the image pull step with:
CannotPullContainerError: pull image manifest has been retried 5 time(s):
failed to resolve reference "123456789.dkr.ecr.us-east-1.amazonaws.com/api:a1b2c3d":
unexpected status: 404 Not Found
The image tag the task definition pointed at did not exist in the registry. At all. But CI had said the build succeeded.
What actually happened
We went back through the CI run for commit a1b2c3d. The build job history told a different story than the Slack message did. There were two pipeline runs for roughly the same commit window, fifty seconds apart, because someone had pushed a quick follow-up fix right after the first push. Our CI's concurrency rule — "cancel in-progress runs on the same branch when a newer commit lands" — had correctly canceled the first build mid-image-build.
The problem was in what happened after the cancellation signal. Our deploy pipeline didn't key off "did this specific job finish and push an image." It keyed off a row in a Postgres table called last_known_good_build, written by a cleanup step that ran in a finally-equivalent block — intended to record metadata for monitoring even on failure. That cleanup step ran during the cancellation, saw a commit SHA, and wrote it to the table without checking whether an image with that tag actually existed in the registry.
The second build (the real one, for the follow-up fix) then ran, succeeded, and pushed its own image under its own correct tag. But the deploy job that fired afterward read last_known_good_build, which still had the first, canceled build's SHA in it, because of an ordering race between the cancellation cleanup and the second build's completion write.
Two systems, each behaving exactly as designed in isolation — a cancellation handler trying to be helpful, and a deploy job trusting a shared pointer — met at a boundary neither of them owned, and production got told to run an image that had never finished building.
The fix
Two changes, both non-negotiable going forward:
1. The deploy step verifies the image exists before touching the task definition — not implied by a database row, but confirmed against the registry itself.
#!/usr/bin/env bash
set -euo pipefail
IMAGE_TAG="$1"
REPO="123456789.dkr.ecr.us-east-1.amazonaws.com/api"
if ! aws ecr describe-images \
--repository-name api \
--image-ids imageTag="$IMAGE_TAG" > /dev/null 2>&1; then
echo "FATAL: image $REPO:$IMAGE_TAG does not exist in registry. Refusing to deploy." >&2
exit 1
fi
echo "Verified $REPO:$IMAGE_TAG exists, proceeding with deploy."
2. We stopped letting a finally-style cleanup block write "success" state. Cleanup writes failure state only, and a separate, un-cancelable final step writes success state, tied to the specific job ID, not the branch.
3. The biggest one: before any image reaches real production, it now gets launched and smoke-tested in an environment that's a true clone of prod, not a staging environment that's drifted for eight months. We spin up a short-lived VM from a snapshot of our actual production image, run the new container on it exactly as ECS would, hit the real health and readiness endpoints, and only then let the deploy proceed. It's disposable — we tear it down the second the check passes or fails. I'm the founder of Krova Cloud, and this is a big part of why we built it this way: full root access, a prod-identical snapshot you can clone in seconds, billing that drops to zero the moment you power the VM off, and no surprise state carried over between runs because every clone starts from the same clean snapshot. It turned "hope the task definition is right" into "prove the task definition is right," for less than the cost of a coffee per month of these checks.
Lessons
- A green CI run is a claim, not a proof. If your deploy step can't independently verify the thing it's about to run actually exists, it isn't verifying anything.
- Be suspicious of any "record success/failure" logic that lives inside a cleanup or
finallyhandler — those paths run on cancellation too, and they rarely distinguish between "finished" and "got told to stop." - Shared mutable pointers (a "latest good build" row, a "current" symlink, a cached tag) are exactly the kind of state that two independently-correct processes will race on. Prefer content-addressed checks (does the artifact exist?) over pointer-based ones (does the row say it exists?).
- The cheapest insurance against this class of bug is a real, disposable environment that runs the actual artifact before production does — not a mocked check, not a staging environment that's stopped resembling prod months ago.
If you've been burned by a "successful" deploy that wasn't, I'd genuinely like to hear the version of this bug you hit — they're never quite the same shape twice.
I'm Rohit, founder of Krova Cloud — fast, disposable cloud VMs with real reserved resources, instant snapshots and clones, and locked-down networking by default, built for exactly this kind of "prove it before it touches prod" workflow. If you want more deep debugging stories like this one, I write regularly over at debugly.dev too.
Top comments (0)