DEV Community

Revin
Revin

Posted on Originally published at revin.com.br

I asked production which commit it was running: 23 behind main, the oldest waiting 19 days

A feature was done. The ticket said done, the pull request had been merged four days earlier, CI was green. The screen did not exist in production.

Nobody had lied to me. The merge happened, the deploy did not. The workflow that ships that service had been failing on a step nobody watched, and the image actually running in the cluster had been built 19 days before.

I found that out because I stopped asking the team and asked the process. Below is the measurement, the commands, and the three sources I tried first that told me nothing useful.

Step 1: make the running process say which commit it is

Most services cannot answer "what are you running?". The fix costs about ten minutes and it is the only reading that does not depend on anyone's memory. Bake the SHA in at build time:

ARG GIT_SHA
ARG BUILD_TIME
ENV GIT_SHA=$GIT_SHA
ENV BUILD_TIME=$BUILD_TIME
Enter fullscreen mode Exit fullscreen mode
// src/routes/version.ts
app.get('/version', (_req, res) => {
  res.json({
    sha: process.env.GIT_SHA ?? 'unknown',
    builtAt: process.env.BUILD_TIME ?? 'unknown',
  });
});
Enter fullscreen mode Exit fullscreen mode
# .github/workflows/deploy.yml
- name: Build image
  run: |
    docker build \
      --build-arg GIT_SHA=${{ github.sha }} \
      --build-arg BUILD_TIME="$(date -u +%FT%TZ)" \
      -t "$IMAGE:${{ github.sha }}" .
Enter fullscreen mode Exit fullscreen mode

That endpoint is the first thing I add when I get access to a service I did not build. Here is what it answered on the day I am describing:

$ curl -s https://api.internal.example/version | jq
{
  "sha": "4f1c9ab",
  "builtAt": "2026-09-11T02:14:07Z"
}
Enter fullscreen mode Exit fullscreen mode

The build date was already the whole answer. I ran it on a Tuesday at the end of September.

Step 2: measure the gap, not the feeling

With a SHA in hand the rest is plain git. Two numbers come out of it: how many commits sit between production and main, and how long the oldest one has been waiting.

$ git fetch -q origin
$ git log --oneline 4f1c9ab..origin/main | wc -l
23
$ git log --format=%cI 4f1c9ab..origin/main | tail -1
2026-09-11T09:33:12+00:00
Enter fullscreen mode Exit fullscreen mode

The commit count is the number people want to put on a slide. The age of the oldest unshipped commit is the one that matters, because it is the actual delay between someone finishing work and anybody outside the team being able to use it.

To see the whole queue instead of just the worst case:

DEPLOYED=$(curl -fsS "$BASE_URL/version" | jq -r .sha)
git fetch -q origin
git log --format='%h %cI %s' "$DEPLOYED..origin/main" | while read -r sha ts subject; do
  age=$(( ( $(date -u +%s) - $(gdate -u -d "$ts" +%s) ) / 86400 ))
  printf '%-10s %3sd  %s\n' "$sha" "$age" "$subject"
done
Enter fullscreen mode Exit fullscreen mode

On macOS that needs gdate from coreutils, since BSD date has no -d. On Linux drop the g.

a19f4c2      1d  fix: null check on invoice pdf
8d02e71      2d  feat: csv export for admin
...
7c3e8b1     19d  chore: bump node to 22
Enter fullscreen mode Exit fullscreen mode

Nineteen days of finished work waiting on a pipeline step. The team's speed was never the problem.

Step 3: the three sources that fooled me first

I did not start with /version. I started with the places that are supposed to answer this, and each one reported something adjacent to the truth.

  • The GitHub Deployments API. gh api repos/:owner/:repo/deployments came back as []. The workflow shipped the service for two years and never created a deployment record, so the API that exists exactly for this question had nothing in it.
  • The last green run of the deploy workflow. gh run list --workflow deploy.yml --json conclusion,createdAt,headSha --limit 5 showed a success six hours old. The final step of that job was aws ecs update-service with no --wait, which means the job succeeded when the API accepted the request. Whether any task came up healthy with the new image was outside the job's knowledge. The pipeline reported what it asked for, not what happened.
  • The registry tag. latest had been pushed three hours earlier. A push is not a rollout, and in this case the rollout had been failing on a readiness probe and silently rolling back.

All three were honest about their own scope. None of them knows what is serving traffic. The running process is the only thing that does.

The number I keep after the one-off check

I do not keep the commit count. I keep the queue age, sampled once a day, written somewhere boring:

#!/usr/bin/env bash
set -euo pipefail
DEPLOYED=$(curl -fsS "$BASE_URL/version" | jq -r .sha)
git fetch -q origin
OLDEST=$(git log --format=%cI "$DEPLOYED..origin/main" | tail -1)
if [ -z "$OLDEST" ]; then
  echo 'deploy_queue_age_days 0'
  exit 0
fi
NOW=$(date -u +%s)
THEN=$(date -u -d "$OLDEST" +%s)
echo "deploy_queue_age_days $(( (NOW - THEN) / 86400 ))"
Enter fullscreen mode Exit fullscreen mode

It runs in CI on a schedule and the output goes to the same place as every other gauge. Over the following two weeks the value sat at 0 or 1 on most days, with one spike to 5 that came from a migration everybody had agreed to hold. A spike with a known reason is fine. A spike nobody can explain is the broken deploy step coming back.

The reason I like this metric over anything derived from tickets: a ticket board records what people said. The /version endpoint records what the machine is doing. Only one of those changes when a YAML step breaks.

Where this breaks down

  • Monorepos. One /version per service, and the git range has to be scoped by path (git log 4f1c9ab..origin/main -- services/api), otherwise the count fills up with commits that were never meant to ship with that service.
  • Trunk based development with feature flags. Code can be live while the behaviour is off. The gauge goes to 0 and says nothing about flag state, which is a separate measurement I have not solved well.
  • Squash and force push on main. If the deployed SHA no longer exists after a rebase, the range query dies with unknown revision or path not in the working tree. I hit that once and the fix was to tag every deploy (git tag deploy/$(date -u +%Y%m%dT%H%M%S) $SHA && git push --tags) so the ancestor survives history rewrites.

It also does not help in the first weeks of something still being discovered. There is nothing in production yet, and asking for this number just creates noise.

The part I am still unsure about

My daily sample is cheap and dumb. It tells me the worst case and nothing about distribution, so a queue of one 19 day old commit looks worse than a queue of twelve commits from yesterday, when the second situation is often the more urgent one.

How do you track what is actually live? Deployments API with real records, a version endpoint like this, image digest comparison against the cluster, something in the service mesh? I am mostly interested in the teams that caught a silently failing rollout and want to know which signal caught it.

· · ·

Originally published on the Revin blog: https://revin.com.br/en/blog/how-to-evaluate-developer-work-without-reading-code

Top comments (1)

Collapse
 
kashif_manzer profile image
Kashif Manzer •

One thing I'd add: the /version endpoint tells you what the app claims it's running, which breaks if an image ever gets swapped outside the normal deploy path or a node runs a stale pull. Since this is Kubernetes, it's worth also stamping the commit into the image's OCI label, org.opencontainers.image.revision, at build time. Then the container imageID from the pod status gives you a platform-level source of truth that doesn't depend on the app answering at all. The version endpoint is still the friendlier interface for humans; the image metadata is the more trustworthy one for automation.