DEV Community

Cover image for I built "boss levels" for my Linux & Docker game: you're on call and prod is down
Doode
Doode

Posted on

I built "boss levels" for my Linux & Docker game: you're on call and prod is down

Every tutorial teaches you how to start a container. Almost none teach you what to do when one is down at 3am and people are waiting.

So I added "boss levels" to DockerLinux, my free game for learning Linux and Docker in a simulated terminal (here's what shipped last time). Type oncall and you take the pager.

A real on-call run: the page, the investigation, the fix and the postmortem

What it looks like

  1. The page. A ticket arrives: SEV-1 · INC-101 · The shop is down. Support says customers can't load the shop; an ops bot mentions the server rebooted at 03:12 for kernel updates.
  2. You're on prod-01. The terminal logs in to a pretend production server where something is broken, and a clock starts (5 minutes for this one). Your own machine is untouched and comes back afterwards.
  3. You investigate with real commands: docker ps -a, docker logs, curl, docker network inspect. There's a runbook.md to cat, and hints if you're stuck (they cost XP).
  4. You fix it, any correct way. Broken containers crash again on restart, like real ones, so "just restart it" doesn't always work.
  5. The postmortem: time, stars, the root cause and the lesson.

The page and the postmortem

There are four incidents, unlocked as you finish the levels that teach what you need:

Incident Unlocks after You'll need
The shop is down Level 5 (container lifecycle) docker ps -a, docker start
Welcome to nginx?! Level 6 (networking & ports) reading the PORTS column
Database unreachable Level 6 docker network inspect, docker network connect
Bad deploy Level 8 (building images) docker logs, rolling back to the previous tag

(Spoiler for the first one only: the server rebooted and shop-web had no restart policy. docker ps hides stopped containers; docker ps -a doesn't.)

How it works under the hood

Everything runs in the browser: a small shell, a virtual file system and a simulated Docker daemon (images, containers, ports, networks, volumes). Incidents are just data:

{
  id: 'inc_shop_down',
  code: 'INC-101',
  title: 'The shop is down',
  severity: 'SEV-1',
  requires: 5,                 // unlocks when Level 5 is complete
  slaMs: 5 * MIN,              // the time target
  setup(machine, now) { /* break a fresh prod-01: shop-web exited, redis still up */ },
  resolved: (m) => !!serving(m, '8080', 'shop:1.0'),
  hints: [ /* where to look → what it means → the fix */ ],
  solution: ['docker ps -a', 'docker start shop-web', 'curl localhost:8080'],
  rootCause: '…', lesson: '…',
}
Enter fullscreen mode Exit fullscreen mode

Two decisions made it work:

  • resolved() checks the state, not your commands. It asks, "is something serving shop:1.0 on port 8080?", not "did they type docker start shop-web?". So docker start, docker restart or a brand-new docker run with the right port all count, just like in real life.
  • The solution is replayed by the tests. Every incident ships with a known fix, and the test suite runs it against a fresh broken server, so an incident can never become impossible after an engine change.

Scoring is deliberately simple: 1 star for fixing it, +1 within the time target, +1 without hints. Replay for three.

Why I think this matters

Most of the job isn't typing docker run. It's reading docker ps -a calmly, finding the one log line that matters, and fixing the right thing, under a bit of pressure. A game is a safe place to practise that before it's real.

Your turn

  • Which real outage should become incident #5? A full disk, a DNS surprise, a volume mounted in the wrong place, an OOM-killed container…
  • If you run Docker in production: where does the simulation behave differently from the real thing in a way that could mislead someone?

👉 dockerlinux.com: free, in your browser, no install. Finish Level 5, then type oncall.

Top comments (0)