Every tutorial teaches you how to start a container. Almost none teach you what to do when one is down at 3am and people are waiting.
So I added "boss levels" to DockerLinux, my free game for learning Linux and Docker in a simulated terminal (here's what shipped last time). Type oncall and you take the pager.
What it looks like
-
The page. A ticket arrives:
SEV-1 · INC-101 · The shop is down. Support says customers can't load the shop; an ops bot mentions the server rebooted at 03:12 for kernel updates. - You're on prod-01. The terminal logs in to a pretend production server where something is broken, and a clock starts (5 minutes for this one). Your own machine is untouched and comes back afterwards.
-
You investigate with real commands:
docker ps -a,docker logs,curl,docker network inspect. There's arunbook.mdtocat, and hints if you're stuck (they cost XP). - You fix it, any correct way. Broken containers crash again on restart, like real ones, so "just restart it" doesn't always work.
- The postmortem: time, stars, the root cause and the lesson.
There are four incidents, unlocked as you finish the levels that teach what you need:
| Incident | Unlocks after | You'll need |
|---|---|---|
| The shop is down | Level 5 (container lifecycle) |
docker ps -a, docker start
|
| Welcome to nginx?! | Level 6 (networking & ports) | reading the PORTS column |
| Database unreachable | Level 6 |
docker network inspect, docker network connect
|
| Bad deploy | Level 8 (building images) |
docker logs, rolling back to the previous tag |
(Spoiler for the first one only: the server rebooted and shop-web had no restart policy. docker ps hides stopped containers; docker ps -a doesn't.)
How it works under the hood
Everything runs in the browser: a small shell, a virtual file system and a simulated Docker daemon (images, containers, ports, networks, volumes). Incidents are just data:
{
id: 'inc_shop_down',
code: 'INC-101',
title: 'The shop is down',
severity: 'SEV-1',
requires: 5, // unlocks when Level 5 is complete
slaMs: 5 * MIN, // the time target
setup(machine, now) { /* break a fresh prod-01: shop-web exited, redis still up */ },
resolved: (m) => !!serving(m, '8080', 'shop:1.0'),
hints: [ /* where to look → what it means → the fix */ ],
solution: ['docker ps -a', 'docker start shop-web', 'curl localhost:8080'],
rootCause: '…', lesson: '…',
}
Two decisions made it work:
-
resolved()checks the state, not your commands. It asks, "is something servingshop:1.0on port 8080?", not "did they typedocker start shop-web?". Sodocker start,docker restartor a brand-newdocker runwith the right port all count, just like in real life. -
The
solutionis replayed by the tests. Every incident ships with a known fix, and the test suite runs it against a fresh broken server, so an incident can never become impossible after an engine change.
Scoring is deliberately simple: 1 star for fixing it, +1 within the time target, +1 without hints. Replay for three.
Why I think this matters
Most of the job isn't typing docker run. It's reading docker ps -a calmly, finding the one log line that matters, and fixing the right thing, under a bit of pressure. A game is a safe place to practise that before it's real.
Your turn
- Which real outage should become incident #5? A full disk, a DNS surprise, a volume mounted in the wrong place, an OOM-killed container…
- If you run Docker in production: where does the simulation behave differently from the real thing in a way that could mislead someone?
👉 dockerlinux.com: free, in your browser, no install. Finish Level 5, then type oncall.


Top comments (0)