DEV Community

Cover image for Claude Code ran for 19 hours 53 minutes without stopping. Here's the setup that made it possible
ShengXing Chi
ShengXing Chi

Posted on

Claude Code ran for 19 hours 53 minutes without stopping. Here's the setup that made it possible

Someone tell me: where is Claude Code's ceiling?

Lagos Life blew up on Oct 1. Everyone in Nigeria was playing it.

I wondered: could one person and an AI build their own Lagos life-sim from scratch, fast?

So on Oct 5 at 8:45pm I gave Claude Code a spec and a set of rules, and let it go.

I went to bed. I woke up. It was still going.

I checked at breakfast. Still going. I checked at lunch. STILL going.

19 hours and 53 minutes. Non-stop. I honestly couldn't believe what I was looking at. ๐Ÿคฏ

By the next afternoon there was a playable multiplayer game: 3D character creator, apartments, jobs, a city map of Agege, an economy, DMs, house visits.

4 days later it's live: 211 commits, ~169k lines of code, 186 test files.

Play both. Then tell me AI can't build games. ๐Ÿ˜

๐Ÿ‘‰ Play Agege Life

How it kept going

So how do you keep an AI agent working for 20 hours without it drifting, faking progress, or stopping to ask you something at 3am? No magic prompt. Four boring things.

1. One file is the loop's memory

Everything lives in BUILD_PROGRESS.md: a numbered list of every item in the build, each with its own check.

- โฌœ 1.8 Shifts for all 5 careers: 1-hour shift, once per day, work modes,
  4 tasks, leave early partial pay, result dialog, promotion rule.
  Check: rules tests + e2e with fake clock.
Enter fullscreen mode Exit fullscreen mode

Every iteration does the same thing: read the file, take the first โฌœ item, build it end to end, verify it, mark it โœ…, append one log line. That's it.

The agent never has to remember what it did three hours ago. The file remembers. A long run is just the same small loop, 36 times.

2. "Done" has a hard definition

An item only gets its โœ… when:

  • typecheck, lint and test all pass
  • the item's own check passes (a test, or a headless browser run)
  • screenshots are taken at five viewports (phone portrait and landscape, tablet both ways, desktop) and looked at โ€” anything that reads as a placeholder gets redone
  • the performance harness is re-run and no budget regressed
  • every command that moves money has concurrency tests (replay, double-spend, 20 tabs at once)

The rules are written down too: no placeholder art, no layout shift between states, one design system, every naira through a double-entry ledger. If it's not written, don't expect it.

3. Blocked? Write it down and move on

The thing that kills long runs is the agent stopping to ask a question while you're asleep. So there's one rule for that:

If blocked on something only the owner can do, write it under Blocked with the exact ask, skip to the next item that does not depend on it.

During the run, the full load test needed a staging deploy only I could do. It wrote the ask down and kept building. I read it in the morning.

4. Every item leaves evidence

Each item gets a folder with a summary and its screenshots, and one line in the log:

2026-10-05 ยท 0.4 Command pipeline ยท same id ร—10 โ†’ one effect,
last-balance race โ†’ one wins, 20-tab race โ†’ exactly 10,
โ‰ค 5 queries/command ยท run log docs/autopilot-runs/...
Enter fullscreen mode Exit fullscreen mode

78 folders later, I can review a night of work over coffee instead of reading 20 hours of transcript.


What it got wrong

It's not magic. Three features I needed (sickness, account deletion, admin numbers) were never clearly on the list, so they weren't built. They were added in the next session, the same afternoon.

The loop builds exactly what the list says. The list is the product.

What my job became

I stopped writing code. I wrote the spec, the rules and the definition of done, made the product calls, and reviewed the evidence.

Then I hit enter and went to bed.

So, again: where is the ceiling? ๐Ÿ‘€

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

The state file pattern solves context drift, but the hard ceiling on runs over 10 hours usually turns into uncommitted artifact pollution.

If step 4 adds a migration or an asset that passes its own unit test, step 28 can pick up that side effect without the test harness realizing it was an undeclared dependency. Tying each checkbox directly to a clean git commit and tagging the commit hash into the evidence folder makes bisecting trivial when step 35 passes but quietly broke step 2.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.