DEV Community

Afterclose dev
Afterclose dev

Posted on Originally published at Medium

A coding beginner shipped a game to 4 stores by directing AI agents: what worked, what broke

Disclosure: I drafted this with help from an AI assistant (Claude). I checked the numbers against my git history and test output before posting.

On October 2nd I was testing my game on a real iPhone and found something strange. After watching a rewarded ad and coming back, the joystick was frozen where my thumb had last been. That night it got worse: after an ad, every sound effect was gone.

At that moment all 1,450 automated tests were passing. I didn't write the code, and I didn't write the tests. AI coding agents did. The thing that found the bug was my thumb on a real phone.

I'm a beginner at software development, working alone in Korea. I didn't program this game myself; I directed AI coding agents (mostly Claude Code, with Codex and Kimi alternating in). The result is Afterclose: Night Delivery, a one-finger mobile roguelite, now submitted to four places: Google Play, Apps in Toss (a mini-app platform inside a Korean super-app), ONE store, and the App Store. This is the honest log of how that went.

The numbers

Git period Sept 24 to Oct 10, 2026 (17 days), commits on 13 of them
Commits 528 (busiest day: 133)
src/ 66 TypeScript files / 13k lines → 229 files / 72k lines
Automated tests about 1,480 (test files: 32 → 162), plus 52 browser E2E spec files
QA rounds 19, with 18 written reports
Stores Google Play live (Oct 8, 178 countries), Apps in Toss live (Oct 9), ONE store listed, App Store submitted Oct 4 and still in review

One correction before anyone quotes this: it is not "built in 17 days". My first commit imported a game I had already prototyped (290 files, about 34k lines). The numbers above are what changed after that, in 17 days of git history.

How the work actually ran

My job wasn't typing code. It was deciding what to build, playing it on real phones, and making the calls that can't be undone. Everything else ran like this:

  • A guardrail file. An AGENTS.md in the repo with hard rules: never press a store submit or review button without my approval, never print secret keys, never commit before bun run check (lint, type check, tests) passes. Any agent that reads the file inherits the same rules.
  • Read-only audits before building. At the start of a big round I had audit agents read the code and report problems without touching anything. Then I split the work by file area, gave each implementation agent its own git worktree, and merged at the end. Round 12 used ten worktrees.
  • Simulated players for balance. Difficulty was measured by test bots playing hundreds of shifts. I wanted numbers, not vibes.
  • Persona QA. In round 19 six agents, each with a different "hand" (slow reactions, a commuter, a sixty-year-old), played in a real browser.
  • Assets. Most of the art was made with AI image tools and then touched up by me. Sound effects come from CC0 packs and the music is synthesized in code.

What worked

Parallel work stays clean if you split by file area. Round 11 had five workstreams (checkout behavior, gear features, character visuals, map expansion, balance). I committed the shared types and effect events first, then gave each agent its own worktree. The only merge conflict was a generated font file, fixed by rerunning one command.

Ask for numbers, not "is it better?" My difficulty rule was "harder as the days go on, easier when you upgrade, and overall neither too easy nor too hard." I turned that into: "the best day reached by 48 simulated runs stays within ±1 day of the baseline." Once that was a sentence, every change got reported in the same table, by whichever agent did it. I also set the rarer-is-stronger rule in numbers (a common item is worth about half a day of difficulty, the rarest tier five to seven).

A handoff document that makes switching tools free. Because I rotate between Claude Code, Codex and Kimi, any of them has to be able to read one file and know the state, the verification commands, the don'ts and what's left. Each QA round also left a report (18 in docs/qa), so a dropped session never cost me an afternoon of "where were we?"

What broke

1. Tests green, phone dead. Four real-device issues showed up within about twelve hours.

  • The joystick froze after an ad. iOS doesn't deliver the touch-end event to the web view while a native full-screen ad is up, so the game believed the old finger was still down and ignored every new touch.
  • All sound effects went silent after an ad. The ad SDK takes the audio session and gives it back, but the web audio context stays suspended, and iOS ignores a resume that isn't triggered by a touch.
  • The sound effects on the level-up, shift-complete and results screens existed in code but never played: a dead wire.
  • Audio stutter.

The agents diagnosed and fixed them. But the person who noticed was me, tapping through TestFlight. Every one of them became a regression test afterwards, and the order was always "found on a real device, then pinned by a test".

2. A store rejection over text my game never used. On October 7th the Apps in Toss review rejected my bundle with (my translation): "a mini-app bundle to be released can't contain a test ad group ID." The ID was only used in development mode, so it never ran in production. But I decided that at runtime, so the literal string ait-ad-test-… still sat in the bundle, and review reads the text, not the logic. I made the value disappear at build time, added a rule to the build script that stops the build if that string survives, and resubmitted. It went live on October 9th, late at night.

3. The usual store traps.

  • Google Play never lets you reuse a versionCode. I found a bug in version (3) while it was in review and had to replace it with (4). A test still expected "3" and failed until I fixed the expectation.
  • Play rejects screenshots whose long side is more than twice the short side, so everything was re-cut to 1080×1920.
  • Upgrading the Toss SDK to 3.x broke 14 end-to-end tests that relied on a fake bridge. The agent rewrote the fake to match the new protocol.
  • The App Store build was submitted on October 4th and is still in review as I write this.

Five lessons

1. Ask an AI to grade its own work and it will be generous. I had the QA scored by the AI. After each follow-up audit the marks went down: 8.9 → 8.5 for round 7, 9.3 → 8.7 for round 8, 9.3 → 8.8 for round 9, 9.4 → 9.1 for round 10, 9.5 → 9.2 for round 11. Five times in a row, 0.3 to 0.6 points. The read-only audit before round 8 found three serious bugs the earlier seven rounds missed: a delayed save restore could overwrite a backup, changing the device date repeated a daily reward, and a refund left a permanent purchase in place. Audit first, then build.

2. Green tests only guard what the AI thought of. The same agents wrote the code and the tests, so they share blind spots. Silence after an ad only appears on a real phone after a real ad. Before launch I needed time with my own hands on the game.

3. Distrust the ruler first. With three simulation runs one metric looked 19% worse; with six runs it was 5% worse. A baseline average of 24.4 over 16 runs turned into 23.7 over 48. Once a bot got stuck for 45 seconds in front of a shelf, which made a new store look hard (63 to 75% first-shift success) when it was actually fine (100% after the pathfinding fix). And when a round made the game suddenly +2.9 days easier, the cause wasn't the kill-score change I suspected; it was gear drops.

4. Review reads your bundle's text, not your intent. Turn rejections into automated checks. After the Toss rejection I had the agent add the scan to the build step. A mistake you only have to be burned by once is a good deal.

5. A beginner's job is judgment and checking the real thing. My round-11 audit found that the phone version has only a joystick and one button, while most of the boots' options were tied to keyboard-only skills and did nothing on a phone. The agents couldn't decide whether to add a dash button; when I asked "why is there no dash?" they built one. And in round 19 an agent reported "the timer runs while you read the cards"; I had the code opened and it wasn't true. Deciding what's true, and not accepting the AI's report at face value, was my work.

What's next

  • Waiting on Apple's review.
  • The German, French, Traditional Chinese and Brazilian Portuguese store text is AI-translated and hasn't been checked by a native speaker.
  • There is currently no crash reporting, so errors on players' phones don't reach me automatically.
  • Balance was measured only with bots. Real player data will change the early-game quota curve.
  • Promotion is where I'm a genuine beginner too. My first video got 230 YouTube views in about 29 hours. There is nothing to brag about yet.

If you're curious about the game

Afterclose: Night Delivery is a one-finger roguelite arcade. A delivery robot, Pico, works the night shift in a darkened store: grab cargo, cash out at the register, hit the quota, escape. Gear enhances up to +20, a failed enhancement never destroys gear (from +11 it only drops one level), and it's free to install.

Questions and "this part is wrong" corrections are welcome. I especially want to hear the store traps you hit when you shipped with AI agents.

Top comments (0)