DEV Community

Cover image for Logbook of a Spec-Driven Developer
Ariel Sidi
Ariel Sidi

Posted on

Logbook of a Spec-Driven Developer

Building a 45k-line app with AI agents: what worked, and what broke

“Now that we have AI, why can’t we do this project in 3 weeks instead of 3 months?”

I had a very good year with AI, achieving things in areas that were new to me much faster than the traditional way. Suddenly everybody was excited: delivery to production had multiplied, and things postponed for years, like technical debt, were finally getting out of the way. It didn’t take long for product managers and executives to notice, and to start asking questions like the one above.

As a senior engineer, I am very clear that I am 100% responsible for the code that reaches production. When someone from product raises their eyebrows and asks to do things very fast, my legs shake. The AI can spit out a mountain of code without breaking a sweat, but it is not the AI who assumes the consequences of a work badly done, of an application that does not align with the business goals, or of a random code without cohesion between its parts.

So we are struggling between 2 dimensions: Speed & Quality. How effective can we be before quality starts to decrease? What I mean by quality is alignment with business goals and alignment with architecture vision. This is where I thought Spec Driven Development (SDD) could solve both problems: it lets you build fast what the team has decided to and how the team has decided. Before answering that question at work, I wanted to find out on my own. So I proposed myself this experience in solitary, to see if that promise holds.

Why Game of Life Studio?

First, one line about me: I have been building software for 25 years, the last ones as senior engineer and engineering manager. For this experiment I used the BMad method, one of the SDD frameworks: you produce a Project Brief, then the Product Requirements, then the Architecture, and only then the AI implements, story by story. The stack is React and TypeScript with a typed-array simulation engine, and the result is live at game-of-life-studio.com.

Why this project? I had run a workshop a year before at my children’s school related to Conway’s game of life. Something to share and have fun with my kids, free user testing, and who knows what this can bring in the future out there. As time passed during the project I was happy to see that it had the right level of complexity needed to prove that the quality expected was not trivial. It was also fun to see how my kids were asking every day, “Is there anything new that we can see?”. They were my Product Managers!

 Four organisms competing on a 200×120 grid in Game of Life Studio: one colony expands across the centre while another holds the corner

The Planning Phase

Shaping a perfectly specified product before starting to build it has always been a big challenge and has required experienced cross functional team roles and responsibilities. Working with SDD methodology helped me a lot in sorting this out. BMad skills forced me to cover each detail of the product requirements, beginning with the Project Brief, following with Product Requirements and lastly with the Architecture. Each phase had a clear Acceptance criteria that could not be skipped until considered done by the Framework. Thanks to this, very few gaps could have sneaked into the Execution phase. AI is incredibly efficient on detecting missing requirements, edge cases and contradictions in a spec.

But this precision has a price. It became hard to wrap up the PRD phase after having the Architecture advanced, the issues kept coming out, regressions, etc. I felt stuck and unable to move on, since it was mandatory to resolve any issue, even the ones marked as low priority. I believe that we typically defer some decisions that are not so relevant, and this was not allowed — it sure is something to improve with SDD. I also had to stop several times to build and optimize a project-context file, since it started to consume enormous amounts of tokens as everything grew.

The other price was communication. The AI is like an employee that is eager for a promotion: it outputs too much content, full of references like AR-33, AR-32 or UX-DR20 in the same paragraph, and poetic terms — it took me weeks to read “Roster” without thinking of an Alice in Chains song. That story deserves its own article.

 The AI's planning review: two sentences packed with references such as AR-33, AR-32, AR-36–38 and UX-DR1–UX-DR20

I remember thinking “I can’t wait to start the execution phase, and see how fast I can get the application implemented with best practices.”

The Execution Phase

Things didn’t turn out too bright at the beginning. In Epic 1 the process was very slow and got me absorbed in front of the keyboard, after all that long planning phase. On one side because of the communication style of the AI, and on the other side because some parts of the new code purpose could not be understood until implementation got further. At some point I found myself blindly approving decisions, and I kept asking myself: what would I do in a real working environment? I am also too proud to let AI do all the coding without my intervention, even if it was following my solution design.

It was an interesting BMad proposal to create the Story specs right before the implementation and after the previous story was done, so that any deferred work can be attended and specs are up to date with it. Also, every part of this process (Create Story > Implement > Review) had to be done in a fresh session, with the review on a different model than the implementation, which made a lot of sense.

After each PR was ready I was doing a thorough review and asking for changes and improvements, although I have to say that from the beginning I was really happy with the resulting code. Then I realized that it was me slowing down the process. The whole idea of SDD is that, on execution phase, this should go very fast.

Key Decision 1: The refactoring (in-between) sprints

My first decision before starting epic 2 was to split the review into two moments. During the epic, I would do only a quick review of each PR before merging it, without requesting changes (at most taking notes), so I wouldn’t slow down the implementation of the whole epic. Once the epic was done, I would take the time for a deeper review: create a refactoring branch and introduce improvements through vibe coding. Not because I did not trust the code, it looked very good, but this became a personal challenge and at the same time a good opportunity to keep up with the details of the implementation. In order to keep track of what were my contributions I would prefix each commit with “HITL refactor:” (HITL = Human In The Loop), and the PR was not merged with Squash, since I wanted to keep these improvements individually. Once I was happy, or I just did not want to keep dedicating time to this, I would create a PR that would be reviewed by Opus. Lovely!

Key Decision 2: Auto-accept + Orchestration Skill

My second decision was to accelerate the process. I had studied the Ralph loop (“let it run overnight and see in the morning”) and that was my original intention for epic 2. But I feel like this is not right, unless you are working in an experimental repository or a PoC. If you are working on important code that will be shipped to production, you should always review story by story and approve. Remember: it is not the AI who assumes the consequences.

So instead I created the “implement next story” skill. What the skill does: it takes the next story of a sprint board and turns it into a pull request, spawning three fresh subagents in turn — create the story, implement it, review it on a different model — and stops with the PR open for one human to merge. It implements nothing itself; it orchestrates, reads artifacts off disk, and refuses to guess.

The implement-next-story pipeline: sprint board, then create story, implement, and review on a different model, each in a fresh subagent, then the PR stops for a human to merge

This was an inflexion point for the project. This is when I finally started to move fast and deliver good work at the same time, what we all want to achieve with AI.

What it measured

At the middle of epic 2 I had the idea to make the skill gather execution stats, taken from the runtime’s own transcripts — measured, not estimated. Thirty stories went through the pipeline. The median from sprint board to open PR was 67 minutes (range 43–106 across the uninterrupted runs): 10 minutes to create the story, 26 to implement it, 32 to review it and open the PR. The median story produced 180k output tokens — and 82.6 million cache-read tokens, 97% of the total, which is what makes the bill reasonable. Thirteen stories were implemented on Sonnet and reviewed by Opus; seventeen on Opus, reviewed by Fable after it came out (the first seven had to settle for a Sonnet review, and I will come back to why that matters). Never the same model twice.

Median per story across 30 stories: 67 minutes from start to open PR (create 10, implement 26, review 32), 180k output tokens, 82.6M cache-read tokens

Key Decision 3: Tasks Parallelization

Things started to move fast and smooth as I got to Epic 3, but there was still a lot of work ahead. I was lucky that Epic 3 and Epic 4 did not depend on each other mostly, because I hadn’t considered parallelization during the planning (my bad!). So I adapted the skill in order to run 2 agents at the same time, one per each epic. I learned about git worktree and iteratively improved the process, dealing with some issues on the way.

What broke

Not everything went smooth, and I think this is the most useful part of this article. Almost every rule the skill has today is there because something went wrong first.

Two agents, one folder. The first day I ran two epics in parallel, I launched both agents from the same folder. Each one checked at the start that there were no pending changes, and there weren’t, because the other one hadn’t written anything yet. A few minutes later the agent working on story 3.8 made its commit, and it included the status lines that the other agent had just written for story 4.1. One story’s commit carried changes that belonged to the other. That’s how I learned about git worktree: one folder per lane. The lesson I take from it applies anywhere: a check that runs at the start only protects you from the state at the start. If something must be true, check it again right before any step you can’t undo, like branching or committing.

A few days later I did it again (my bad!). I opened two terminals in the main folder, even though one of the epics already had its own worktree. Which epic lived in which folder was only in my head. Now the skill figures that out by itself, and it locks the folder so a second session stops before writing anything.

Before: two agents in one folder, and one commit sweeps in the other's status lines. After: one folder per lane, each with a lock

The reviewer reviewing itself. The whole point of the review step is that a different model looks at the code. The first version of the skill always used Opus as reviewer, so when Opus implemented a story, Opus reviewed its own work. Nothing failed and nothing warned me; it just silently stopped being a second opinion. My first fix, “use the other model”, created a new problem. The most important stories, the ones I sent to Opus because the rest would copy their patterns, ended up reviewed by the weakest model. Now it’s a table: Sonnet is reviewed by Opus, Opus is reviewed by Fable. Before starting the review, the skill has to say which model implemented and which one will review. A silent mistake became a visible one.

The 65-hour review. When I added the stats table, one story showed a review phase of 65 hours. The AI didn’t spend 65 hours reviewing. The clock was counting nights, usage-limit resets, and the time the PR was waiting for me to make a decision. Now the stats count only active time and list every gap they left out. It was also the first sign of something I’ll come back to at the end: the slowest part of the pipeline was not the AI anymore.

What I’d tell someone starting tomorrow

This was for me a very good start as a Spec-Driven Developer, and I am sure this experience will help me to get the best of AI and achieve better results faster. However I worked on my own, and I am now eager to try this in a real working environment, with a team and an active business going on. I can tell you that this is not a comfortable path: setting all up for the execution takes time and patience, and I am sure you will be asked several times, “When is this going to be ready?”. One advice for each aspect of the process:

  • Communication with AI: Learning to read the AI instead of making it talk simpler is a good approach. At least look for a balance, specially if you are not a native English speaker; this will prepare you for keeping up with strong communication professional environments.
  • Architecture: I’ve chosen the RFC format because I had been working with it for a long time and felt comfortable with it. Later on I realized that some parts did not make much sense when working with AI — the risk and mitigation section, the alternatives considered — all very artificial, because they were all AI slop.
  • Architecture, again: Where has my intervention been most valuable? When creating the RFCs I was making much better decisions than the AI did. I had a clear idea to build the core with 3 abstraction layers: a generic RulesEngine, an Adapter for Game of Life, and a Simulation Engine with pluggable evaluation strategies. These were the decisions that had the bigger impact on the end result.
  • Planning: Probably one of my bigger mistakes was not planning for parallelization, even when I knew that I was the only developer. My advice here is simple: always plan for parallelization!
  • Execution: The lack of the skill in the BMad method is not a defect of that framework; this is just what I needed and what worked for me. You gotta do the same: adapt the tools to what you, and most important your team, actually need in your own circumstances.

Closing

I couldn’t have ever built an app like this in this short time, and I am impressed with how the resulting code looks. But there is one irony to close with: by the end, the slowest part of the pipeline was not the AI anymore. It was me, deciding on open PRs, and the CI queue behind them. That is a very different problem from the one I started with, and I believe it is the problem many teams are about to meet.

So, to the product manager asking for 3 weeks instead of 3 months: yes, the code can be written much faster, but only after the specs are done, and not faster than the team can review and decide. That is where the time goes now.

The app is live at game-of-life-studio.com, the code is at github.com/sidiar/game-of-life-studio, and the skill has its own repo at github.com/sidiar/implement-next-story. If your team is meeting this problem too, I would love to hear how you are dealing with it.

Top comments (0)