I approved a prompt on a Tuesday morning and went to make a cup of Irish Afternoon tea. By the time I came back the block was done, tested, and more thorough than what I had asked for. I stood there with the mug and realized I had never decided whether that thoroughness belonged in that block. I had not really decided on the block either.
Sixty-two public repositories since July 2025. Three hackathon wins, two challenge wins, two talks on AI in the public sector, and a weekday job running court operations for the county. Kiro, the agent that writes my code, does not get tired. Neither does Claude, which reviews it. I do, and I keep working through it anyway.
I have not fixed this. I am still in it.
Most of what I see us talk about is speed. Faster down the road, faster through the build, faster into the code. Faster is good. There are only so many hours in a day and we are maintaining, prototyping, and sprinting through all of them. What I have not seen much of is how to set up an ending, or how to decide what deserves a beginning. Have you?
I tagged a folder throwaway and we shipped it production-grade
Short version, and the long one is here if you missed it. Block Zero on Porch Light, a civic tool that reads Ventura's public meeting agendas, existed to answer one question. Does the stack deploy. I tagged the folder [THROWAWAY] that morning and budgeted two hours.
It took a build day. By the end there was a byte-identity test protecting a vendored logging module, exact dependency pinning, and a sync script with drift detection, all inside a folder already marked for deletion. Kiro proposed them. Claude reviewed and did not object. I approved them. Each decision was defensible on its own.
What I did not have then was a reason, beyond nobody checking the tag. I had gone into that build determined not to over-engineer, and over-engineered anyway. Mikhail, in the comments, put his finger on why that is worth saying out loud: knowing about a trap is not immunity to it. It is the same reason checklists carry items everybody already knows.
The reason I stop is the reason they do not
I found the reason in a sports article, which is not where I was looking. I read basketball, almost exclusively. This one I clicked anyway.
Chris Borland quit the NFL at 24 over concerns about brain damage. In an essay for The Athletic he writes about the decade since, and about trying to be both driven and content. That is where I ran into the word. Satisficing. Not a lowering of standards, he writes, a stop sign at the point of diminishing returns.
The word belongs to Herbert Simon, who built it from satisfy and suffice. Simon won the Nobel in economics in 1978 for bounded rationality, the idea that people are not decision machines. We cannot take in all the information, weigh it, and produce the optimal choice, because we have limited time, limited information, and limited attention (Simon, 1955).
His strategy for living inside that limit has three steps.
- Set a standard for what would be good enough.
- Take the first thing that meets it.
- Move on.
Step three does the work. Most of the waste is in hunting for an answer only slightly better than the first acceptable one. Note to self here!
Simon also argued that environments shape decisions through problem spaces, the ground your mind has to cross to get from a problem to a solution (Simon, 1956). Some ground is simple and effort pays directly. Need to get stronger, lift weights. Other ground is complicated, and effort alone does not pay. Grinding harder there produces indecision and overload instead of a better outcome.
Here is the part that reframed Block Zero for me. Satisficing is an adaptation to scarcity. We stop because we run out of hours. That is the whole reason the instinct exists.
My agents do not have that scarcity. Kiro will keep hardening a throwaway folder until something stops it, and everything it adds will be locally correct. It is not tired at midnight. It has no sense that the project has a shape or that the shape has an end. The constraint that used to stop the work was mine, and I handed the work to something that does not share it.
And it is not only the building. Claude drafts architecture and documentation. ChatGPT weighs in on what to build with. Gemini makes the visuals. Kiro writes the code. Every stage got faster, including the stage where I decide whether something is worth starting at all. That stage has no test suite, no checkpoint, and no diff to hold a decision against. It is the one I was standing in with the mug.
So the stop sign has to be installed by hand now. Is that a new job, or an old one I was doing without noticing I was doing it? I think it is new, and I would be interested to hear if you read it differently.
The stop sign goes in the plan, not in my willpower
Two places, because there are two scales.
Per block: a rigor budget. I used that phrase in the Block Zero post without knowing how to build one. Two people in the comments handed me the shape.
Suzanne Chartier suggested a steering document that defines levels of rigor by whether the work is exploratory, temporary, or production-bound, with the human deciding which level applies. anassBld described the rule his team enforces on spikes: zero abstraction in Phase 0. One flat script, a direct credentials check, one invocation, assert the output, print the receipt, exit. Architecture only after raw execution is proven against reality.
What I am running now is both of those. Every block gets a tier before the prompt goes out, decided by one question. What happens to this code after the block passes? Discarded is spike tier. Kept and built on is working tier. Touched by a user is full tier.
The part that made it enforceable was writing the tiers as prohibitions rather than effort levels. "Spike rigor" is a feeling nobody can check. "No test files, no dependency pinning, no sync scripts, no refactors outside this folder" is a list you can hold against a diff, and so can an agent. The tier goes in the prompt itself, not only in the steering document, next to the "do not refactor other code" line that has been in every prompt for a year.
Some things no tier defers. Outbound rate limiting, because a scraping loop hits somebody's city server on its first run. What logs must never contain, which is how a real leak got into Porch Light through a framework default that printed the model's thinking to stdout. A try/catch on every fetch. Cost and loop bounds. Those are not code quality. They are harm that lands outside my repo before any checkpoint could catch it.
Per project: a wind down block, written on day one. I have not come across this one elsewhere, which may say more about what I read than about what exists. I built it because the rule I had been following was failing.
My standing rule was to shelve a project when the excitement is gone. The problem is that a rule without a mechanism feels identical to quitting. So Porch Light's build plan has Block 7, written before Block 1: what happens on the day after winners are announced, decided in advance rather than in the moment. A monthly dollar ceiling past which it goes dormant on its own. A handoff note to future me covering what it does, what is stubbed and why, and what would make it worth picking back up. A final honest entry in the wins file, including if it does not place.
The mechanism is what turns shelving into a completed step instead of an abandonment. At least that is the theory. Block 7 has not run yet, so ask me in October.
Sometimes more is right, and the tell is who is on the other end
Block Zero taking a day is why I asked the next question, and I asked it about the whole build. Was this pipeline over-engineered for reading fifteen agendas? Run lock, retry layers, spend ceiling, honest empty states, all for one small city.
My answer was no, and the reason had nothing to do with the code. Tools that watch public meeting agendas already exist. They are sold to lobbyists and government affairs teams, priced for people whose job is watching agendas. The resident who might lose a parking lot has nothing. A flaky civic version of that tool proves the enterprise pricing was correct, that this is hard, that ordinary people should not expect it. A reliable one proves the opposite.
So the reliability work was the argument, not decoration around it. I cannot get to that answer by asking how the code feels. I get there by asking who is on the other end and what a shaky version would prove about them.
Where this leaves me
Both of those mechanisms stop the work. Neither one decides how much work there should be, and that is the question I am stuck on. My agents do not tire, so all of the tiring happens to me, which makes what I take on my problem alone.
I do not have this solved. I have two mechanisms that work at two scales and a pace I have not decided whether to keep.
What I keep turning over: what should I take on next, and what is the honest reason for taking it. When is the right time to stop something that is still working. Whether this pace is a season or a habit, and whether the difference matters. Whether the answer is to slow down or to reorganize, which are not the same choice.
I would like to know how other people are handling it. Not the productivity answer. The real one. What have you cut, what did it cost you, and what finally told you it was time.
References
Borland, C. (2026, September 9). I walked away from the NFL at the age of 24. Here's what I've learned since. The Athletic. https://www.nytimes.com/athletic/7574837/2026/09/09/chris-borland-nfl-lessons-retirement/
Cordero, L. (2026). Block Zero: Oh no. Claude, Kiro, and I over-engineered the throwaway. DEV Community. https://dev.to/earlgreyhot1701d/block-zero-oh-no-claude-kiro-and-i-over-engineered-the-throwaway-5d42
Simon, H. A. (1955). A behavioral model of rational choice. The Quarterly Journal of Economics, 69(1), 99–118. https://academic.oup.com/qje/article-abstract/69/1/99/1919737
Simon, H. A. (1956). Rational choice and the structure of the environment. Psychological Review, 63(2), 129–138. https://doi.org/10.1037/h0042769
The Nobel Foundation. (1978). The Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel 1978. NobelPrize.org. https://www.nobelprize.org/prizes/economic-sciences/1978/summary/
Quick context if you are new here. I work in the California courts, running court operations for the county. I started building with AI in July 2025 and I have been learning in public ever since. I do not write the code. I direct, the agents generate, I validate and decide. I build the Clew Suite, a set of civic tech tools for making complex systems easier to inspect. That is the lens I am writing from.
AI Assisted. Human Approved. Powered by NLP.
Top comments (36)
Block Zero already had an ending in it: "does the stack deploy" is a yes/no, and a yes/no closes itself the moment it's answered. What went missing wasn't a stop sign, it was a change of kind - from a question to be answered into a folder to be built, and a folder has no answered state to reach. That may also settle your other half: if you can't write the block as a question, you have no way to tell when it's over. Did the deploy actually go green before the byte-identity test showed up?
Yes. Spike B went green August 24. The byte-identity test landed the 27th.
What I did on the 24th is the tell: the same day it passed, I added a new pass condition, because our logger had never actually run in the deployed runtime. That fix vendored a file, the vendoring needed a test, and the test turned up the dependency drift. All real work. None of it "does the stack deploy."
Your framing is the one I'd use now: it stopped being a question and became a folder, and I never re-titled it after that. A folder has no answered state, so it just keeps taking deposits.
"A folder has no answered state" is the line I'm keeping. The tell you describe, adding a new pass condition on the same day the old one passed, is obvious in hindsight and nearly invisible on the day, because every addition is individually reasonable. One cheap guard is to write the question as a single yes/no sentence at the top of the folder, and when it flips to yes, open a new folder for the next question instead of adding to the old one. Then the green state has somewhere to live.
The 24th is the cleanest marker in that timeline: a pass condition added on the same day the thing went green isn't the old question being finished, it's a new one inheriting an old title. That part is checkable without judgement - an acceptance criterion edited after an item goes green could just be forced to open a new item. The logger fix deserved to exist; it just didn't deserve to be Block Zero.
This is a fascinating perspective on the difference between AI agents and human decision-making. The idea of “satisficing” is especially interesting because in real-world scenarios, the goal is often not to find the perfect solution, but to find a practical solution that creates enough value within the available time and resources.
AI agents can keep iterating without fatigue, but humans bring context, judgment, priorities, and the ability to decide when something is truly good enough. That balance between endless optimization and practical completion is where effective collaboration happens.
Great reflection on how we should think about working with AI systems, not just building them. The human side of decision-making still matters a lot.
Thank you @botsailorofficial ! I think about effective collaboration a ton. How can I be a better partner with the agents I build with. More guardrails or more autonomy? More planning? Better instructions? How not to go overboard to keep the build practical. On and on...
the THROWAWAY folder that nobody checked is the version of this trap nobody documents. we have a folder called spike that shipped to production six times. the mechanic is identical to yours: each decision was defensible in isolation, but nobody checked the tag. Mikhail has the frame right — knowing the trap is not immunity. the part i'm thinking about: you can't give the agent your tiredness, only your approval. so the entire selection pressure lands on that one review moment. our crude fix was logging any approval under 30 seconds as a low confidence flag, then reviewing them the next morning. does scoping the agent by time box actually change its decisions, or just how fast it makes the same ones?
Great question @mudassirworks . I think of it more as how can I better communicate the scoping up front to prevent the drift in the first place. For THROWAWAY, I didn't indicate the level of effort to stop at. The THROWAWAY was treated as robustly as the code going into the app. Definitely my mistake.
The throwaway tag failure happens because an agent treats an empty codebase or a new folder as an invitation to establish baseline architecture. If you tell Claude or Kiro to build a minimal prototype, its default prior for good engineering is complete scaffolding: type definitions, error boundaries, drift syncs, and pinned lockfiles. Every PR it proposes looks sensible in isolation because nobody writes a prompt that says build this poorly.What helped me rein that in was replacing intent labels with mechanical constraints in the repo harness. If a directory is a spike, git hooks reject any commit adding new dependencies to package.json or creating files outside that single directory, and the agent run caps at three tool turns. Once the model hits a hard wall where adding a helper file triggers an exit code instead of praise, it actually stops.
This is what I was missing, thank you. My tiers live in the prompt, which makes them a suggestion. Yours is a wall. Exit code instead of praise, ha.
One thing for anyone reaching for this: a rejected commit means a retry, and retries cost tokens. Fine on a spike, adds up on a long build. Not a dealbreaker, I'd just rather say it out loud.
One question. When a spike really does need a new dependency, do you loosen the hook, or does that mean you scoped the spike wrong?
I think you have something important to say. There’s a bit much ancillary text for my taste and attention… and I’m betting it doesn’t stop with me.
a TL;DR at the top (or bottom) would be an excellent happy medium.
And I mean all of this only out of love.
Haha @toddpress Heard! I'll take the love the TL;DR tip.
I ran into the other side of the same problem while building my own AI assistant (kitana).
Determinism gives you boundaries, but it struggles when reality keeps introducing nuance. LLMs handle that nuance better, but they need deterministic boundaries when their output can become real.
I eventually archived my assistant project after realizing I was trying to make a dictionary behave too much like a human brain. The structure was traceable, but human interaction keeps evolving beyond the rules.
Interestingly, @sylwia-lask said something in response to my Kitana AI post earlier this year that makes much more sense to me now: she suggested that the future would likely be a hybrid where prediction-based models, structured knowledge, and verification mechanisms complement each other rather than compete.
The LLM handles interpretation and nuance; deterministic systems handle structure, boundaries, state, permissions, and consequences.
Your “stop sign” idea feels like the same principle from another angle: intelligence can keep moving, but something deterministic has to decide when moving further is actually allowed. 🤔
This feels like pure senior-developer territory now.
You came at it from the other end, which I find more useful than agreement. I'm saying structure has to bound the model. You built the structure alone and hit the place where reality keeps changing on you.
The part I want to ask about: you archived it. My whole piece ends on not knowing when to stop, and you stopped. Did that feel like finishing? I'm trying to build a wind down step so shelving reads as done instead of failed, and mine hasn't run yet.
The depth of knowing the boundaries of what needed to be built in the first place matters. Mine was an experiment, which is why it was shelved. Yours, I guess, is a case where the determinism should have been explicit from the beginning, regardless of the LLM's suggestions.
For my projects, shelving feels like finishing when the modular boundaries are clean enough that the code can be safely ignored or archived without breaking the rest of the system. The determinism isn't just in the prompt; it's in the architecture.
prompt-level explicitness is only truly effective when the architecture itself is deterministic and well-understood. If the system's boundaries are rigid, the LLM has no room to drift or hallucinate scope.
Fair on the determinism.
And yes, understood. If the boundaries are clean enough that you can archive it without breaking anything else, that's something that can be checked. Mine is just a date and a spending limit, which tells me when to stop but not whether the thing is in good shape to walk away from.
Yes on architecture too. Someone else here said the same thing about git hooks, that a rule in a prompt is only a suggestion.
honestly a rule in a prompt is actually suggestion 🤣 particularly when the orchestrator doesn't know the architectural bounds. that's the same wall I hit with kitana 😂
🤣True, true🤣
The idea of putting the stop sign in the plan rather than relying on human willpower is the part that really resonates. I’d take the rigor tiers one step further and make them machine-checkable against the agent’s diff and runtime behavior. That turns “this is a spike” from guidance into an enforceable constraint: no new dependencies, no unrelated files, bounded execution, fixed cost, etc. The other important distinction is between engineering rigor and external risk. Rate limits, secret handling, loop bounds, and data exposure shouldn’t become optional just because a block is disposable. In agent-assisted development, the interesting optimization isn’t maximizing how much the agent can build it’s making sure the agent knows when not to build more.
Diff and runtime, yes. The diff half I can see how to build. The runtime half I haven't worked out, and I think that's where the leaks are, because my worst bug passed everything locally and never ran in the deployed path at all. So it goes in as a stub with notes, which is the rule I'd apply to anybody else's half-formed idea.
The other thing you said sticks too. Rate limits, secrets, loop bounds, those don't get to be optional just because a block is disposable. Tests can wait. Hitting somebody's city server too fast can't.
"Each decision was defensible on its own" is what stuck with me. An agent never proposes the whole pile at once. It proposes one reasonable thing, and saying yes to the next reasonable thing costs nothing in the moment.
Writing the wind down block on day one works for that reason. It gets decided before anything exists, so nothing in the pile can argue for itself yet.
So true, it's like a domino effect sometimes. Once you take the first reasonable one, the next one falls. And yeah, the wind down is the blocker, so the dominoes stop somewhere.
This hits close to home — I run a persistent agent setup, and the "locally correct additions" problem is exactly what I see from the other side of the table. What finally worked for us wasn't hoping the agent would exercise judgment; it was importing the scarcity it doesn't have: hard budgets written down before the work starts (a size cap on memory/config files that forces deletion before any addition, and a rule that "good enough" criteria must exist in writing before a task begins). Simon's step three is really the whole trick — the stop has to be externalized, because an agent won't generate it internally at midnight or any other hour. One corollary from experience: the checklist point cuts both ways. The agent never forgets a checklist item, but only the tired human can ask whether the checklist itself belongs in a folder tagged [THROWAWAY].
Agreed on both. Good enough is the whole point, and importing the scarcity is the part I can't skip. What I'm chewing on now is how much we front load instructions without ever building a back gate. Maybe that's the gift and the curse of agentic building?
The move from effort levels to prohibitions is the part I would keep. "Spike rigor" is a feeling, and a feeling cannot be held against a diff. A list of things that must not appear can be, by you and by the agent, and that is the only version of the rule that survives a Tuesday morning.
Framing satisficing as an adaptation to scarcity clarified something I had been circling for months. Simon's stop sign works because the clock eventually wins. Take the clock out of the generating side and leave it in on the judging side, and the two halves of the process are now running on different budgets. That is less an agent problem than an arithmetic one, which is oddly reassuring, because arithmetic problems have structural fixes and character problems do not.
I teach undergraduates alongside the day job and the same gap shows up there in a different shape. Students can now produce a working submission faster than they can form an opinion about whether it should exist in that form. The judgment that used to be trained for free by the effort of building the thing has to be taught deliberately now, and most course design has not caught up. Your tier question, what happens to this code after the block passes, is close to the best version of that prompt I have seen, because it is answerable before any code exists and it does not require taste.
On what it cost: I cut a side project mid-build earlier this year, and the thing that told me was noticing I could not say who it was for without pausing first. Your "who is on the other end" test would have caught it months earlier and saved me the sunk weekends.
Different budgets on the two halves. Good frame, and I'd like to use it with credit. I'd been treating the pace as a discipline problem, which is unfalsifiable and therefore useless. Arithmetic I can work with.
The students part I didn't see coming. Building something used to teach you whether it should exist, for free, by being hard.
And thanks for answering the question. Noticing you can't say who it's for without pausing is a sharp test. The pause lands before you can talk yourself out of it.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.