DEV Community

Cover image for Spec-Driven Development: What Should Survive After Shipping?
Markus
Markus

Posted on Originally published at the-main-thread.com on AI-assisted

Spec-Driven Development: What Should Survive After Shipping?

Open a repository after six months of spec-driven agent development and you may find another system beside the code. Requirements, research notes, designs, plans, task lists, and review reports all explain what the application is supposed to do.

The code changed last Tuesday. Several of those documents did not.

I understand how teams get there. Coding agents produce an implementation quickly, so we try to settle more decisions before they start. We ask for requirements, a design, and acceptance criteria. Then we make every change follow the same workflow. Before long, a small pull request carries enough generated Markdown to make finding the relevant decision difficult.

I prefer a change brief: a short description of what should change, what must stay stable, and what evidence will let us accept it. Most of that brief can leave the active context after release. The facts that still apply need a maintained home.

Describe the change you intend to make

Existing code shows behavior. Tests, runtime data, and connected systems help explain which parts of that behavior people depend on. A design document can add rationale, but it cannot replace that evidence.

The code also cannot decide what the next version should do. It may show that an account-closing operation deletes a record without explaining an obligation to retain part of the history. It may expose an export endpoint without telling you which partner still uses it.

A change brief brings that outside context into a decision. It should identify the intended outcome, non-goals, affected boundaries, unresolved questions, and acceptance evidence. Add a technical design when the change contains a decision that will be expensive to reverse.

For a small Java application, the brief might fit on one screen:

# Reject blank task titles

Outcome: Creating or editing a task with a blank title returns 422
and leaves the stored task unchanged.

Contract: Validate the title after trimming. Accept 3–100 characters
under the application's existing Java String length convention.
Return an HTML error fragment for the HTMX form.

Non-goals: No change to authentication, storage, or task completion.

Boundaries: POST /tasks, PUT /tasks/{id}, and both HTML forms.

Acceptance evidence:
- Valid titles can be created and edited.
- Blank and too-short submissions return 422 through direct HTTP calls.
- The edit form displays the error and retains the entered value.
- A rejected update does not modify the stored title.

Open question: Confirm the product's title-length convention before
changing it to count Unicode code points or grapheme clusters.
Enter fullscreen mode Exit fullscreen mode

This is an illustrative brief, not a report of tests I ran. Its value is the connection between an observable outcome, a contract, and evidence. “Improve task validation” leaves all three open to interpretation.

Let discovery change the plan

A brief will be incomplete. The agent may find another caller, an unusual transaction boundary, or a test fixture that makes the proposed design unsuitable. A prototype may reveal that the interaction needs a different error response.

Keep the brief open to those findings. Before implementation, inspect the relevant code and identify decisions that need human judgment. During implementation, record deviations that change the outcome or risk. After implementation, compare the actual patch with the intent and run the checks that exercise it.

For the task-title example, discovering that the create form handles a 422 response while the edit form ignores it changes the implementation scope. The endpoint can return the correct status and still fail the user-facing requirement. Reviewing only the server method would miss that.

The brief helps everyone see the missing behavior. It cannot establish that behavior before anyone builds and checks it.

Decide where each surviving fact belongs

At release, read the brief once more. For each statement that will still apply next month, choose its authoritative home.

Fact that must survive Maintained home
Response shape and compatibility API schema and contract tests
Valid task title Validation code and boundary tests
Authorization rule Access policy and allow/deny tests
Module dependency boundary Module structure and architecture checks
Reliability target Service objective, telemetry, and alerts
Release requirement CI and deployment policy
Reason for an expensive trade-off A short, owned decision record

A schema check participates in delivery. An alert reaches an operator. A paragraph buried in an old planning directory usually waits for someone to remember it exists.

Some facts need prose. A business policy, an architectural trade-off, or the reason a rejected approach failed may not fit into code. Keep that explanation close to the artifact it supports, with an owner who can recognize when it becomes stale.

The temporary brief can remain in a ticket or pull request for historical traceability. It does not need to remain an active instruction to every future coding-agent session.

Match the process to the risk

A copy edit and an authorization change need different evidence. Requiring a research report and an architecture review for both consumes attention without making the hard decision easier to see.

For a familiar change, a brief, implementation, focused tests, and ordinary review may be enough. Unfamiliar code may need factual research first. Unclear user experience may need a prototype. A change across an architectural boundary needs explicit alignment. Behavior with serious consequences needs stronger independent evidence and approval.

The team still defines the outcome, boundaries, and acceptance authority. Inside those boundaries, the agent can choose tactics and report what it discovers. When a new finding changes the consequences, it should surface that decision instead of silently expanding its authority.

That is how I would grow the process: add a step because it resolves a specific uncertainty. A standard list of fifteen stages tends to create fifteen artifacts whether anyone needs them or not.

Leave room for the code

Every plan, instruction file, tool definition, and retrieved note competes for the agent's working context. Duplicated requirements and stale explanations also make it harder to find the current source.

Give the agent a small entry point and links to deeper material. Retrieve the relevant schema or module guide when the change reaches that boundary. Read the actual implementation when reviewing the result.

An approved plan is an agreement about intent. Generated code still needs review. The implementation may diverge from the plan, and a coherent patch can contain a product decision nobody approved. Compilation and passing tests only establish what those checks cover.

In the task-title example, a generated test that checks the same validation helper as the endpoint may say little about the browser's handling of a 422 response. The acceptance evidence needs to reach the boundary where the user experiences the behavior.

Modernization needs an extra decision

An older application mixes business policy, published interfaces, platform workarounds, incident fixes, and defects that have survived for years. Translating all of it faithfully can produce a cleaner implementation of the same problems.

Before modernizing a path, classify the observed behavior:

  • Preserve business invariants and externally required behavior.
  • Verify behavior that appears active but lacks clear ownership or evidence.
  • Redesign logic tied to obsolete architectural constraints.
  • Remove confirmed dead paths, duplicated logic, and defects.

Code analysis reveals dependencies. Runtime data shows use. Tests establish a baseline. Business and operational context help decide which behavior belongs in the target system.

The brief records that decision for the change. The new code, contracts, and checks carry it forward after release.

Close the brief when the change ships

My closing review is fairly simple. Compare the implementation with the intended outcome. Check the evidence. Move surviving obligations into maintained artifacts. Leave a short explanation where future engineers need the rationale. Remove superseded planning material from active agent context.

For the task-title change, the surviving pieces are validation, endpoint tests, the two form behaviors, and a documented length convention if that convention needs explanation. The exploratory notes and implementation checklist have finished their job.

This makes a specification easier to trust. It describes a change while the team is deciding and implementing it. After release, the repository contains the behavior and the checks that keep it stable.

What does your team do with a spec after shipping: maintain it, archive it, or keep feeding it to the next agent session?


Adapted from my original Main Thread article with AI assistance for editing. Cover illustration generated with AI.

Top comments (0)