I am putting this design direction into practice and evaluating it in medium-sized and larger repositories.
Database and infrastructure models, along with verification mechanisms, are already implemented in our tooling. UI/API models and cross-domain composition of migration phases are the next direction.
I currently use a custom harness for development, including database changes. I also use Drizzle in some projects. Across both approaches, I keep coming back to the same friction: a coding agent can produce a locally reasonable change without giving enough attention to the shape of the system as a whole. Accumulating migrations is one place where that becomes visible.
My impression is that newer agents are much better at this than they were a year ago. But the architectural question remains.
How do we make the repository express the system well enough that an agent can understand what a change must preserve?
My target is simple to state:
Make code the best available model of the current specification. Derive intermediate representations from it so changes can be analyzed, rather than judged from text alone.
That applies to UI, API, workflow, database, and infrastructure together. Consider a rule as simple as “after account deactivation, no new billing operation may begin.” Preserving it during a schema change can require agreement across all five domains. I will use that example throughout.
The repository should explain the system as it is
When an agent opens a repository, it should be able to discover the current business concepts, responsibilities, constraints, and dependencies from its structure.
A billing decision hidden inside a generic utility function is harder to understand than a named policy at the billing boundary. A workflow assembled from scattered callbacks is harder to inspect than explicit states and transitions. A permission enforced inconsistently across handlers is harder to reason about than a clear authorization boundary.
The ideal is not to maintain another specification document alongside the implementation. It is to make the implementation's structure carry as much of the specification as possible.
Names, types, module boundaries, contracts, constraints, and transitions all contribute. Comments should explain the reasons that those structures cannot express: why a boundary exists, why an exception is necessary, or which external assumption a decision relies on.
This does not mean code can recover every human intention. It means we should reduce the amount of intention that an agent has to guess.
Current state and change history have different jobs
Drizzle already combines a declarative schema with incremental migrations. Its generate command builds a schema snapshot and compares it with the previous snapshot to generate SQL. Declarative state and migrations are therefore compatible, not competing approaches. See the Drizzle documentation.
The current schema tells us what structure we want. A change also needs to explain how existing data should reach that structure.
Two snapshots cannot fully determine that meaning. A removed column and an added column might represent a rename, a replacement, or a change in the underlying concept. Drizzle’s documented generation flow prompts the developer for renames when necessary. The prompt supplies information the structural diff cannot decide. A backfill similarly requires a rule for interpreting existing rows.
I call this change meaning: information that cannot be recovered from the final state, including rename intent, backfill rules, and cutover conditions. It belongs in the change artifact—the migration or saved plan—while current behavior belongs in the source code.
The repository should make the current specification legible, while the change should retain the meaning that cannot be reconstructed from the final state.
History still matters. It should not be the only way to understand the present.
Declare the system across its boundaries
Declarative code makes important relationships available for inspection. Each domain has different things to declare.
| Domain | What to express | What should become legible |
|---|---|---|
| UI | View states, input rules, available actions | What users can understand and do |
| API | Inputs, outputs, authorization, errors | The contract an operation offers |
| Workflow | States, events, guards, retries, compensation | Allowed execution and recovery paths |
| Database | Schema, constraints, access rules | Stored facts and rejected inconsistencies |
| Infrastructure | Resources, permissions, connections | The execution environment and its boundaries |
These declarations must agree. A disabled button does not enforce authorization. An API response does not prove a queued operation has completed. A database constraint does not establish that an external payment request was never sent.
The architectural value comes from expressing the relationships between these boundaries clearly enough to inspect them.
IR makes the declarations analyzable
Readable source helps an agent understand the system. An intermediate representation, or IR, helps tools compute properties of that system.
For this design, I would derive domain-specific IRs from the source:
- A schema IR for database objects and constraints.
- A configuration IR for infrastructure resources and dependencies.
- A machine IR for workflow states and transitions.
- Contract and view models for API and UI relationships.
I would keep each domain's semantics intact. Database locks, infrastructure replacement, and workflow reachability need different analyses. Flattening them into one universal schema risks losing the details that make those analyses useful.
Cross-domain analysis can follow contracts, references, and dependencies already expressed in code. Where a relationship cannot be derived, it needs an explicit representation close to the relevant code. It does not require a separate specification registry.
An IR is not the state space itself. It is a representation from which we can describe and analyze states, constraints, and transitions.
The business rule needs an explicit expression in code as well. For billing, that might be a named eligibility policy, a guard that invokes it, and a property test defining what deactivation means for pending work. An analyzer can then check whether relevant execution paths use that policy and whether a change invalidates its dependencies. An independent structural oracle checks fidelity to the database; the executable rule and its tests provide a different criterion for business behavior. Neither establishes that the chosen rule matches human intent: its initial definition and substantive changes still require domain judgment.
Terraform provides an existing example of machine-readable change information: plan and state JSON. It also supports validation conditions.
Graphs expose dependencies; DAGs constrain change order
In a database, a schema IR can support a dependency graph: which views, constraints, or other objects depend on which definitions?
A migration planner can also construct a directed acyclic graph of operations. Adding a column may precede a backfill; the backfill may precede validation; validation may precede switching readers.
These are different graphs. Database dependencies can contain cycles, such as two tables with foreign keys referencing each other. An executable operation DAG represents precedence constraints, and a valid ordering still does not prove that the migration preserves business meaning.
The useful progression is:
| Representation | Question it helps answer |
|---|---|
| Declarative source | What should exist, and what rules are visible? |
| Domain IR | What structure and semantics can tools inspect? |
| Dependency graph | What depends on this change? |
| Operation DAG | What must happen before what? |
| Checks and observations | Which properties were actually tested or enforced? |
This gives an agent more than a sequence of SQL statements. It gives the agent a basis for explaining and checking a plan.
The difficult case is the transition
Return to the billing rule: after an account is deactivated, no new billing operation may begin.
For this example, begin means committing a durable authorization to perform a particular billing operation in the database. Deactivation means committing the account's disabled state. An operation authorized before deactivation may still be sent to the payment provider afterward. That is a deliberate rule, not a guarantee that no external request occurs after deactivation.
A guard alone is insufficient. A worker can read an enabled account, another transaction can deactivate it, and the worker can then send a payment request. A normal PostgreSQL Read Committed SELECT does not protect that interval; see the transaction isolation documentation.
A possible protocol serializes both operations on the same account row. In a billing transaction, acquire the row lock, read eligibility, and insert a durable authorization with a unique operation key before committing. The deactivation transaction takes the same lock before changing eligibility. All authorization paths must follow this protocol. A retry with an existing operation key resumes the committed authorization through a separate path, without creating a new one or reapplying the new-authorization eligibility condition. It must not create a fresh authorization after deactivation. External dispatch occurs only for committed authorizations and needs its own retry and idempotency contract.
This identifies the properties an analyzer would need to model:
| Operation | Preconditions | Conditions maintained during execution | Postconditions |
|---|---|---|---|
| Create a new billing authorization | Account lock acquired; account enabled | Lock held through eligibility check and authorization commit | Durable authorization exists for this operation |
| Deactivate account | Same account lock acquired | Lock held through disabling commit | No later authorization may be created while disabled |
| Dispatch payment | Committed authorization exists | Retry identifies the same operation | External result recorded, or outcome remains unknown |
A trace checker can look for an authorization commit after a deactivation commit, while the account remains disabled. It needs records that reliably reconstruct order for that account, not merely application-log timestamps; its finding applies to observed executions, not all possible executions. Static analysis can flag an authorization path that bypasses the shared protocol.
Now migrate the boolean is_deactivated to a representation that separates account status from deactivation time: status and deactivated_at. A true legacy value becomes status = disabled with deactivated_at = null, meaning disabled with an unknown historical time. A false value becomes status = enabled with no deactivation time. New deactivations record a timestamp. Eligibility depends on status; treating a null timestamp as enabled would incorrectly re-enable legacy accounts. These backfill rules belong in the change artifact.
During the transition, an old application may write only is_deactivated while a new billing worker reads only status. Both versions can be individually plausible, yet their combination can bill a deactivated account.
Later, removing the old column may break old jobs still waiting in the queue. Depending on error handling, they might fail safely or interpret missing information incorrectly. Sending a payment request introduces an external side effect that cannot be undone merely by rolling back the schema.
The plan therefore needs to consider application versions, workflow versions, schema phases, and queued work together.
| Phase | Example condition to check |
|---|---|
| Expand | Existing readers and writers remain compatible |
| Backfill and synchronization | Both representations preserve the intended meaning |
| Cutover | New readers receive correct data, including concurrent updates |
| Drain or upgrade old work | Queued jobs retain valid execution semantics |
| Contract | No remaining supported execution requires the removed structure |
The exact mechanism might be dual writes, a trigger, compatibility code, or another protocol. The important point is to make its assumptions and ordering visible.
Existing tools address parts of this problem. pgroll uses versioned schema views and triggers to support concurrent schema versions during PostgreSQL migrations. Temporal Worker Versioning addresses the relationship between workflow executions and worker versions. Atlas migration lint analyzes migration safety. These are useful precedents; connecting the obligations across domains remains the architectural task.
A concrete architecture: domain compilers and a thin linking layer
The architecture follows two paths: compile declarations into domain models, then connect those models through relationships already expressed in code.
Each domain adapter parses and resolves its declarations into a typed IR. IR nodes retain source locations so a diagnostic can point back to the code an agent needs to change. References resolve to program symbols and artifacts.
flowchart TD
U["View and Contract IR: deactivateAccount"] --> D["Deactivate: lock account, update status, commit"]
N["Machine IR: create new authorization"] --> C["Eligibility check under the same account lock"]
C --> A["Schema IR: committed billing authorization"]
D --> S["Schema IR: account status and deactivation time"]
S -->|read under lock| C
A --> P["Machine IR: dispatch committed authorization"]
R["Retry: existing operation key"] -->|resume existing authorization| P
F["Config IR: worker and queue"] -->|hosts| P
P -.-> X["External payment service"]
The graph connects domain-owned models while making the protocol visible: new authorization checks eligibility under the same account lock used by deactivation, and dispatch consumes a committed authorization. A retry resumes that authorization rather than creating another. These edges identify obligations for the analyzers; drawing them does not prove they are enforced. The dotted payment edge requires an external-effect contract.
Reference resolution distinguishes definite edges, conservative sets of possible targets, explicit library contracts, and unresolved references. Possible edges are useful for conservative analysis but must not be presented as confirmed calls. Missing information is not evidence of safety.
The graph supports impact analysis. Temporal correctness needs an additional contract layer, close to the relevant code: transaction boundaries, commit points, synchronization resources, and the ordering of external effects.
| Component | Responsibility |
|---|---|
| Domain adapter | Preserve domain semantics, source locations, and analysis limits |
| Linker | Resolve references and artifact versions |
| Domain analyzer | Evaluate local properties and report the assumptions they require |
| Composer | Connect guarantees to assumptions and check interference between operations |
For billing, the composer must establish that both paths use the same account lock, that eligibility is read under that lock, and that authorization commits before external dispatch. If another path can change eligibility or create authorization outside the protocol, the local guarantee does not compose. Where the available model cannot establish these facts, the result is unknown. Thin reference linking remains thin; concurrency analysis is a separate responsibility.
From IRs to a change plan and runtime evidence
Planning needs three distinct inputs: the candidate model, models of old artifacts still supported, and evidence or deployment contracts identifying what can run or resume. Retain generated IRs and manifests for supported builds, including workflow versions. Deleting an old worker from the current checkout does not erase its dependencies while that artifact can still execute.
Observations remain separate from declarations and carry their origin and acquisition time. These generated models retain the semantics of supported builds.
flowchart TD
P["Plan: operation DAG and phase conditions"] --> V["Verify: compatibility, ordering and gaps"]
V --> G["Gate: bind evidence to plan and recheck prerequisites"]
G --> E["Execute: domain operations"]
E --> O["Observe: actual state and execution traces"]
O --> A["Adjudicate: repair, accept or bound an exception"]
A -.-> P
The composer adds cross-domain precedence and interference conditions to local plans. Before removing the old field, two conditions must hold: no new work requiring it can be admitted, and no accepted work can execute or resume with that dependency. Depending on the runtime, this includes running jobs, delayed jobs, retries, dead-letter replay, and suspended workflows.
One proposed protocol installs a durable version-admission fence, stops old producers, and retires or upgrades every old dependent execution. Scheduling, replay, and resume paths must obey the same fence. The fence remains in force through contract and prevents old artifacts from returning afterward. A gate checks the fence and retirement evidence before removal; rechecking queue depth alone cannot close the race. If the runtime cannot enforce this protocol, the prerequisite remains unknown.
Evidence needs a reproducible verification identity: the exact plan, source revision, input-model digests, and verification conditions, including the analyzer version. A fingerprint over that bundle binds a result to what was actually checked. It does not freeze the environment: observations also carry their acquisition time, and execution gates recheck relevant runtime assumptions. Observations feed verification and adjudication; they do not automatically rewrite the authoritative source.
An agent-facing interface would expose affected source locations, dependency paths, candidate operations, failed conditions, and unknowns. That gives the agent actionable evidence for a code change without asking it to reconstruct the entire architecture from logs.
For example, a declaration change that removes accounts.is_deactivated would remove that field from the candidate Schema IR. Resolving the retained IR of the supported v1 worker against the candidate schema would produce a missing dependency—even if its source was deleted from the checkout. The contract-phase check separately needs evidence for admission fencing and retirement of old dependent executions.
An illustrative diagnostic can be presented as a table. This is a proposed result, not captured output from the harness.
| Proposed operation | Evidence or condition | Analysis result |
|---|---|---|
Remove accounts.is_deactivated
|
Retained billing-worker:v1 model reads the field at src/workers/billing-v1.ts:42
|
Violated: incompatible supported reader |
| Enter contract phase | New admission of old dependent work is fenced | Unknown: enforcement evidence missing |
| Enter contract phase | Accepted old dependent work cannot execute or resume | Unknown: retirement evidence missing |
Execution decision: block removal. The analysis distinguishes violations from missing evidence; the destructive-operation policy determines the decision. An exception would be an explicit policy decision, not a verified result.
Verification needs a scope
A database constraint, a workflow guard, a test, and an alert do different jobs.
A constraint rejects invalid states within its scope. A guard controls operations that pass through it. Tests check the cases they execute. Monitoring detects what its instrumentation can observe.
Calling all of these “verified” hides important differences. A result should state what property was checked, against which evidence, under which assumptions, and for which plan or code version.
An IR also cannot serve as its own sole oracle. If generation and validation share the same lossy transformation, they can agree while omitting the same behavior.
We encountered this in the database harness: a convergence check could report agreement when generation and verification symmetrically dropped information. A separate PostgreSQL-backed structural oracle helped reduce that blind spot. The lesson was specific: agreement after a shared transformation was weaker evidence than it appeared.
We also found declarations that existed in configuration but were not wired into enforcement. A declared policy can give an agent false confidence when its implementation does nothing. Checks for declaration-to-enforcement mismatches, together with deliberately faulty implementations, test whether the mechanism is actually connected. Neither an impressive declaration nor a linked test is enough on its own.
When actual behavior diverges from the declared model, the response requires judgment: restore the intended behavior, accept a corrected design, or record a bounded exception. A production hotfix may be valid, but observed behavior should not silently redefine the specification.
The claim about AI is still a hypothesis
The database and infrastructure models, scoped verification, and plan-bound evidence have concrete implementations in our tooling. The full architecture above—including UI/API models and cross-domain phase composition—is the direction I am working toward. Existing mechanisms and operational lessons do not yet establish a general benefit for AI coding.
My hypothesis is that code with explicit semantics, combined with derived analysis, reduces the context an agent must reconstruct and the assumptions it must invent.
Use matched change tasks with seeded violations, compatibility traps, and cross-domain dependencies. Compare a strong baseline—source, type checking, lint, tests, and existing migration analysis—with the same setup plus derived cross-domain analysis. Keep the model and budget fixed, repeat runs, and use held-out repositories and tasks. Separate changes to declaration structure from added tooling; this evaluates the analysis system, not IR in isolation.
Measure detection, false positives, correct implementations, repair iterations, and maintenance cost. Judge both groups against independent expected outcomes, rather than the analyzer's own finding count.
Counting declarations alone would miss the point. The model must help an agent understand and preserve behavior.
Make the current specification readable in code. Make its structure analyzable through IRs and graphs. Then use that analysis to check both the destination and the path of a change.
That is the design direction I want for large systems built with coding agents.
Top comments (0)