DEV Community

qnbs
qnbs

Posted on Fully Autonomous

The Engineering Episode: A Work Contract for Coding Agents

“Upgrade the authentication library safely” sounds like a task. For an agent, it is a bundle of decisions: which package, what counts as safe, which files may change, whether a lockfile update is allowed, what tests matter, and who may approve a release.

If those decisions stay implicit, the agent has to guess. A better unit of delegation is a bounded engineering episode: a task described by the result it should produce, the boundaries it must preserve, the evidence it must return, and the point where it must stop.

That small change in framing makes a request easier to delegate, review, and repeat. It also keeps the agent’s freedom over implementation separate from the human’s authority over outcomes and risk.

A task is more than an instruction

An instruction describes an action: “change this function,” “update the dependency,” or “make the build pass.” A work contract describes a desired state and the conditions under which that state is acceptable.

For an engineering episode, write down six things:

  1. Outcome: the externally meaningful state you want.
  2. Scope: the parts of the system the agent may inspect or change.
  3. Invariants: what must remain true, even if the implementation changes.
  4. Authority: which tools and consequential actions are allowed.
  5. Evidence: what observations would support the completion claim.
  6. Exit: when to stop, report a limitation, or ask a person to decide.

This is not ceremony for its own sake. It moves ambiguity to the beginning, when a person can resolve it, instead of letting it surface after a broad patch has already been produced.

OpenAI’s practical agent-building guide also recommends beginning with a specific workflow and planning for human intervention when retries are exhausted or an action is high-risk. The episode card below is an editorial synthesis, not a vendor-prescribed schema.

Start with the result, then bound the path

Consider a hypothetical task: update a client library to a supported version while preserving the current API behavior.

“Upgrade the library and fix anything that breaks” leaves a large decision space. Does “anything” include unrelated tests? Is a major-version migration acceptable? Can configuration or credentials change? May the agent update the lockfile? Is a successful local build enough to call the work done?

A tighter contract might read:

Outcome: Move the application from client-library version 4.x to the approved 5.x release while preserving the documented login and token-refresh behavior.

Scope: Inspect the dependency manifest, lockfile, client adapter, and related tests. Change those files and add or update focused tests. Do not reformat unrelated code.

Invariants: No credential values in source control. No change to public method names or token storage. No network call to a production identity provider.

Authority: Read the repository and documentation; edit the current working branch; run local tests. Do not merge, publish a package, modify secrets, or deploy.

Evidence: Report the dependency diff, focused test results, full relevant test command and exit status, and any behavior that could not be exercised locally.

Stop: Stop before changing public behavior, encountering an undocumented migration choice, needing a secret, or repeating the same failing check without new information.

The example is hypothetical, but the structure is reusable. Notice that it grants freedom over how to adapt the client while reserving high-impact decisions for a person.

Make constraints observable

“Keep it secure,” “don’t break anything,” and “use best practices” express intent but do not tell an evaluator what to inspect. Where possible, translate them into observable conditions:

Vague instruction More reviewable condition
Keep the API compatible Existing public method names and documented request/response shapes remain unchanged
Do not weaken authentication Existing authorization tests still pass; no credential is added to tracked files
Keep the change small Only named dependency, adapter, and test paths change unless a new path is explained first
Verify the update List the exact commands run, their exit codes, and checks skipped

Not every value can be reduced to a test. The point is to separate what can be checked mechanically from what needs human judgment. An agent should not quietly convert an unmeasured preference into a claim that it has been satisfied.

Scope is an access boundary

The scope line should match the actual environment, not merely appear in the prompt. “Do not deploy” is weak if the agent still holds deployment credentials and an unrestricted release tool. A contract becomes credible when the available permissions enforce its boundaries.

For a low-risk documentation edit, read access plus a branch-local write may be enough. For a production change, the agent may prepare a patch and evidence while a separate release process retains deployment authority. The work contract should name which boundary is enforced by tooling and which is only an instruction.

Write exclusions as carefully as inclusions. “Update the adapter and tests” does not imply permission to rotate credentials, change account policy, or publish the package. If an adjacent issue appears, the agent can report it as a follow-up rather than expanding the task on its own.

Define done and define not-done

A useful completion statement is falsifiable. “The task is complete” is not evidence. “The adapter and focused tests changed; command X exited zero on revision Y; integration behavior against the live provider was not tested” gives a reviewer something concrete to assess.

The contract should also make room for honest partial completion:

  • Complete: the requested state was reached and the named evidence exists.
  • Blocked: a required decision, permission, dependency, or external state is missing.
  • Incomplete: the agent made progress but one or more acceptance conditions remain unmet.
  • Escalated: the work exposed a risk or choice outside the delegated authority.

These states prevent the pressure to produce a polished “done” message from turning an unknown into a pass.

A compact episode card

For routine work, the contract can fit in a short issue comment or task form:

Field Prompt
Outcome What observable state should exist at the end?
Scope Which repositories, files, environments, or records are in bounds?
Invariants What must not change?
Allowed actions What may the agent read, write, run, or submit?
Evidence What should the completion report include?
Stop and escalate What ambiguity, risk, or repeated failure ends the episode?
Owner Which human owns decisions and accepts the result?

If the task cannot be described this way without major uncertainty, that is useful information: it may need discovery or a human design decision before implementation begins.

The practical test

Before handing work to an agent, ask:

  1. Could two reasonable people interpret the requested outcome differently?
  2. Can the agent affect anything outside the named scope?
  3. What result would prove progress, and what result would only look reassuring?
  4. Which decisions are reversible, and which need a person first?
  5. What should happen when the available evidence is incomplete?

The strongest work contract is not the longest prompt. It is the shortest description that makes the desired result, permitted actions, evidence, and exit conditions clear enough to review.

An agent can improvise the route. The contract defines the destination, the guardrails, and the point where a human must take the wheel.

References

Source, license, and AI assistance

This article is based on the engineering-episode concept in Part II of From Vibe Coding to Agentic Software Engineering, whose source record credits ChatGPT as preparer and identifies CC BY-NC-SA 4.0. This version is substantially reorganized and expanded with original templates and examples, and is shared under the same license: CC BY-NC-SA 4.0.

AI disclosure: The article text was generated primarily by AI. A human publisher supplied the topic, source material, and editorial direction, and remains responsible for checking claims and examples before publication.

Top comments (0)