DEV Community

Cover image for Your agent brief and your human runbook are two different documents
Sam Novak
Sam Novak

Posted on Originally published at wagglet.com

Your agent brief and your human runbook are two different documents

More and more tasks on a team board now have two readers:

  1. The coding agent that does the work.
  2. The teammate who starts the agent, watches what it does, and decides if the result is good enough to ship.

Most task descriptions are written for neither of them. They are too short for the agent, and they mix context with instructions, so a person can't scan them quickly during a run.

The usual fix is to write more. A better fix is to write two things:

  • a detailed brief for the agent
  • a short runbook for the person

What each reader needs

The agent: a long, specific brief

An agent doesn't remember last week's standup. It doesn't know which files are fragile or which parts of the code it should leave alone. But it will read every word you give it, so its brief can be long.

Section What it answers
Outcome What should exist or work differently when the task is done?
Relevant context Which files, systems, earlier decisions and links matter?
Out of scope What must not change?
Evidence required Which tests, commands or screenshots count as proof?
Stop and escalate when At what point should it hand back to a human instead of guessing?

The person: a short runbook

The person already knows the project. They need a short list they can keep in their head while the agent works.

Section What it answers
Only you decide Which judgment calls must the agent not make?
Inspect What do you have to check with your own eyes?
Do not proceed when Which red flags mean stop and ask?

If these three answers are buried inside a 600-word agent brief, the person skims, and the one check that was theirs gets missed.


A concrete example

The bug: in a game UI, after crafting an item, the inventory count doesn't update until the player closes and reopens the menu.

Agent brief

## Outcome
The inventory count updates as soon as crafting completes,
without reopening the menu.

## Relevant context
src/ui/inventory/InventoryList.ts binds to `inventory.items`.
Crafting emits `craft:complete` in src/crafting/CraftQueue.ts.

## Out of scope
Save format, item definitions, economy values.

## Evidence required
- `npm test -- inventory` passes
- a new test covers the craft-then-count case
- a before/after screenshot of the menu

## Stop and escalate when
The fix needs changes outside src/ui/ or touches save data.
Enter fullscreen mode Exit fullscreen mode

Human runbook

## Only you decide
Whether the count animation still matches the design.

## Inspect
Craft one item and watch the count.
Try a three-digit stack and the empty state.

## Do not proceed when
The build won't start after two tries,
or the agent asks for anyone else's credentials.
Enter fullscreen mode Exit fullscreen mode

The runbook is under 40 words. That's the point: a person can read it once and remember it for the whole run.

Two things to notice:

  • The runbook has no file paths, event names or test commands. The person can find those in the brief if they need them.
  • The brief doesn't ask "does the animation feel right?" An agent can't answer that. Asking only invites a confident guess.

Four rules that keep the split honest

1. Keep the runbook short on purpose.
If the human section grows past a short paragraph, it's turning into a second agent brief. Move the detail back into the brief.

2. Make "evidence" something you can check.
❌ Tests pass
✅ npm test -- inventory passes, and a new test covers the craft-then-count case
The more precise this line is, the faster the review.

3. Write a stop condition for both readers.

  • The agent's stop condition is about scope: the fix is spreading into areas it was told to leave alone.
  • The person's stop condition is about the run: a broken build, a strange request, or a result that doesn't match the outcome.

4. Never put credentials in the task.
It's tempting to paste a token so a colleague "can just clone it". That token then ends up in the agent's prompt and in the task history. The person running the task should use their own access, and the runbook can say so.


Where a normal tracker gets awkward

You can do all of this today in a plain issue tracker, using a template with two headings. It works, but both halves still live in one text field. When someone copies the task into their agent, the human notes go along with it. Now the agent has instructions that were never meant for it.

That's worth keeping in mind when you choose tools. If you're looking at a Linear alternative for human and AI teams, ask whether it stores the agent's instructions and the human's instructions as separate fields. Then only the agent's half goes into the prompt, and the person gets a clean checklist of their own.


Honest limits

  • It doesn't make a vague task clear. If you don't know what "done" looks like, splitting the task gives you two vague documents.
  • It costs a few minutes per task. That's not worth it for a one-line typo fix.
  • It pays off when: the agent could easily wander out of scope, or a human judgment call decides whether the work is acceptable.

For the full pattern, with a reusable template for each side, see The Dual Prompt: one task for the agent and the human supervising it.

Top comments (0)