The team meets. Everyone is excited. Someone opens a chat window and types "build us a dashboard that tracks our renewals." Twenty minutes later there is a dashboard. It has charts. It has a dark mode. It looks amazing.
Then the person who actually handles renewals every month looks at it and says, "That number is wrong."
Nobody can say why and nobody wrote down what right was supposed to be.
Why this step matters so much
A.I. is really good at making things that look finished. That's the magic, and it's also the trap.
What it can't do is know what "right" means for your work. It doesn't know that a charge on the 3rd is normal but a charge on the 30th means someone forgot to cancel. It doesn't know which column in the export can't be trusted. You know that, or someone on your team does.
That knowledge is the most valuable thing your team brings. Getting it out of people's heads and onto the page before anything gets built is where most of the win comes from.
What it looks like in practice
Say the project is tracking software subscriptions. You have a list of what your team pays for, and a card statement of what actually got charged. The goal is to catch surprises.
Before anybody prompts anything, write three examples by hand.
A normal one. "Design tool, $15 a month, charged on the 5th, matches the list. Nothing to flag."
A wrong one. "Survey tool. The list says $20 a month. The card says $340. It renewed annually at a new price. Flag it."
A weird one. "A $9 charge from a vendor nobody recognizes. Not on the list at all. Flag it and find out who owns it."
Three examples, ten minutes. Now you have something to hand the A.I., and something to check its work against. When the tool says "everything looks good," you can run your three and know whether that's true.
The weird one is usually the most useful. It's where the real work lives, and it's the case the A.I. may quietly skip.
Tips for getting started
1. Start with a small slice of real data. Twenty rows beats twenty thousand. You can read twenty rows yourself and spot what the A.I. gets wrong. Stick to data and tools you're allowed to use, and leave out anything sensitive you don't need.
2. Write your examples before you prompt. One normal, one wrong, one weird. If you have issues here, you may need to spend more time sorting your data and approach. That's a great thing to find out on day one instead of week three.
3. Decide what "wrong" looks like. One sentence is enough: "This is wrong if a charge shows up that isn't on the list." I wrote more about this in Checking A.I. Work is Hard. If you can't name what broken looks like, you'll end up with misleading results.
4. Pick the smallest version that helps someone next week. Skip the full vision for now. Find the one piece that would save somebody real time on Monday. A spreadsheet of flagged charges is a win. The dashboard can come afterwards.
5. Let the person who knows the work be the judge. On a mixed team, the most important person in week one often isn't the one writing code. It's the person who can look at an output and say "that's wrong, and here's why." If that's you, everyone on the team will need you.
6. Make the A.I. ask questions first. Give it your sample and your examples, and tell it not to build anything yet. Let it interview you. The questions it asks will show you what your team hasn't decided.
Tips particularly for developers
Turn the examples into tests. Your team's hand-written cases can become your first test fixtures. Every change then gets checked against what the team said right looks like.
Don't let the agent edit the examples. Keep them in their own file and mark it read only. If a test fails, the code is likely wrong, not the example. Agents love to "fix" a failing test by changing what it expects.
Add an example every time something surprises you. Every bug the domain expert catches becomes a new case. By the end of the project you have a real test suite written by the people who know the work best.
Prove the logic before the interface. A small script that reads the sample and prints the flags is a great first deliverable. Once the flags are right, the UI is the easy part.
Wrapping up
The tools keep getting faster at building. Knowing what right looks like is still on us, and it's the part that makes everything else worth building.
So before your team opens a chat window this week, grab a whiteboard, pull twenty rows, and write down three examples. It takes 10 minutes, and will be very worth it.
Good luck to everyone kicking off something new.
Example prompts
Before you start
Here is a small sample of our data and three examples of what correct
output looks like, including one that should be flagged. Do not build
anything yet. Read everything, then ask me the questions you need
answered to get this right. Ask them one at a time.
Here is what we want to build: <one or two sentences>. What is the
smallest version of this that would still be useful to someone next
week? What would we leave out, and what would we lose by leaving it out?
Writing your examples
Here is a sample of our data. Help me write examples of correct output:
one normal case, one that is clearly wrong, and one edge case that is
easy to miss. Don't guess at our rules. If you are not sure how a case
should be handled, ask me.
Before it builds
Before you build anything, restate in plain language what "done" means
for this project, and list how each of my examples should come out. If
anything I've told you conflicts, point it out.
After it says it’s done
Run my examples through what you built and show me the result for each
one next to what I said it should be. Do not change the examples. If
any don't match, tell me which ones and why.
What kinds of cases did my examples NOT cover? Give me three new
examples I should add, and tell me how you currently handle each one.
For developers
Turn examples.md into test fixtures. Treat examples.md as read only.
If a test fails, fix the code, never the expected result. If you think
an example itself is wrong, tell me instead of changing it.


Top comments (0)