The demo worked. Congratulations. Now what?
We got very good at making AI demos. Give a model a document, explain seventeen things it should do, add “you are a world-class expert”, and watch it produce something convincing. Apparently the expert part is important. Without it, we might accidentally get the intern.
I have written those prompts too. Some worked surprisingly well, which is exactly why it is easy to keep adding things. One more rule, another exception, a little check at the end. Eventually we have half an application written inside a string, and we are still calling it a prompt.
The trouble starts when “it worked on my example” becomes “let's let customers use it”. Now the inputs are messy, information is missing and rules that looked obvious to us compete with everything else we asked for. The answer can still look excellent. It just happens to promise a service the customer has not paid for.
Before improving the wording again, I think we should stop and count what we put in there.
“Just answer the ticket”
Take a normal support prompt. Read the ticket, decide what kind of problem it describes, choose its severity, check the customer's plan, send it to the right team and write a reply. Consult the documentation when available, do not invent a workaround, ask for missing information and put everything into the right fields.
Also be warm. We would not want the unsupported promise to sound unfriendly.
None of these instructions is strange. Together, though, they ask the model to do several different jobs, and we often judge all of them by reading the final reply and deciding whether it sounds right. A good sentence can hide a bad decision remarkably well.
This is why I built Prompt Spider, a small open-source pet project for taking prompts apart. It breaks the text into chunks, highlights different kinds of instruction and helps identify the work hidden inside them. The spider view gives that structure a shape you can explore, instead of another wall of text explaining your original wall of text.
The complex example on the page is that support task. Its breakdown identifies 32 responsibilities: inputs to read, conditions to apply, decisions to make, fields to return and checks to perform.
That does not mean the model executes exactly 32 separate mental operations. Some categories overlap, and another way of counting would produce another number. What matters is seeing how much we have packed into “answer the ticket”. Whether you count 28 or 35, it is clearly doing more than writing a polite paragraph.
Some of this is just an if-statement
One rule says the highest severity is allowed only when the problem affects revenue and the customer is on a qualifying plan. There are two different questions here. Understanding the business impact may need judgement. Checking the plan does not.
The application already knows what the customer bought. It can enforce that restriction directly. Asking the model to remember and apply it while also interpreting the complaint and composing a nice reply gives us another thing to check afterwards. The humble if-statement was available the whole time. It just does not look very exciting in a pitch deck.
Moving the rule into code does not magically make the whole decision correct. The model can still misunderstand the impact. But now we know what to test: whether it understood the ticket, rather than whether it also remembered the subscription policy while writing the third sentence.
The same applies to basic checks. Is the ticket empty? Does the plan exist? Are the required fields present? We do not need an opinion on these things. We need the application to check them.
Other requirements need more than checking a box. If a reply cites documentation, we should be able to find that documentation and see whether it supports the answer. A successful search is evidence that the model looked something up. Anyone who has opened a Stack Overflow page and immediately made things worse knows that looking something up is only the beginning.
And when the information is missing, let the answer say so. Requiring a field is fine; requiring an actual citation without giving the model a source is where we create trouble. An empty list of sources is less impressive than three convincing titles, but at least nobody has to spend the afternoon discovering that the third paper exists exclusively in the reply.
Not every prompt needs a committee
The easy example in Prompt Spider asks for a LinkedIn post with a few rules. One call is perfectly reasonable. We do not need a writer agent, an editor agent and a senior emoji compliance officer to announce a feature.
The overloaded example goes in the other direction. It asks for a 1,200-word article, Italian and Spanish translations, five tweets, a LinkedIn post, a newsletter introduction, a summary, keywords, image prompts and a grammar review. Everything must fit under 2,000 words.
The full article and its two translations already leave that budget in trouble. Before the model has made a mistake, we have given it a request that cannot reasonably fit as described. Something will be shortened, skipped or quietly reinterpreted. Then we will complain that it did not follow instructions.
It also asks for three studies with exact figures, without providing sources. That deserves attention. “Playful but suitable for enterprise buyers” is a normal writing request. “No emojis except in the tweets” is a clear exception. A useful review should distinguish those from missing evidence and a word budget that does not add up.
This is also where the tool needs a little humility. Its English word-and-pattern checks can miss meaning. The model-assisted review, which requires a Claude API key in the Model tab, adds another assessment, not privileged access to what a model will actually do. Its risk indicators point to things worth checking; they are not measured odds of failure.
You still have to run the prompt on realistic examples. Preferably including some you did not choose because they make the demo look good.
Split where it helps
Once we can see the jobs inside a prompt, we can decide which ones belong together. Perhaps classifying a ticket and writing its reply work well in one call. Perhaps we need to check the classification before allowing the reply to be written. That is a decision we can test.
Separate calls cost time and money, and mistakes can travel between them. Adding more parts does not automatically improve anything. I would split when it gives me a concrete benefit: a rule I can enforce, a decision I can inspect, or a failed step I can retry without doing everything again.
The useful change is that we stop treating the final answer as the only thing worth examining. We know which decisions produced it and which checks stand between that answer and a customer. We can change a policy without hoping that another paragraph of instructions will do the job.
Count before adding another IMPORTANT
We spend a lot of time asking how to make models follow instructions better. Fair enough. But we should also look at the instructions we keep handing them. Sometimes the prompt is unclear. Sometimes it asks for information we never supplied. Sometimes it is doing work that a few lines of code could handle more reliably.
Open Prompt Spider and try the prompt you are most proud of. Inspect the breakdown, disagree with it where necessary, and look for the decisions hiding between the writing instructions.
Before adding another IMPORTANT in capital letters, ask how many jobs that prompt already has, and why the model got all of them.
Top comments (0)