`"Prompt Engineering Examples"
Prompt Library · Part 1 of 2 · Updated August 2026
The first prompt I ever "engineered" was for a product description. I typed "write a product description that converts," and Claude handed back four sentences so generic they could've described a candle, a CRM, or a kayak. Nothing was wrong with the words. The structure simply wasn't there to catch a specific outcome.
That's the gap this post is built to close. Not with adjectives ("write something amazing") but with the actual mechanics: what changes in a prompt's structure, and why the model responds differently when you change it. Below are 48 paired examples — a generic version next to the engineered rewrite — across five categories, plus the research behind the patterns that hold up across models.
Quick answer
The fix is almost never vocabulary. It's naming the audience, capping the scope, and locking the output format — five structural moves cover most of what "prompt engineering" actually means in practice.
Examples teach format more than content. One well-formatted example usually beats three sloppy ones.
Where you put information changes whether the model uses it — critical instructions belong at the start or the end, never buried in paragraph four.
The discipline itself is shifting. By mid-2026, most practitioners had folded "prompt engineering" into the broader work of context engineering — curating what the model sees, not just how you phrase the ask. The patterns below still hold; several matter more at that scale, not less.
One honest caveat
2026 is an awkward year to write this kind of article, and I'd rather say so than pretend otherwise. Andrej Karpathy's framing of the model as a CPU and the context window as RAM picked up real traction through 2025 and into this year, and a growing share of practitioners now treat prompt wording as one input into a larger context-management workflow rather than the main lever.
A February 2026 study out of HxAI Australia ran 9,649 experiments across 11 models and four context formats and found something that should temper any prompt-format advice, including some of what's below: format choice had no statistically significant effect on aggregate accuracy, while the gap between frontier and open-source model capability was 21 percentage points — by far the largest factor the study measured. Read plainly, that means which model you're using will usually matter more than how cleverly you format the context around it. The patterns in this article are still worth knowing. Just don't expect prompt polish to out-run a capability gap.
Table of Contents
Writing & Content (10)
Code & Engineering (11)
Data & Spreadsheets (9)
Support & Sales (9)
Marketing & SEO (9)
FAQ
Note: This is Part 1: 48 examples across the five categories above. Part 2 — research & decisions, image/video/design, agentic workflows, education, and legal/HR/admin, plus a section on prompts that failed and why — is in production now. We split it rather than publish all eleven categories as one long scroll, because a single post with 111 examples stops being a reference and starts being a wall.
The five mechanics that show up everywhere
Four things show up over and over in the 48 rewrites below, and all four are backed by published research rather than vibes:
- Role + Stakes "You are reviewing this before it ships to 40,000 subscribers."
- Context First Background, data, constraints — placed before the task.
- One Task, One Sentence Not three goals stacked into one ask.
- One Worked Example Shows format and judgment, not just instructions.
- Output Lock Format, length, and what to do if information is missing. Order matters: items 1 and 2 anchor the start, item 5 anchors the end — see "lost in the middle" below. What the research actually says
- Where you place information changes whether the model uses it. Stanford and Google researchers tested how language models handle long inputs and found a consistent U-shaped curve: accuracy is highest when the relevant fact sits at the very start or very end of the context, and drops measurably when it's buried in the middle — even in models built for long contexts. Anthropic's own current documentation notes that newer Claude models have meaningfully improved on this, but the safe habit — critical instructions first, critical instructions last, never buried in paragraph four — still costs nothing and still helps.
- Examples teach format as much as content. A widely-cited 2022 study from the University of Washington and Meta AI found something genuinely counterintuitive: replacing the correct labels in few-shot examples with random ones barely hurt performance, because the model was leaning on the example's structure and label space more than the literal correctness of each one. Practically, this means a single well-formatted example often does more work than three sloppy ones — which is why so many rewrites below include exactly one.
- Structure beats vocabulary, and role-play preambles are weaker than they used to be. This is the part most "100 prompts" lists skip. Phil Schmid at Hugging Face has been blunt about it: most production agent failures aren't bad prompts, they're context failures — the wrong documents retrieved, too much history stuffed into the window, missing tool definitions. Several practitioners who build production systems now report that "You are an expert…" identity priming barely moves accuracy on frontier models and mostly spends tokens you'd rather budget elsewhere — which is why the "role + stakes" pattern in this article leans on concrete stakes (a real audience, a real deadline) rather than a job title alone.
- Verification is part of the prompt, not an afterthought. The rewrites that ask the model to check its own output against a stated constraint before finishing consistently produce fewer silent errors than ones that just ask for the output. This isn't from a paper — it's the single most reliable lever I've found in three years of doing this for a living, and it's the cheapest one to add. One measured result, for scale: Microsoft Research ran a controlled trial where developers wrote a JavaScript HTTP server with and without GitHub Copilot. The group with Copilot finished 55.8% faster, with a 95% confidence interval of 21% to 89% — a real, replicated effect, but a wide enough range that "AI makes you faster" and "AI makes you twice as fast" are both defensible readings of the same study, depending which end of the interval you land near. That's the kind of nuance most "AI productivity" headlines strip out, and exactly why the rule on this site is: no stat without the conditions attached. How to read the examples Table Format Each card shows the generic prompt, the engineered rewrite, and the one mechanic that changed. Portability Most examples work on Claude, GPT-5, and Gemini with no edits. A few model-specific notes are flagged. Use them Copy, swap the bracketed details for your own, and keep the structure intact. Symptom → Fix Map If your prompt is producing these problems, here's the fix: Table Symptom Fix Output is generic, could describe anything Specificity lock + role & stakes Model ignores half your instructions Decomposition (one ask at a time) Format changes every time you re-run it One worked example + output schema Sounds confident, factually wrong Verification pass + source grounding Works once, breaks on the next input Edge-case forcing + few-shot anchor Keep that table nearby. Almost every "after" prompt in the 48 below is one of these five fixes applied to a specific situation. Writing & Content 10 examples. Most "make it better" requests fail for one of two reasons: the model doesn't know who's reading, or it doesn't know what to leave out. Every example here fixes one of those two things. Example 1: Specificity lock + negative constraint Generic: plain Write an engaging introduction for a blog post about remote work productivity. Engineered: plain You're writing the opening 80 words of a blog post for mid-level managers who just inherited a fully remote team and don't trust it yet.
Context: the post argues remote teams underperform not because of distance,
but because managers replace trust with monitoring — which backfires.
Open with a specific, concrete scene of a manager doing this. Not a statistic.
Not a question. End the paragraph on the tension, not the resolution.
No "in today's remote-first world."
Why it works: Naming the three openers the model defaults to (statistic, question, that exact cliché phrase) rules them out before it writes a single word, instead of you discovering and rejecting them after.
Example 2: Audience lock
Generic:
plain
Write a product description that converts for a ceramic pour-over coffee dripper.
Engineered:
plain
Write a 60-word product description for a ceramic pour-over dripper, $38,
aimed at people who already own a $200+ espresso machine and are buying
this as a slower, deliberate alternative.
Lead with the specific friction this solves — not "great taste," they already
have that at home. One plain sentence on what it's not good for.
End with the actual brew time.
Why it works: Naming who already owns an espresso machine rules out the generic "rich, full-bodied flavor" copy that fits any coffee product ever sold, and forcing one honest tradeoff in instead of a disclaimer reads as confidence, not weakness.
Example 3: Constraint budget + self-critique loop
Generic:
plain
Give me 10 subject lines for an email about our Black Friday sale.
Engineered:
plain
Generate 8 subject lines for a Black Friday email to past customers who
bought once and never returned.
Under 45 characters. No emoji, no "last chance," no exclamation points.
At least 2 should reference that they haven't been back, framed as curiosity
rather than guilt.
Mark your top 2 picks and say why, one line each.
Why it works: Banning the four phrases every inbox is already full of removes the model's safest, laziest defaults, and asking it to rank its own output hands you a decision instead of 8 near-duplicates you still have to judge.
Example 4: Role + stakes · scope fence
Generic:
plain
Make this more professional: [paste text]
Engineered:
plain
Rewrite this for a CFO audience: shorter sentences, no hedging ("I think,"
"maybe," "it seems"), lead with the number that matters, cut anything that
isn't a decision or a risk.
[paste text]
If a sentence doesn't change a decision the CFO will make this week,
delete it rather than rephrase it.
Why it works: "More professional" is a tone request with no real target; naming the reader's actual job this week gives the model a concrete filter instead of its default — longer words, more formal, equally vague.
Example 5: Decomposition
Generic:
plain
Write a 1,500-word article about why people quit their jobs in 2026.
Engineered:
plain
Before writing anything, give me a 6-point outline for a 1,500-word article
on why people are quitting jobs in 2026. One sentence per point stating the
argument, one sentence naming what evidence would support it (a study, a labor
statistic, a named example).
Stop after the outline. I'll tell you which points to keep.
Why it works: Splitting "write the article" into "plan, then check, then write" catches a weak or unsupported argument while it's six bullet points, not after 1,500 words are built on top of it.
Example 6: Negative constraint + verification pass
Generic:
plain
Turn these meeting notes into a paragraph.
Engineered:
plain
Turn the bullet notes below into 3 short paragraphs for a non-technical
stakeholder update. Preserve every number exactly as written — don't round,
estimate, or soften a figure into "significant" or "modest." If a bullet
is ambiguous, flag it in brackets rather than guessing its meaning.
[notes]
Why it works: Forbidding number-softening in writing is the single fix for the most common way AI rewrites quietly introduce factual drift into a status update nobody re-checks against the source.
Example 7: Scope fence + verification pass
Generic:
plain
Make this shorter.
Engineered:
plain
Cut this to 200 words by removing redundant sentences and weak qualifiers —
not by summarizing the ideas into vaguer versions of themselves. If two
sentences make the same point, delete one entirely rather than blending them.
List what you cut, separately, below the rewrite.
Why it works: "Shorter" alone usually produces a thinner, vaguer draft; specifying that cutting means deletion, not compression, and asking for a visible list of cuts, makes the editing decisions checkable instead of invisible.
Example 8: Persona transfer + audience lock
Generic:
plain
Write a short bio for my About page.
Engineered:
plain
Write a 90-word About page bio. Audience: potential freelance clients deciding
whether to email me, not peers in my industry.
Facts to use: 10 years in UX research, worked at two startups that got acquired,
now solo, based in Lisbon.
Lead with what a client actually cares about — can I trust this person with
a real project — not a chronological job list. One sentence of personality at
the end, not a hobby list.
Why it works: Naming who's reading (a buyer, not a peer) changes which facts get foregrounded; without it, bios default to CV order, which quietly answers the wrong question.
Example 9: Audience lock + failure mode naming
Generic:
plain
Simplify this paragraph for a general audience.
Engineered:
plain
Rewrite this for someone with no background in the topic, same length or shorter.
Replace every term a non-specialist wouldn't recognize — but with the concrete
thing it means, not a vaguer word. If a sentence has no accurate plain-language
equivalent, keep the term and define it in 6 words or fewer in parentheses.
[paragraph]
Why it works: Naming the actual failure mode — "don't replace it with a vaguer word" — heads off the most common way AI "simplifies" text: by making it less precise instead of less technical.
Example 10: Self-critique loop + audience lock
Generic:
plain
Give me 5 headline options for this article.
Engineered:
plain
Generate 6 headline options for the article below. For each, name the specific
reader situation it's written for (not a demographic — e.g. "someone who just
got passed over for a promotion").
Then pick the one you'd actually run if this were your own newsletter, and say
what you'd A/B test against it.
[article summary]
Why it works: Asking for a reader situation instead of a demographic filters out the generic "10 Tips" instinct, and forcing a single committed choice produces a decision instead of a menu you still have to make yourself.
Code & Engineering
11 examples. In Microsoft Research's randomized trial, developers using GitHub Copilot finished a standardized task 55.8% faster, 95% CI 21–89% than the control group. The prompts below are the layer on top of that: the difference between an AI coding assistant that's a faster autocomplete and one that's an actual collaborator.
Example 11: Edge-case forcing + scope fence
Generic:
plain
Write a function that validates email addresses.
Engineered:
plain
Write a Python function that validates email addresses for a signup form.
Before writing it, list 5 edge cases you'll handle (plus-addressing, subdomains,
missing TLD, etc.) and 2 you'll explicitly NOT handle, and why. Then write the
function with a docstring noting those decisions. Standard library only, no new
dependencies.
Why it works: Naming edge cases before writing code surfaces the validation gaps that normally show up as a bug report three weeks later, and capping dependencies stops a 6-line function from becoming a regex-library import nobody asked for.
Example 12: Reasoning trace + scope fence
Generic:
plain
Fix this bug: [code + error]
Engineered:
plain
This function throws [error] on this input: [code]. Before changing anything,
state your hypothesis for the root cause in one sentence, and name one alternative
you're ruling out and why. Then make the minimal fix — don't refactor surrounding
code unrelated to the bug.
Why it works: Asking for a stated, ruled-out alternative catches the common failure where a model patches the symptom at the call site instead of the actual cause two functions upstream.
Example 13: Scope fence + negative constraint
Generic:
plain
Review my code.
Engineered:
plain
Review this code for one thing only: data races in the shared state between
these two functions. Don't comment on naming, style, or anything unrelated to
concurrency. If you find none, say so directly — don't pad the response with
minor style notes to look thorough.
[code]
Why it works: An unscoped review produces twenty minor nitpicks and buries the one finding that matters; this is also the exact failure mode Anthropic's own current prompting documentation flags when it notes newer Claude models can over-deliver unless the scope is capped explicitly.
Example 14: Scope fence
Generic:
plain
Refactor this to be cleaner.
Engineered:
plain
Refactor this function for readability only. Don't change its behavior, its public
interface, or add error handling for cases that can't currently occur. Don't
introduce a new abstraction unless it's already reused at least twice in this file.
Why it works: "Cleaner" with no boundary tends to produce more flexible, more abstracted code than the task needed — stating the limit up front heads it off instead of catching it in review.
Example 15: Edge-case forcing + output schema
Generic:
plain
Write tests for this function.
Engineered:
plain
Write unit tests covering: the happy path, an empty input, a boundary value at
the function's stated limit, and one case that should raise an exception. Use pytest.
One assertion focus per test — no test checking five unrelated things at once.
Why it works: Naming the four cases up front stops the model defaulting to three near-identical happy-path tests that all pass and tell you nothing about the boundary that actually breaks in production.
Example 16: Audience lock
Generic:
plain
Explain this code.
Engineered:
plain
Explain this code to a developer who knows Python but has never seen this codebase.
Assume they understand the language, not our domain. Walk through what it does in
execution order, not file order, and flag the one part that isn't obvious from
reading it line by line.
Why it works: With no audience specified you get either a useless line-by-line restatement or a one-paragraph summary with no traction; naming the actual reader fixes the altitude of the explanation.
Example 17: Scope fence + verification pass
Generic:
plain
Convert this to TypeScript.
Engineered:
plain
Convert this JavaScript function to TypeScript. Preserve the exact runtime behavior,
including existing error handling — don't "improve" it with stricter checks that
change what inputs are accepted. Add types only; flag any place the original behavior
is ambiguous enough that you had to guess a type, instead of guessing silently.
Why it works: Language conversions are where models most often slip in unrequested behavior changes disguised as type safety; asking for flagged ambiguity instead of silent resolution keeps the actual decision with you.
Example 18: Specificity lock
Generic:
plain
Optimize this function.
Engineered:
plain
This runs on ~50,000 rows once per day, not in a hot path. Optimize for readability
and maintainability, not raw speed — don't introduce caching, memoization, or
algorithmic complexity unjustified at this scale. If a genuinely faster approach
exists, note it in a comment, don't implement it unless asked.
Why it works: "Optimize" is undefined until you say what for; without scale and frequency stated, models default to the most impressive-looking optimization, which is often the wrong tradeoff for a once-a-day batch job.
Example 19: Few-shot anchor + verification pass
Generic:
plain
Write a regex to match phone numbers.
Engineered:
plain
Write a regex matching US phone numbers in these formats: (555) 123-4567,
555-123-4567, 5551234567. Then give me 5 strings that should match and 3 that
look similar but shouldn't, so I can verify it against real input before using it.
Why it works: The exact formats act as anchoring examples, and asking for both true and near-miss test strings turns a regex you'd otherwise debug against production data into one you can verify in ten seconds.
Example 20: Reasoning trace + decomposition
Generic:
plain
Why does this test fail?
Engineered:
plain
This test fails intermittently, not every run: [test + code]. List your top 2
hypotheses for why it's flaky, ranked by likelihood, with the evidence in the code
supporting each. Don't propose a fix yet — I want the cause first.
Why it works: Separating diagnosis from fix matters most exactly when a bug is intermittent, because the instinct to "just fix it" usually means a sleep() or retry that masks a race condition instead of resolving it.
Example 21: Scope fence + output schema
Generic:
plain
Add documentation to this code.
Engineered:
plain
Add docstrings only to public functions, not private helpers prefixed with underscore.
Each docstring: one line on what it does, params, return type, one line on what it
raises and when. No inline comments explaining obvious lines.
Why it works: Capping documentation to the public surface and naming the exact fields prevents the common over-delivery where every line gets a comment and the file becomes harder to scan, not easier.
Data & Spreadsheets
9 examples. A model that's never seen your spreadsheet will still confidently write you a formula for it. The fix isn't a smarter model — it's telling it which version of Excel-flavored truth you're working in before it guesses.
Example 22: Output schema + verification pass
Generic:
plain
Write me a formula to calculate commission.
Engineered:
plain
Excel formula: commission = 8% of sales in column D, but 12% on the portion
above $10,000, for each row. Sales are in D2:D500, some cells are blank
(treat as 0, not error). Give me the formula for one cell plus a one-line
explanation of how it handles the blank-cell case, since that's where my last
version broke.
Why it works: "Calculate commission" has no fixed meaning — tiered or flat, blanks as zero or error, single rate or marginal — and naming the exact tier structure plus the blank-cell rule is the difference between a formula that works on row 2 and one that works on all 499 rows.
Example 23: Audience lock + scope fence
Generic:
plain
Summarize this sales data.
Engineered:
plain
Summarize this quarter's sales data for a regional manager who already knows
the numbers roughly — she wants what changed and why, not a recap. Three bullets
max: the biggest swing, one number that looks fine but isn't (explain why), and
one thing you can't explain from this data alone. Don't restate totals she already
has in the report.
[data]
Why it works: "Summarize" with no audience defaults to restating the table in sentences; naming what the reader already knows forces the model to surface the delta instead of the dataset.
Example 24: Edge-case forcing + decomposition
Generic:
plain
Clean up this messy data.
Engineered:
plain
Before changing anything, list every issue you see in this data: duplicate rows,
inconsistent date formats, trailing whitespace, mixed casing in the category column,
and anything else. Then propose a fix for each as a numbered rule I can approve or
reject — don't apply any fix until I've seen the list. Flag any row you'd delete
rather than fix, separately, since deletions aren't reversible.
[data]
Why it works: "Clean up" with no approval step means silent deletions you only discover when a report comes up short; separating the diagnosis from the fix turns an irreversible action into one you sign off on first.
Example 25: Comparison frame
Generic:
plain
What chart should I use for this data?
Engineered:
plain
I have monthly churn rate for 5 customer segments over 18 months. I'm deciding
between a multi-line chart and a small-multiples grid (one mini chart per segment).
Give me the actual tradeoff for this specific data — not chart theory in general —
and which you'd pick if the audience is execs skimming on a phone.
Why it works: Chart-type questions almost always get a generic "bar charts for comparison, line charts for trends" answer; naming the two real candidates and the actual viewing context gets a decision instead of a taxonomy.
Example 26: Reasoning trace + source grounding
Generic:
plain
Why did this metric spike?
Engineered:
plain
Signups jumped 340% on March 11th, visible in this data: [data]. List only
explanations the data itself can support — a referrer spike, a specific day-of-week
pattern, a duplicate-row artifact. Don't invent a marketing campaign or external
event I haven't mentioned. If the data can't tell you why, say that explicitly
instead of guessing.
Why it works: Asked "why" with no constraint, a model will often narrate a plausible-sounding cause — a launch, a press mention — that simply isn't in the data; telling it to stay inside what's actually there turns a confident guess into either a real finding or an honest "I don't know."
Example 27: Specificity lock + verification pass
Generic:
plain
Write a SQL query to get active users.
Engineered:
plain
Write a PostgreSQL query: users who logged in at least once in the last 30 days
AND haven't been marked deleted (deleted_at IS NULL). Table is users, columns
last_login_at and deleted_at. After the query, tell me what it would return if
last_login_at is null for a user who's never logged in — I want to make sure that
case is handled, not assumed.
Why it works: "Active users" is a business term with no fixed SQL meaning, and not naming the dialect is how you get TOP 10 syntax in a Postgres file; asking what happens to the null-login case catches the silent exclusion bug before it ships, not after a stakeholder asks why the count looks low.
Example 28: Audience lock + format lock
Generic:
plain
Turn this dataset into a report.
Engineered:
plain
This goes to a board that reads for five minutes, not five pages. One paragraph
of context, one table with the 4 numbers that matter (not all 14 columns), one
sentence on what you'd watch next quarter. No chart unless a number alone is
misleading without it.
[dataset]
Why it works: "A report" defaults to comprehensive, which is the wrong target for a five-minute reader; naming the actual reading context and capping the table to four columns forces a real editorial choice about what matters instead of dumping everything and calling it thorough.
Example 29: Verification pass + failure mode naming
Generic:
plain
Did variant B win the A/B test?
Engineered:
plain
Variant B converted at 4.1% vs 3.8% for control, n=1,200 per arm, 14 days.
Before saying which "won," tell me whether this difference is large enough to
trust given that sample size, and name the most likely way this result reverses
itself if I ran it for another 2 weeks. I'd rather hear "too early to call" than
a confident answer that's wrong.
Why it works: A small percentage gap on a modest sample is exactly the setup that produces a confident-sounding wrong answer — the same gap the Microsoft Research Copilot study reported as a 55.8% productivity gain came with a 95% confidence interval running from 21% to 89%, and that width is the honest part most write-ups leave out.
Example 30: Constraint budget + negative constraint
Generic:
plain
Forecast our revenue for next year based on this data.
Engineered:
plain
Project revenue for next quarter only, not the year — further out than that isn't
a forecast, it's a guess wearing a forecast's clothes. Show the trend-line projection
and state the two assumptions it depends on. Don't smooth over a seasonal dip in the
data to make the line look cleaner.
[data]
Why it works: I no longer ask a model for anything past one quarter out, and you shouldn't either — the visible confidence of the output doesn't shrink as the horizon grows, even though the actual reliability does, so capping the ask is the only honest move.
Support & Sales
9 examples. A support reply that sounds like every other support reply is the fastest way to make someone feel like a ticket number. The fix usually isn't tone — it's giving the model the one fact that makes the situation specific instead of generic.
Example 31: Role + stakes + negative constraint
Generic:
plain
Write a reply to this angry customer email.
Engineered:
plain
This customer's order arrived broken for the second time in a month — this is their
second email, not their first. Write the reply as someone who can see that history
and is genuinely annoyed on their behalf at our process, not apologizing for them
being upset. No "we sincerely apologize for any inconvenience." Offer the specific
fix (replacement shipped today, no return needed) before anything else.
[email]
Why it works: Naming that this is the second failure, not the first, changes the entire register of the reply — a generic apology to a second-time complaint reads as not having read the ticket, which is usually true.
Example 32: Context first + audience lock
Generic:
plain
Write a follow-up email to a sales lead.
Engineered:
plain
Lead context: demo'd our product 8 days ago, asked twice about pricing for teams
over 50 seats, went quiet after I sent the quote. Write a follow-up that doesn't
chase ("just checking in!") — it should reference the specific 50-seat question and
either answer something they likely didn't ask, or name the probable reason a 50-seat
quote goes quiet (budget approval cycle). Under 90 words.
Why it works: "Just checking in" is the email equivalent of a hedging cluster — it asks nothing and offers nothing; naming the actual stall point (budget approval, not interest) gives the model something specific to write toward instead of a content-free nudge.
Example 33: Scope fence + persona transfer
Generic:
plain
Explain our refund policy to this customer.
Engineered:
plain
Explain why this specific purchase falls outside the 30-day window (bought day 34)
as a support rep who's allowed to offer a one-time courtesy exception, not just
recite the policy. Lead with the exception offer, then the policy reason, in that
order — don't make them read the rejection before the resolution.
Policy: [policy text]. Purchase date: [date].
Why it works: Reciting policy first and the resolution second is technically accurate and reads as a wall; reordering so the good news lands before the explanation changes nothing about the actual decision but changes how the customer experiences it.
Example 34: Comparison frame
Generic:
plain
How do I respond to "it's too expensive"?
Engineered:
plain
"It's too expensive" can mean "I don't see the value yet" or "I genuinely can't
afford this" — give me a question that distinguishes which one I'm dealing with
before I respond to either, since the two responses shouldn't be the same.
Why it works: Most objection-handling advice gives you a rebuttal to the words, not the underlying reason — and the same line means two different things depending on the buyer, so the actually useful output is a diagnostic question, not a script.
Example 35: Few-shot anchor + format lock
Generic:
plain
Make this canned response sound less robotic.
Engineered:
plain
Here's our current macro for "how do I cancel": [macro]. Here's an example of a
reply we wrote from scratch that customers responded well to: [example]. Match the
second one's structure — short sentences, no "we understand your frustration" — but
keep the macro's actual steps, since those are correct.
Why it works: "Less robotic" is a vibe with no anchor; giving a real example that already worked tells the model exactly which register to copy instead of guessing at what "more human" means to you specifically.
Example 36: Failure mode naming + reasoning trace
Generic:
plain
Write an email to win back a customer who might churn.
Engineered:
plain
Usage dropped from daily to zero over 3 weeks, no support tickets, no cancellation
request yet. Before writing the email, name the two most likely reasons usage drops
silently like this (not "they're busy" — something more specific), then write toward
whichever is more likely rather than a generic "we miss you."
Why it works: A silent drop-off with no complaint usually means the product stopped fitting a workflow, not that the customer forgot it exists; naming the real candidate causes first stops the email from defaulting to a discount offer that doesn't address why they actually left.
Example 37: Specificity lock
Generic:
plain
Write a personalized cold email to this prospect.
Engineered:
plain
Prospect: VP Ops at a 200-person logistics company, posted last week about warehouse
turnover hitting 40%. Open with that post specifically, not "I noticed you work in
logistics." Connect it to one concrete thing our product does for onboarding speed,
not a feature list. Three sentences, no "I hope this finds you well."
Why it works: "Personalized" without a specific, recent detail produces a mail-merge with the company name swapped in; naming the actual post they wrote is the difference between a cold email that gets a reply and one that gets reported as spam.
Example 38: Role + stakes
Generic:
plain
Write a calm response to defuse this situation.
Engineered:
plain
This customer is threatening to post about us publicly over a billing error that was,
on review, actually our mistake. Write the reply as someone who's already fixed the
billing error (refund processed) and is now writing only to acknowledge the error
plainly — not to talk them out of posting, and not to over-apologize for something
already corrected.
[message]
Why it works: De-escalation prompts often produce something that sounds like it's managing the threat instead of the actual problem; separating "fix the error" from "respond to the anger" keeps the reply from reading as damage control.
Example 39: Constraint budget + negative constraint
Generic:
plain
Write an upsell email for our premium tier.
Engineered:
plain
This customer hit their plan's usage limit twice last month — that's the only reason
to write this email, so lead with it. One sentence on what the upgrade actually removes
(the limit), one on price difference, nothing about features they haven't used. If they
hadn't hit the limit, don't send this email at all; say so if the data doesn't support it.
Why it works: Upsell emails sent on a schedule rather than a trigger read as exactly that; anchoring the email to a real usage event, and telling the model to refuse the premise if the event isn't there, keeps "we think you'd love premium" from going out to someone with no reason to care.
Marketing & SEO
9 examples. Most marketing prompts fail because they ask for the finished asset instead of the decision behind it — which platform, which intent, which version wins. Make the decision explicit and the asset gets easier to write.
Example 40: Output schema + specificity lock
Generic:
plain
Write an SEO meta description for this article.
Engineered:
plain
Meta description, max 155 characters, for an article about [topic]. Target keyword
"[keyword]" must appear in the first 60 characters, naturally, not stuffed. It should
make someone choose this result over a near-identical competing title — name the one
thing this article has that a generic version wouldn't.
Why it works: A character limit and keyword position without a differentiation ask gets you a technically correct description that reads exactly like the nine others on the results page; the differentiation clause is what actually earns the click.
Example 41: Audience lock + format lock
Generic:
plain
Write a social media caption for this post.
Engineered:
plain
Same announcement, three separate captions, not one caption reused: LinkedIn
(professional context, can be longer, lead with the implication for their work),
Instagram (visual-first, caption supports the image, doesn't repeat what's visible
in it), X (under 200 characters, one idea, no hashtags). Don't write a generic
version and tell me to "adapt as needed."
Why it works: One caption copy-pasted across platforms is the most common tell of an unmanaged account; naming what each platform's format actually rewards forces three genuinely different pieces of writing instead of one with the line breaks changed.
Example 42: Comparison frame + constraint budget
Generic:
plain
Write 5 versions of this ad copy.
Engineered:
plain
Write 3 versions, each testing a genuinely different angle, not a rephrase: one leads
with price, one leads with the specific problem it solves, one leads with social proof
(a number, not a vague claim). Same length, same CTA, only the opening line changes —
that's the only way the test tells you anything.
Why it works: Five variants that all say the same thing in different words don't test anything; capping it at three forces each one to isolate a real variable, which is the only way an A/B result means something afterward.
Example 43: Decomposition + scope fence
Generic:
plain
Write a content brief for a blog post.
Engineered:
plain
Brief for a freelance writer who knows the industry but not our specific angle.
Include: the one belief we want the reader to leave with, two sources they should
NOT lean on (too generic, already overdone on page one of search), one detail only
we'd know to include, target word count with a reason for that number, not a round
guess. Don't include a keyword list — that's not their job.
Why it works: A brief that's just a topic and a word count produces writing indistinguishable from the rest of the search results; naming what to avoid and the one detail that makes it ours is the actual brief, the rest is paperwork.
Example 44: Reasoning trace + audience lock
Generic:
plain
What keywords should I target for this page?
Engineered:
plain
For each of these keywords, tell me whether the likely searcher wants to learn,
compare, or buy — and flag any where this page's current content matches the wrong
intent. Don't recommend keyword volume, I have that data already; I need the intent
match, since that's what's actually missing.
[keyword list]
Why it works: A page can rank for a keyword and still get zero conversions because it answers the wrong question for that search — ranking for the right keyword with the wrong intent match is worse than not ranking, because it burns the impression for nothing.
Example 45: Self-critique loop
Generic:
plain
Give me 10 headline options for this article.
Engineered:
plain
Give me 4 headline options. Then, before I pick, tell me which one you'd cut and
the specific reason — vague, overpromises relative to the actual content, or just
generic — and which one you'd actually run if it were your budget.
Why it works: A list of ten headlines with no opinion attached offloads the actual decision back onto you; asking the model to argue against its own weakest option and commit to a favorite gets you reasoning you can disagree with instead of an undifferentiated list.
Example 46: Decomposition + format lock
Generic:
plain
Turn this blog post into a Twitter thread.
Engineered:
plain
Pull the single strongest claim from this post, not a summary of all of it — a thread
that tries to cover everything reads like a table of contents. 6 posts max, each one
standalone enough to be quoted on its own, building to the claim rather than restating
the headline at the top.
[post]
Why it works: A thread that compresses an entire article loses the thing that made any one part worth reading; picking the sharpest single claim and building toward it gives the thread its own reason to exist instead of being a worse version of the link.
Example 47: Comparison frame + few-shot anchor
Generic:
plain
Does this sound like our brand voice?
Engineered:
plain
Here are 3 pieces we've published that we consider on-voice: [examples]. Here's a
new draft: [draft]. Point to the specific sentences in the draft that don't match —
not a vague "tone feels off" — and say what's different about them structurally,
not just the word choice.
Why it works: "Does this sound on-brand" with no reference produces a yes/no guess based on the model's own idea of your brand; giving real anchor examples turns a subjective check into a structural comparison it can actually point to.
Example 48: Role + stakes + negative constraint
Generic:
plain
Write a brief for an influencer partnership.
Engineered:
plain
Brief for a creator who knows their audience far better than we do. State the one
outcome we actually need (a specific action, not "awareness"), the one fact about
the product they must get right, and explicitly: no required script, no mandated
phrasing, no "must mention" list beyond that one fact. Tell them what success looks
like, not what to say.
Why it works: A brief that scripts the creator's words is how you end up with content that performs worse than their normal posts — naming the actual goal and getting out of the way of the phrasing is the trade most brands say they'll make and then don't.
What could be wrong with this post
It's Part 1, not the full set. The remaining six categories — research, image/video/design, agentic workflows, education, legal/HR/admin, and a "prompts that failed" postmortem — aren't in this draft. If you found this post looking for those, check Part 2 or ask for them directly.
Portability across models is a claim, not a guarantee. Most of these patterns hold on Claude, GPT-5, and Gemini as of mid-2026, but model updates ship faster than articles do. If a rewrite underperforms the generic version on your model, that's real signal — don't assume you did it wrong.
The context-engineering framing is genuinely contested. Some practitioners argue prompt engineering was never really "replaced," just absorbed as one layer of a bigger stack. This article leans on the newer framing because it matches what I see in my own work, not because it's the only defensible read.
None of this was tested with a formal before/after benchmark. These are patterns from repeated use, not a controlled study — the two cited papers and the Copilot RCT are the only claims here with that level of rigor.
Frequently Asked Questions
Is prompt engineering still worth learning in 2026?
Yes, but as one layer, not the whole skill. The structural patterns in this article — audience lock, scope fence, verification pass — still change output quality in a single message. What's changed is that they're no longer the ceiling on what determines whether an AI system works; for anything multi-step or agentic, what you retrieve and feed into context matters at least as much.
Do these prompts work the same on Claude, GPT-5, and Gemini?
Most of the structural fixes here — scope fences, audience locks, output schemas — are model-agnostic, because they're really about what information the model has, not phrasing tricks specific to one vendor. A few examples note model-specific behavior where it's known to matter, like Claude's documented tendency to over-deliver on unscoped reviews.
Why do so many of these examples add negative constraints ("don't do X")?
Because most AI output problems are over-delivery, not under-delivery: extra caveats, extra scope, extra "helpfulness" you didn't ask for. Telling the model what to leave out is often more load-bearing than telling it what to include, since the default behavior already covers the basics.
Should I use one giant prompt with every rule I can think of?
No. The Microsoft Research Copilot trial and the "lost in the middle" research both point the same direction: a shorter prompt with the two or three constraints that actually matter, placed at the start or end, tends to outperform an exhaustive one where the important instruction is buried in the middle.
What's the single highest-leverage change I can make to my prompts today?
Add a verification step: ask the model to check its own output against one stated constraint before it finishes. It's the cheapest addition in this entire list and the one with the most consistent effect on catching silent errors, in my own use.
Is "context engineering" just a rebrand of prompt engineering?
Partly, and reasonable people disagree on how much. The practical distinction is scope: prompt engineering is about the wording of a single instruction, while context engineering also covers what gets retrieved, what history is kept, and what tools are exposed across a multi-step task. For a one-off request to a chat interface, that distinction mostly doesn't matter yet.
When is Part 2 of this series coming?
It's in production now, covering research & decisions, image/video/design, agentic workflows, education, and legal/HR/admin prompts, plus a postmortem section on prompts that failed and why. Follow here for updates.
About the author
Tom Morgan writes about applied AI tooling and prompt structure. This article draws on roughly three years of writing and testing prompts for content, support, and data workflows — mostly small-to-mid-size teams, mostly B2B and e-commerce. It doesn't cover enterprise-scale agentic deployments in depth; that's a different practice with different failure modes.
Disclosure: No tool in this article is sponsored. Where a specific model (Claude, GPT-5, Gemini) is named, it's because a pattern was verified to behave differently on it, not as an endorsement.
A good prompt doesn't sound smarter. It just leaves the model fewer ways to guess wrong.
📚 More from this series: Part 2 — Research, Design, Agentic Workflows & Failed Prompts | Full Prompt Library`
Top comments (0)