DEV Community

draftkit
draftkit

Posted on

10 AI Prompts I Use for UX Research and Design (Interviews, Personas, Usability Tests)

I've spent the last year running UX research and design work with an AI assistant in the loop. Not to replace research — to speed up the parts that eat hours without adding insight: turning 12 interview transcripts into themes, writing a usability test script that doesn't lead the witness, drafting personas that don't read like stereotypes.

Below are 10 prompts I keep coming back to. Each one has a fixed structure — Role, Context, Constraints, Output — and each has been tested against real inputs. I'll show the prompt, then what it produced, then why it works.

If you want the full library (50+ prompts across UX, research, design ops, and product work), I keep an updated pack here. But these 10 are the ones I use most weeks.


1. Interview synthesis that preserves verbatim quotes

Role: You are a senior UX researcher.
Context: I'm pasting 6 interview transcripts (30 min each) from a study on how freelance designers choose project management tools.
Constraints:

  • Group findings into themes. Each theme needs ≥3 quotes from ≥2 different participants.
  • Preserve quotes verbatim — no paraphrasing, no "cleaning up."
  • Flag any theme supported by only one participant as "single voice."
  • Do NOT invent quotes. If a theme has no quote, say so. Output: A themes table: Theme | Quote 1 (P#, verbatim) | Quote 2 | Quote 3 | Confidence (high/med/single-voice).

Why it works: The "no paraphrasing" constraint kills the thing LLMs do worst — smoothing quotes into bland summaries that lose the participant's actual words. The "single voice" flag prevents one loud participant from dominating the synthesis.


2. Persona built from research, not stereotypes

Role: You are a UX researcher writing a persona.
Context: Based on these 6 transcripts, build a persona. Avoid stock persona language.
Constraints:

  • Name, role, and one specific quote that anchors them.
  • Goals: 3, each tied to a verbatim quote.
  • Frustrations: 3, each tied to a verbatim quote.
  • Banned phrases: "tech-savvy," "time-poor," "busy professional," "values efficiency," "wants things to just work," "demands seamless experiences."
  • Do NOT invent demographics not in the data (age, location, income). Output: One-page persona markdown.

Why it works: Personas go wrong when they drift into fiction. Banning the stock phrases forces the model to reach for the actual words participants used. Anchoring every goal/frustration to a quote keeps it honest.


3. Usability test script that doesn't lead the witness

Role: You are a moderated usability test moderator.
Context: I'm testing a new onboarding flow for a budgeting app. 5 participants, 45 min each.
Constraints:

  • Task-based, not question-based. Each task starts with a goal, not "click here."
  • Neutral language. No words like "easy," "simple," "obviously," "as you'd expect."
  • After each task: one open probe ("What were you expecting?"), one specific probe ("What made you click there?"). Never "Why did you do that?" — it sounds accusatory.
  • End with one question the participant must answer without help: "If you had to do this again tomorrow, what would you remember?" Output: Full script: intro, 6 tasks with probes, post-task questions, wrap-up.

Why it works: Leading questions are the #1 usability test sin. Banning the loaded words and forcing neutral probes catches the subtle bias that creeps in. The "remember tomorrow" question surfaces mental models better than "how was that?"


4. Heuristic review that cites the specific heuristic

Role: You are a Nielsen Norman-style heuristic evaluator.
Context: Here are 12 screens from a checkout flow. Review against Nielsen's 10 heuristics.
Constraints:

  • For each finding: Severity (1-4) | Heuristic violated (name it specifically) | Screen | What's wrong | Why it matters for the user | Suggested fix (one sentence).
  • Do NOT list "violations" that are actually preferences. If it's a preference, label it "Polish, not violation."
  • Group by severity, not by heuristic. Output: Severity-ordered findings table + a 3-bullet "top priorities" summary.

Why it works: Heuristic reviews turn mushy when the evaluator waves at "consistency" without naming which consistency. Forcing the specific heuristic name and a severity score keeps it actionable. The "polish vs violation" split stops bikeshedding.


5. Survey that avoids double-barreled and leading questions

Role: You are a survey designer.
Context: I need a 12-question survey to understand why users churn after the free trial.
Constraints:

  • No double-barreled questions (one concept per question).
  • No leading questions (no "How much do you love…").
  • No loaded wording ("powerful," "intuitive," "frustrating").
  • Mix of scale (1-5), multiple choice, and one open-ended at the end.
  • The open-ended question must be neutral: "What almost stopped you from continuing?" not "What did you dislike?" Output: Survey in order, each question labeled with type.

Why it works: Survey design has well-known failure modes. Naming them explicitly as constraints makes the model police itself. The neutral open-ended reframe is the single biggest lever — "what almost stopped you" gets more honest answers than "what did you dislike."


6. Wireframe annotation a developer can build from

Role: You are a designer writing wireframe annotations for engineering.
Context: Here's a low-fi wireframe of a settings page. Write the handoff notes.
Constraints:

  • Per component: Purpose | States (default/hover/active/disabled/error) | Empty state | Loading state | Edge cases (max char, null, permission denied).
  • No design-system jargon without a link. If you say "use the primary button," link to the component.
  • Flag any state you can't infer from the wireframe as "ASSUMPTION — confirm with designer." Output: Component-by-component annotation, grouped by section.

Why it works: Handoff docs fail when they describe the happy path and forget states. Forcing every component to enumerate its states catches the gaps. The "ASSUMPTION" flag makes the model admit uncertainty instead of inventing specs.


7. Competitor teardown focused on decisions, not features

Role: You are a UX strategist.
Context: Here are 3 competitor apps in the habit-tracking space. I want a teardown, not a feature list.
Constraints:

  • For each competitor: One decision they made that's interesting (not "they have streaks" — everyone has streaks). Why it's interesting. The trade-off it implies.
  • No "feature comparison table." Decisions, not checkboxes.
  • End with: "What I'd steal" (1 thing) and "What I'd avoid" (1 thing) per competitor. Output: Per-competitor analysis + a cross-competitor "patterns" section.

Why it works: Feature lists are useless — they tell you what exists, not why. Forcing "one interesting decision + the trade-off" surfaces the actual design thinking. "What I'd steal / avoid" forces a point of view.


8. Accessibility review beyond color contrast

Role: You are an accessibility specialist (WCAG 2.2 AA).
Context: Here's a Figma-exported screen description. Review for accessibility.
Constraints:

  • Cover: color contrast, focus indicators, alt text, heading hierarchy, touch target size, form labels, error identification, motion/animation, keyboard nav.
  • For each finding: WCAG criterion number | Issue | Who it affects | Fix.
  • Do NOT only flag contrast. If contrast is the only issue, say "contrast is the only automated-detectable issue; manual testing needed for the rest." Output: Findings by category + a "needs manual testing" list.

Why it works: LLMs default to "check contrast" and stop. Enumerating the categories forces a fuller pass. The honesty caveat ("manual testing needed") prevents false confidence — automated review has real limits.


9. Research recruitment screener that filters without biasing

Role: You are a UX research recruiter.
Context: I'm recruiting 8 freelance designers who use project management tools weekly. Write a screener.
Constraints:

  • Disqualifying questions must be behavior-based, not demographic ("How many hours per week do you use a project management tool?" not "Are you a designer?").
  • No "killer questions" that telegraph the right answer ("Do you love trying new tools?").
  • Include one attention-check question ("Select 'sometimes' for this question").
  • Incentive, time commitment, and recording consent stated upfront. Output: Screener survey + a one-line recruitment blurb for each channel (Twitter, Slack, email).

Why it works: Screeners go wrong when they signal what you're looking for. Behavior-based questions and the "no killer questions" rule keep the filter honest. The attention-check catches panel-study junk responses.


10. Stakeholder readout that leads with the decision needed

Role: You are a UX researcher writing a readout for a product VP.
Context: Here are the findings from a 6-person usability study. The VP has 5 minutes.
Constraints:

  • Line 1: The decision this research enables or blocks (one sentence).
  • Then: 3 findings max, each with a quote and a recommendation.
  • Banned: "interesting insights," "rich data," "participants generally," "users expressed."
  • End with: "Open question" — the one thing this study couldn't answer. Output: One-page readout. No executive summary separate from the body — the body IS the summary.

Why it works: Researchers bury the lede. Forcing line 1 to be the decision reframes the whole doc around action. Banning the filler phrases ("interesting insights") forces specificity. The "open question" builds trust — it says you know the limits of your own research.


How I test these prompts

Three steps, every time:

  1. Run it 3× on the same input. If the output varies wildly, the prompt is under-specified. Tighten the constraints.
  2. Generalize to a second domain. Take a prompt built for UX research and run it on, say, customer support transcripts. If it breaks, the constraints are too narrow. If it holds, it's portable.
  3. Read it aloud. If it sounds like an AI wrote it — "delve into," "in today's fast-paced world," "it's important to note" — rewrite the constraints to ban that register.

The pattern across all 10: constraints do more work than instructions. Telling the model what NOT to do (no paraphrasing, no stock phrases, no leading questions) produces sharper output than telling it what to do.


If these were useful, I keep a larger library — 50+ prompts covering UX research, design ops, product discovery, and stakeholder communication. It's here ($9). There's also a deeper developer productivity library and a SaaS marketing copy pack if you're wearing multiple hats.

Happy researching.

Top comments (0)