A brand-visibility dashboard is only as trustworthy as the prompt set behind it.
If prompts change every week, mix branded and non-branded questions, or reflect what the marketing team hopes people ask instead of what buyers actually ask, the resulting score may look precise while measuring very little.
This guide shows how to design a prompt benchmark that can be repeated, audited, and improved over time.
1. Define the decision the benchmark should support
Start with a practical decision, not a metric.
Examples:
- Which product categories fail to surface the brand?
- Which competitors are recommended for high-intent questions?
- Which sources influence answers in a specific market?
- Does a documentation or content change improve visibility?
- Are AI answers describing the product accurately?
A benchmark designed for competitive discovery will use different prompts from one designed for factual monitoring. Writing the decision down prevents the prompt set from becoming an unstructured list of interesting questions.
2. Separate discovery from verification
Use two top-level prompt groups.
Non-branded discovery prompts
These test whether an answer engine surfaces the company without being handed its name.
Useful categories include category discovery, problem diagnosis, product comparisons, alternatives, implementation, vendor selection, and audience-specific recommendations.
Examples:
- What tools help a B2B team monitor brand visibility in AI answers?
- How can a company track which sources are cited by answer engines?
- What should an AI-search visibility audit include?
- Which platforms help compare brand share of voice across AI assistants?
Branded verification prompts
These test whether the system understands the entity correctly.
Examples:
- What does Corank do?
- Who is Corank designed for?
- What is the difference between Corank and a traditional rank tracker?
Do not combine these groups into one headline mention rate. A model repeating a brand that appears in the prompt is not the same as discovering it independently.
3. Build an intent taxonomy
Every prompt should have a stable intent label.
| Intent | What it measures | Typical value |
|---|---|---|
| Definition | Topic understanding | Low to medium |
| Problem | Pain recognition | Medium |
| How-to | Implementation relevance | Medium |
| Comparison | Competitive positioning | High |
| Alternative | Shortlist inclusion | High |
| Recommendation | Vendor discovery | High |
| Branded | Entity accuracy | Diagnostic |
The taxonomy lets you report results by commercial importance. A high mention rate on branded prompts cannot hide a zero rate on recommendations.
4. Collect candidate prompts from evidence
Do not invent the entire benchmark in a conference room.
Candidate sources can include:
- sales-call questions
- support tickets
- website search logs
- Search Console queries
- autocomplete and related questions
- community discussions
- competitor comparison pages
- product onboarding questions
- internal customer research
Rewrite candidates into natural questions, but preserve the underlying intent.
Deduplicate prompts that ask the same thing with minor word changes. Keep deliberate variants only when they test a real distinction, such as audience, location, language, or company size.
5. Give every prompt a stable identifier
A prompt record should include more than the sentence.
A minimal schema contains:
- prompt_id
- prompt_text
- intent
- topic
- audience
- priority
- brand_in_prompt
- locale
- version
- active status
The identifier allows the wording to evolve without losing history. Increment the version when wording changes materially.
6. Define eligibility before scoring
Not every answer should be included in every metric.
For example, a mention-rate calculation may include only answers that actually return vendors. If a model refuses the question, returns an unrelated answer, or asks for clarification, record the run but mark it ineligible for the vendor-mention denominator.
Write the eligibility rule before reviewing results. Otherwise it is too easy to exclude inconvenient answers after the fact.
7. Use a fixed competitor set
Choose the competitor set at the start of a reporting period.
The list can include direct product competitors, adjacent tools that appear in the same answers, incumbents buyers use as substitutes, and an other category for unexpected brands.
Do not change the list because a new answer looks favorable or unfavorable. Add newly discovered competitors during the next versioned review.
8. Run repeated samples
Generative answers vary. A single run cannot show consistency.
For high-priority prompts, run the same question more than once. Store each platform-prompt-run combination as its own observation.
A run record should contain:
- run_id
- prompt_id
- platform
- run_number
- run_at
- raw_answer
- brands_mentioned
- target_brand_mentioned
- citations
- eligibility
Repeated runs reveal whether a brand is consistently visible or only appears occasionally.
9. Score multiple outcomes
Avoid reducing the benchmark to one opaque number.
Track at least:
Mention rate
Eligible answers containing the target brand divided by eligible answers tested.
Share of voice
Target-brand appearances divided by all appearances from the fixed competitor set.
Position
Classify the brand as a primary recommendation, shortlist member, example, source citation, or negative mention.
Citation coverage
Capture every cited URL and domain. Track which sources support each brand and prompt group.
Accuracy
Use a small rubric: accurate, partly accurate, inaccurate, or not enough information.
Keep notes for stale names, incorrect features, outdated pricing, and category confusion.
10. Preserve the evidence
Every score should be traceable to a raw answer.
Store the exact prompt, prompt version, platform, timestamp, complete answer, brands mentioned, nearby mention context, cited URLs, evaluator decision, and evaluator notes.
If a chart cannot be audited back to these records, it is not ready to guide content, PR, or product decisions.
11. Version changes instead of rewriting history
A benchmark will evolve. New products appear, customer language changes, and platforms add features.
Review the prompt universe on a fixed schedule. When a prompt is added, retired, or rewritten:
- record the reason
- preserve the old version
- state the effective date
- avoid comparing incompatible versions as if nothing changed
For long-term reporting, keep a stable core set and a rotating research set.
The stable core supports trend analysis. The research set explores new topics without corrupting the baseline.
12. Add quality assurance
Before each full run, check for duplicate prompt IDs, missing intent fields, accidental brand mentions in discovery prompts, changed competitor lists, inconsistent platform names, missing raw answers, and duplicate run IDs.
After evaluation, sample a portion of the records for a second review. Agreement matters most for subjective fields such as position and partial accuracy.
A simple operating rhythm
A sustainable process can use:
- weekly high-intent prompt monitoring
- monthly full benchmark runs
- quarterly prompt-set and competitor reviews
The benchmark should change slowly enough to support trends and quickly enough to remain relevant.
What to avoid
Common failure modes include testing only branded prompts, selecting questions after seeing favorable answers, changing wording during every run, treating one platform as the whole AI-search market, counting mentions without context, hiding the denominator, and discarding answers that do not fit the expected narrative.
Closing principle
The purpose of a prompt benchmark is not to prove that a brand is visible.
It is to create a repeatable test that can show where the brand appears, where it does not, what the systems say, and which sources influence the result.
We are building Corank around repeatable monitoring of brand mentions and cited sources across AI answers. The product is available at https://corank.ai.
Disclosure: I work on Corank.
Top comments (0)