Most teams "check AI visibility" the way they used to check Google: open a tab, type something that feels representative, screenshot the answer, declare victory or panic.
That is not measurement. That is anecdote with better typography.
A scoreboard without a fixed prompt set is storytelling, not measurement.
If you want mention rate, share of voice, sentiment, and citation mix you can manage, you need a prompt set—a fixed portfolio of buyer questions re-run on a schedule across engines. BrandAEO is built around that idea. This guide is how to design the prompts themselves so the dashboard is not lying to you.
What a good prompt set is for
A prompt set is not a content calendar. It is a sample of demand.
You are trying to approximate the questions that create shortlists in your category—the moments when an assistant chooses a few brands and (sometimes) a few sources. If your prompts never create shortlists, you will understate competition. If they are all branded vanity queries ("Is Acme good?"), you will overstate your own presence.
| Property | What it means |
|---|---|
| Stable | Same wording week to week so trends mean something |
| Buyer-shaped | Language from RFPs, sales calls, and community forums—not internal product names |
| Comparable | Peers can plausibly appear alongside you |
| Covered | Category discovery, comparison, and recommendation intents |
Thirty to eighty prompts is a workable band for most B2B and consumer brand teams. Fewer than fifteen is usually noise. Hundreds without taxonomy is theater.
Three prompt families that matter

| Family | Buyer question | What it stresses |
|---|---|---|
| Category | Who belongs in the set? | Corroboration and category language |
| Comparison | Who wins the narrative? | Framing—adjectives as carefully as presence |
| Best-of | Who gets the crown? | Shortlist scarcity—fourth in a paragraph is often worthless |
1. Category prompts
These are the "who belongs in the set?" questions.
Examples of the shape (adapt to your market):
- Best [category] for [job / segment] in [region]
- Top [category] tools / brands for [use case]
- Alternatives to [incumbent] for [job]
Category prompts stress corroboration and category language. If your site and third-party pages never use the nouns buyers use, you disappear here first.
2. Comparison prompts
These are the "who wins the narrative?" questions.
- [Brand A] vs [Brand B] for [job]
- [Brand] compared to [peer list]
- Is [Brand] better than [peer] for [constraint]?
Comparison prompts stress framing. You can be mentioned and still lose: "capable but expensive," "legacy," "hard to implement." Track adjectives as carefully as presence.
3. Best-of / recommendation prompts
These are the "who gets the crown?" questions.
- Recommend a [category] for [scenario]
- What should I buy if I need [outcome] under [constraint]?
- Shortlist for [buyer persona] evaluating [category]
Best-of prompts stress shortlist scarcity. Being fourth in a fluent paragraph is often worthless. Being in the top three with a clean citation is the game.
Design rules that keep metrics honest
Freeze the wording. Changing "best CRM for startups" to "top CRM platforms for early-stage SaaS" mid-quarter resets your baseline. Version the set; do not silently edit live prompts.
Separate branded and unbranded. Keep a small branded slice for reputation ("What do people say about [Brand]?"), but do not let it dominate the scorecard. Unbranded demand is where discovery happens.
Name peers deliberately. A peer set that is too weak flatters you. A peer set that is pure mega-cap giants can hide mid-market wins. Align peers with how Brand Hub and your sales team already define the competitive field.
Localize on purpose. If you sell in multiple languages or regions, duplicate the intent, not a machine-translated mash. One English global set plus a few market-critical local sets beats one messy multilingual blob.
Tag every prompt. Module (visibility / sentiment / authority), funnel stage, persona, and product line. Without tags, you cannot answer "where did we drop?" without re-reading hundreds of answers.
Common traps
| Trap | Why the dashboard lies |
|---|---|
| The demo prompt | You pick questions your content already answers perfectly—buyers asking harder questions never see you |
| The SEO keyword dump | Exact-match keyword salad does not mimic how people talk to ChatGPT, Gemini, Claude, or Perplexity |
| The weekly rewrite | Marketing "optimizes" prompt text after every bad week—you destroy trend integrity |
| The single-engine habit | One model's personality is not the channel; cross-engine disagreement is itself a signal |
A practical build sequence
- Pull 20 real questions from sales, support, and community.
- Expand into category / comparison / best-of variants until you hit a stable set size.
- Lock peer names and geographies.
- Run a baseline week in BrandAEO.
- Review where you are absent, weakly framed, or cited to stale sources.
- Fix the evidence graph (public facts, docs, encyclopedic coverage via BrandWiki workflows)—not the prompt wording.
What "good" looks like after 30 days
You should be able to say, without hand-waving:
- Mention rate on unbranded category prompts is X, up or down vs last period
- Share of voice vs named peers is Y on comparison prompts
- Citation domains concentrating risk are Z
- Three prompt tags explain most of the movement
That is measurement. Screenshots of a friendly lunchtime chat are not.
FAQ
Who should own prompt-set hygiene?
Name an AEO lead or brand research DRI who approves wording changes, maintains the version log, and blocks mid-sprint edits. Product marketing owns buyer language inputs; the DRI owns what enters the locked set.
What if marketing wants to rewrite prompts after a bad week?
Treat it as a process failure, not a measurement fix. Log the request, schedule changes for the next planned refresh (usually quarterly), and diagnose the movement on the current set. Rewriting prompts to manufacture green is how scoreboards become fiction.
How do we know the peer set is still honest?
Review peers monthly against how sales and Brand Hub define the competitive field—not how the dashboard looks. A peer set that flatters you or hides mid-market wins is a measurement bug, not a strategy win.
What should we never do with branded prompts?
Never let branded vanity queries dominate the scorecard. Keep a small reputation slice, but do not use "Is [Brand] good?" spikes to declare AI visibility victory while unbranded category prompts still show absence.
Where BrandAI fits
- BrandAEO — run the locked set, track mention rate, SOV, sentiment, and citations across engines
- Brand Hub — align peer definitions and investigate where prompts show gaps
- BrandWiki — encyclopedic and reference hygiene that category prompts actually detect
Design prompt sets like research instruments—and treat your brand story as something the instruments must be able to detect.
