Competitive benchmarking dies in two opposite ways: you compare yourself to nobody, or you compare yourself to everyone famous. In AI search, both produce vanity metrics—numbers that flatter or frighten without supporting a decision.
What "fair" means
A fair benchmark has four properties:
- Shared demand — prompts that all peers could win
- Honest peers — brands buyers actually shortlist
- Locked instruments — stable prompt text and scoring rules
- Decision linkage — each metric implies an owner and a next action
If any property is missing, you have theater.

Build the peer set first (not the chart)
Start from go-to-market reality:
- Who appears on RFPs and sales battlecards?
- Who shows up in analyst / retailer / marketplace shelves?
- Who wins the "alternative to X" content war?
Cap the set. Six to twelve peers beats forty logos. Mix one aspirational leader, your true cluster, and one disruptive specialist if relevant. Revisit quarterly—not weekly.
Hub workflows help here: category and value cohorts are clues, not autopilot. A valuation peer is not automatically an answer-engine peer.
Choose metrics that survive interrogation
Prefer
- Unbranded mention rate by prompt family (category / comparison / best-of)
- Share of voice vs the named peer set on those families
- Framing flags (price, complexity, trust adjectives)
- Citation concentration and churn
- Engine disagreement rate on strategic prompts
Avoid as primary KPIs
- Branded-only mention rate ("people know our name when asked about our name")
- Single-prompt screenshots
- Unweighted averages across junk prompts
- "Share of voice vs the entire internet"
- Sentiment without presence (vibes on zero mentions)
Prompt hygiene for benchmarks
- Same wording for the quarter
- Explicit geography / language
- No competitor names in category prompts unless the family is comparison by design
- Tags for segment and product line so losses are diagnosable
When leadership asks "are we winning?" answer with a prompt family, not a single lucky chat.
How to present without getting destroyed
Bad slide: "AI SOV is 18%."
Better slide: "On 24 unbranded mid-market category prompts, SOV vs peer set is 18% (was 12%). Gap is concentrated in implementation-focused prompts where Peer B owns documentation citations."
Then show the tickets.
Connect benchmarks to brand structure
If you always lose globalization-tagged prompts, check BrandSight globalization and regional evidence—not just ads. If you win momentum but lose stability-framed adjectives, stop celebrating spikes. Use BrandSight and BrandValue as context layers so AEO benchmarks do not float free of brand strategy.
A 30-day benchmarking reset
- Freeze peer set and prompt set version.
- Run two weekly snapshots; discard week-one instrumentation bugs.
- Publish a one-page baseline with families and framing notes.
- Pick three repairs that would change the baseline.
- Re-measure; report deltas only on the frozen set.
FAQ
Who owns the peer set and prompt lock?
The AEO or competitive-intelligence DRI proposes changes; marketing leadership approves quarterly revisions with written rationale. Never let a bad week silently rewrite peers or prompts to protect a vanity number.
How do we attach benchmarks to work we already do?
Freeze the instrument first, then route losses to the evidence repair backlog with prompt-family tags—not single-screenshot panic. Each baseline slide should end with tickets, not just deltas.
What is the most common vanity failure mode?
Celebrating SOV against forty logos, the whole internet, or branded-only prompts that never test category shortlists. Fair benchmarks name 6–12 rivals and locked unbranded families.
What should leadership NOT ask for in a benchmark readout?
Single lucky chats, unweighted averages across junk prompts, or sentiment scores on zero mentions. If the number cannot survive "so what do we do on Tuesday?" it does not belong on the scoreboard.
Where BrandAI fits
- BrandAEO — locked prompts, honest peers, and deltas that survive a leadership meeting
- Brand Hub — investigative context when a loss is structural, not just a weak URL
- BrandSight / BrandValue — keep AEO benchmarks from floating free of brand strategy
Bottom line
Vanity benchmarks optimize for meetings. Fair benchmarks optimize for shortlists. BrandAI gives you the measurement surface (BrandAEO) and the investigative context (Hub and sister products). Your job is to keep peers honest, prompts locked, and metrics cruel enough to be useful.
If a number cannot survive the question "so what do we do on Tuesday?" it does not belong on the scoreboard.
