Measurement / Guide
How to measure AI visibility without misleading scores
Use a repeatable GEO measurement protocol for mentions, citations, factual accuracy, referrals, and leads, with explicit denominators and limitations.
The short answer
Measure AI visibility with a fixed prompt set, repeated observations, saved answers, and explicit denominators. Report mentions, citations, factual accuracy, and business outcomes separately. A sampled appearance rate is not market share or a universal AI ranking.
Define the question before the metric
A discovery test asks whether a business appears when the user does not name it. A branded accuracy test asks whether an answer about a named business is correct. Keep these groups separate.
Write the exact prompts in advance. Record the service, town, intended buyer need, platform, and collection method. If a prompt changes, give it a new version rather than silently replacing it in the historical comparison.
The GEO research paper studies visibility using a specific evaluation framework. Your local reporting method should describe its own population and limits instead of borrowing a research result as a business benchmark.
Use a repeatable observation protocol
Our suggested starting protocol is ten buyer questions, tested three times per platform across separate sessions. This is a manageable operational sample, not a statistically representative survey.
- Use a fresh conversation for each observation.
- Record the date, time, platform, visible mode, and location context.
- Save the prompt and complete first answer, including citations.
- Mark whether the answer is valid, unavailable, or a product error.
- Record business mentions, links to your domain, and factual issues.
- Preserve the raw evidence before summarizing it.
Keep platform results separate. Avoid running until the answer you want appears. If you repeat a failed request, retain the failure record and identify the retry.
Calculate rates with explicit denominators
For a defined platform and prompt group:
Observed mention rate = valid answers naming the business / all valid answers.
Observed site citation rate = valid answers linking to your website / all valid answers.
Count each answer once per metric, even if it mentions the company or links to the site more than once. Record third-party citations about the business separately.
For factual accuracy, count the checked claims rather than all answers: verified correct claims / checked factual claims. Keep unverifiable claims in a separate category and report their count. A response without a checkable fact is not automatically accurate.
A worked example
Suppose you plan 30 discovery observations in one platform. Two fail, leaving 28 valid answers. Seven name the business; four link to its site.
- Observed mention rate: 7 / 28 = 25%.
- Observed site citation rate: 4 / 28 = 14.3%.
- Collection failures: 2 / 30 = 6.7%.
These numbers are hypothetical. They do not describe Smoketown GEO or a client. They also do not mean the business reaches 25% of all users.
If a later run shows 9 mentions in 30 valid answers, report 30% and both denominators. Do not call this proof that a content edit caused an improvement.
Separate platform reports from manual samples
Bing’s AI Performance report provides information about citations in supported AI experiences. Preserve the reporting scope, date range, and metric definitions when using it.
A platform report and a manual prompt test observe different populations. Put them side by side rather than adding their counts together. Also distinguish citations from referral visits in analytics.
For business outcomes, record qualified inquiries and the attribution method. A customer’s self-reported discovery source is useful but incomplete; analytics may miss visits or lose referral information. Do not assign every direct visit to AI.
Make comparisons honest
Keep the prompt set and conditions stable. Note changed interfaces, location settings, source pages, site edits, and missing observations. A series can show movement without establishing why it happened.
Repeated answers are not necessarily independent samples. Avoid confidence claims that assume random sampling when you collected a convenience sample.
A useful report ends with decisions: which errors need correction, which gaps need content, and what to recheck. Use the audit checklist to connect those decisions to owners and completion evidence.
Check the evidence
Sources & review notes
Primary references checked for this edition on . Platform documentation can change. Recommendations and illustrative examples are Smoketown GEO’s editorial guidance, not guaranteed outcomes.
- Bing Webmaster Tools: AI Performance
- Aggarwal et al.: GEO — Generative Engine Optimization (KDD 2024)
How we source and update these guides · Suggest a correction