Methodology
How we measure AI visibility
Most AI-visibility tools show you a number that changes every time you run it. Ours comes with a confidence interval. Here's exactly how the score is built, so you can trust it, and so a skeptic can verify it.
1. The questions we ask
We generate a library of customer questions from your category and a short description of what you do, the kinds of questions real customers type into ChatGPT, Claude, Perplexity, and Gemini when deciding what to buy, who to book, or where to go. One hard rule: we never put your brand or a competitor's name inside a question. Naming a brand guarantees the engine echoes it back, which would measure nothing. We measure whether you surface organically, the way a real customer's question would surface you.
Every generated question then passes a validation gate: a second model checks that its natural answer would actually name specific companies, providers, or products. Questions that would only return generic types or how-to advice (which can't surface any brand, and would silently drag a score down) are dropped and regenerated. Your set is validated, not just generated.
The set deliberately spans two kinds of question, because engines answer them differently: list-style questions (“what are the best X for Y?”), where we measure whether you appear and how near the top, and first-choice questions (“if you had to pick one X, what would it be?”), where the answer names a single winner. So the score reflects both your list visibility and your first-choice visibility.
You review, edit, add, or remove any question. The score is measured against the questions you approved, consistency matters more than perfection.
Every generated question then passes a validation gate: a second model checks that its natural answer would actually name specific companies, providers, or products. Questions that would only return generic types or how-to advice (which can't surface any brand, and would silently drag a score down) are dropped and regenerated. Your set is validated, not just generated.
The set deliberately spans two kinds of question, because engines answer them differently: list-style questions (“what are the best X for Y?”), where we measure whether you appear and how near the top, and first-choice questions (“if you had to pick one X, what would it be?”), where the answer names a single winner. So the score reflects both your list visibility and your first-choice visibility.
You review, edit, add, or remove any question. The score is measured against the questions you approved, consistency matters more than perfection.
2. The score (0–100)
For each question, on each engine, we detect whether you were mentioned, where you ranked, and the sentiment of the mention. Those roll up into five components, each worth up to 20 points:
Every component is a continuous rate, not a yes/no flag: appearing in 4 of 25 questions scores proportionally, not identically to appearing in all 25. The score is a deterministic function of the stored results, so it's reproducible without re-querying the models.
- Mention rate, how often you appear at all
- Top-3 rate, how often you appear near the top
- First-mention rate, how often you're named first
- Sentiment, how positively you're described when mentioned
- Cross-engine coverage, how many engines carry you
Every component is a continuous rate, not a yes/no flag: appearing in 4 of 25 questions scores proportionally, not identically to appearing in all 25. The score is a deterministic function of the stored results, so it's reproducible without re-querying the models.
3. Why your score is stable (the confidence interval)
AI models are stochastic: ask the same question twice and the answer can differ. Run a single scan and your score bounces a few points from sampling noise alone. Showing you that single noisy number would be false precision.
Instead, paid plans scan daily, and your headline score is pooled over a rolling window of recent scans. From how much those daily scores actually vary, we compute a 95% confidence interval using the Student-t distribution, so your score reads like “62 ± 3”, not a falsely-exact 62. When we say your score moved, it moved by more than noise explains.
Instead, paid plans scan daily, and your headline score is pooled over a rolling window of recent scans. From how much those daily scores actually vary, we compute a 95% confidence interval using the Student-t distribution, so your score reads like “62 ± 3”, not a falsely-exact 62. When we say your score moved, it moved by more than noise explains.
4. When your questions change
If you edit your question set, we don't blend the old and new results. That would be comparing two different measurements. Each scan records the version of the questions it ran, and the pooled score only averages scans on your current version. After an edit, the window shrinks to post-edit scans and grows back as new daily scans land. Every number stays attributable.
5. How we query the models, and an honest caveat
We query the official APIs (OpenAI, Anthropic, Perplexity, Google), never scraped consumer apps. On Growth and Scale we enable each provider's live web-search tool so answers reflect current information, not just training data.
We deliberately do not lower the models' temperature to fake a smoother number. The variance is part of what we measure, it's what real customers experience at default settings. We handle it with sample size and intervals, not by suppressing it.
The honest caveat: API responses aren't identical to what a specific person sees in the ChatGPT or Gemini app, which layer on search and personalization. No measurement tool can be. What we guarantee is a consistent instrument, the same validated questions, the same way, every day, so the trend and the competitive gap are real even where an absolute number can't be.
We deliberately do not lower the models' temperature to fake a smoother number. The variance is part of what we measure, it's what real customers experience at default settings. We handle it with sample size and intervals, not by suppressing it.
The honest caveat: API responses aren't identical to what a specific person sees in the ChatGPT or Gemini app, which layer on search and personalization. No measurement tool can be. What we guarantee is a consistent instrument, the same validated questions, the same way, every day, so the trend and the competitive gap are real even where an absolute number can't be.
6. Local vs. global measurement
You choose the competitive arena we measure. Global runs category questions with no geography. Local anchors every question to an explicit place you name (“best [category] in Portland”). We deliberately never use the phrase “near me”, a real customer's “near me” depends on their own location, which no tool can know, so it isn't measurable or reproducible. Explicit geography is the honest, repeatable way to measure local visibility.
7. How long a fix takes to show up
Longer than most tools admit. The engines answer from live search results, so a change you make only counts once it has been crawled, indexed, and ranked well enough to be retrieved into an answer. Those are three separate delays stacked on top of each other, and none of them are ours to control.
In practice: claiming a profile or getting listed somewhere the engines already cite usually lands within days to two weeks, because those domains get crawled constantly. New pages on your own site typically take several weeks, sometimes longer for a small site with little crawl history. Technical fixes, like publishing an llms.txt file or unblocking AI crawlers in robots.txt, we verify directly by fetching your site on the next scan.
Your pooled score then moves deliberately, not instantly. It averages a rolling window of scans, so a genuine improvement shows up partially at first and firms up as more scans agree. That is the same stability guarantee that stops a single lucky answer from reading as progress. If you want to watch a change land sooner, the per question view and the raw single scan score both move faster than the pooled headline.
In practice: claiming a profile or getting listed somewhere the engines already cite usually lands within days to two weeks, because those domains get crawled constantly. New pages on your own site typically take several weeks, sometimes longer for a small site with little crawl history. Technical fixes, like publishing an llms.txt file or unblocking AI crawlers in robots.txt, we verify directly by fetching your site on the next scan.
Your pooled score then moves deliberately, not instantly. It averages a rolling window of scans, so a genuine improvement shows up partially at first and firms up as more scans agree. That is the same stability guarantee that stops a single lucky answer from reading as progress. If you want to watch a change land sooner, the per question view and the raw single scan score both move faster than the pooled headline.
8. Your data
We query the engines only about your brand and the questions you choose. We don't sell your data, share your tracked queries or dashboards, or use your data to train any model. See our privacy policy.
Questions about the method? Get in touch , we'll answer in detail.