Methodology
How we measure AI visibility — and what we refuse to claim
Three tools can give a brand three different AI-visibility answers, and the usual reason is method, not data. This page is the method: what runs, how often, how it aggregates, and where we deliberately hold back. Agencies are welcome to share it with their clients.
What we measure
- Five answer engines: Perplexity, ChatGPT, Gemini, Claude, and Google AI Overviews. Every scan asks a real question and records the full answer.
- Two facts per answer: was the brand mentioned in the text, and was its site cited as a source. A citation weighs more than a mention.
- Prompts are chosen per site — a buyer-journey mix of informational and commercial questions, editable by the agency before anything runs.
Why we sample instead of trusting one run
- AI engines do not answer identically twice. A single scan is a snapshot, not a score — treating it as a score is how tools end up contradicting each other.
- We run each prompt repeatedly on a schedule and report the rate: the share of sampled answers that mentioned or cited the brand over a rolling window.
- Every rate carries a 95% confidence interval. Small samples say "collecting" — they are never dressed up as verdicts.
How we call a change real
- A movement counts as significant only when the confidence intervals of the two periods separate — bigger than sampling noise, not just different.
- Before/after impact of published content splits the prompt's run stream at the publish date and compares the two sides with the same test.
- Engines that errored during a run are excluded as unmeasured. An outage is not evidence a brand was left out of answers.
Leading and lagging, labeled
- Citation-rate trend, AI crawler activity and competitor share move first — we report them as leading indicators.
- Traffic and leads follow, typically on a 60–90 day delay — we report them as lagging, and never promise them early.
- AI referral counts are a measured floor: parts of AI traffic strip the referrer, so the true number is higher, never lower. We say so wherever a count appears.
What we refuse to report
- No fabricated placeholders: when data is missing, the surface says what is missing and what fills it.
- No universal conversion multipliers: the evidence for AI-referral conversion is vertical-dependent, and we present it that way.
- No single-run "scores" as headline numbers, no engines we did not run, and no significance claims the intervals do not support.
Questions about the method?
The measurement statistics — sampling, intervals, significance — run identically for every account, on every plan. If something here doesn't match what you see in the product, that is a bug and we want to know. Contact us.
