Methodology

How we measure AI visibility — and what we refuse to claim

Three tools can give a brand three different AI-visibility answers, and the usual reason is method, not data. This page is the method: what runs, how often, how it aggregates, and where we deliberately hold back. Agencies are welcome to share it with their clients.

What we measure

  • Five answer engines: Perplexity, ChatGPT, Gemini, Claude, and Google AI Overviews. Every scan asks a real question and records the full answer.
  • Two facts per answer: was the brand mentioned in the text, and was its site cited as a source. A citation weighs more than a mention.
  • Prompts are chosen per site — a buyer-journey mix of informational and commercial questions, editable by the agency before anything runs.

Why we sample instead of trusting one run

  • AI engines do not answer identically twice. A single scan is a snapshot, not a score — treating it as a score is how tools end up contradicting each other.
  • We run each prompt repeatedly on a schedule and report the rate: the share of sampled answers that mentioned or cited the brand over a rolling window.
  • Every rate carries a 95% confidence interval. Small samples say "collecting" — they are never dressed up as verdicts.

How we call a change real

  • A movement counts as significant only when the confidence intervals of the two periods separate — bigger than sampling noise, not just different.
  • Before/after impact of published content splits the prompt's run stream at the publish date and compares the two sides with the same test.
  • Engines that errored during a run are excluded as unmeasured. An outage is not evidence a brand was left out of answers.

Leading and lagging, labeled

  • Citation-rate trend, AI crawler activity and competitor share move first — we report them as leading indicators.
  • Traffic and leads follow, typically on a 60–90 day delay — we report them as lagging, and never promise them early.
  • AI referral counts are a measured floor: parts of AI traffic strip the referrer, so the true number is higher, never lower. We say so wherever a count appears.

What we refuse to report

  • No fabricated placeholders: when data is missing, the surface says what is missing and what fills it.
  • No universal conversion multipliers: the evidence for AI-referral conversion is vertical-dependent, and we present it that way.
  • No single-run "scores" as headline numbers, no engines we did not run, and no significance claims the intervals do not support.

Questions about the method?

The measurement statistics — sampling, intervals, significance — run identically for every account, on every plan. If something here doesn't match what you see in the product, that is a bug and we want to know. Contact us.