Skip to main content
Methodology · Reviewed

A visibility score is a dice roll unless you show the distribution

Ask ChatGPT the same buyer question twice and you can get two different answers. Most tools report whichever one they happened to catch. Collimer runs the probe enough times to report a calibrated proportion and the confidence interval around it, so you know what's signal and what's noise.

Visibility score yourcompany.com
58 ±7

Mid

Cited in 6/10 answers across every engine on your plan. The interval is the credibility: your score drifts as models update.

What a scan actually does

Engines
ChatGPT, Claude, Gemini, Perplexity, Grok (by plan tier)
Probes per scan
Dozens of real buyer questions, per engine
Runs per probe
Repeated: the same engine answers differently twice
Interval
Wilson score interval on the citation proportion
Refresh
Models drift; we re-scan and report the delta

Why the interval is the point

The confidence interval is the credibility signature. It tells you whether a change between two scans is real or within the noise. When we recommend a fix and you ship it, the before/after delta is only meaningful against the interval. A move from 47 ±8 to 50 ±8 is noise; a move to 58 ±7 is signal. We will never report a single fixed visibility number as if it were exact.

Limitations we'll say out loud

  • Model providers change behavior without notice; your score drifts and we show how much.
  • Engines ground differently: Perplexity always cites; ChatGPT triggers web search on a minority of queries. We report per engine, not a false average.
  • Citation overlap between engines is low; being cited by one is not being cited by all.
  • A scan measures presence in answers, not downstream conversion. It's the leading indicator, not the sale.

Want the evidence behind the recommendations? The guides cite the research each lever is built on.