A visibility score is a dice roll unless you show the distribution
Ask ChatGPT the same buyer question twice and you can get two different answers. Most tools report whichever one they happened to catch. Collimer runs the probe enough times to report a calibrated proportion and the confidence interval around it, so you know what's signal and what's noise.
Mid
Cited in 6/10 answers across every engine on your plan. The interval is the credibility: your score drifts as models update.
What a scan actually does
- Engines
- ChatGPT, Claude, Gemini, Perplexity, Grok (by plan tier)
- Probes per scan
- Dozens of real buyer questions, per engine
- Runs per probe
- Repeated: the same engine answers differently twice
- Interval
- Wilson score interval on the citation proportion
- Refresh
- Models drift; we re-scan and report the delta
Why the interval is the point
The confidence interval is the credibility signature. It tells you whether a change between two scans is real or within the noise. When we recommend a fix and you ship it, the before/after delta is only meaningful against the interval. A move from 47 ±8 to 50 ±8 is noise; a move to 58 ±7 is signal. We will never report a single fixed visibility number as if it were exact.
Limitations we'll say out loud
- Model providers change behavior without notice; your score drifts and we show how much.
- Engines ground differently: Perplexity always cites; ChatGPT triggers web search on a minority of queries. We report per engine, not a false average.
- Citation overlap between engines is low; being cited by one is not being cited by all.
- A scan measures presence in answers, not downstream conversion. It's the leading indicator, not the sale.
Want the evidence behind the recommendations? The guides cite the research each lever is built on.