How to tell whether a GEO fix actually worked
A single re-check cannot distinguish a real visibility change from run-to-run noise. The honest method: freeze the prompt panel, re-probe with a confidence interval, and report one of five outcomes, including the null.
By Nick Bair
A single re-check cannot tell you whether a GEO fix worked. AI answers vary run to run, and a one-point score comparison cannot distinguish a real change from ordinary noise. The honest method is to freeze your prompt panel before the change, re-probe after with enough repetition to carry a confidence interval, and only call a delta real when the two intervals stop overlapping. Everything else is reading tea leaves.
Why can’t one re-check tell me if my AI visibility improved?
Because AI engines do not return the same answer twice. The same prompt, sent to the same engine, one minute apart, can produce different citations, different phrasing, and a different set of sources. That variance is structural; you cannot engineer around it. A score built from a single run has no error bar, which means you cannot tell whether a two-point move reflects your fix or the engine’s mood that afternoon.
The analogy that holds: a single temperature reading does not tell you whether the weather changed. You need enough readings, taken consistently, to see a trend that sits outside the measurement noise.
What does “the same measurement” actually mean?
A score is only comparable over time against a fixed set of prompts, a fixed engine set, and a fixed repetition count. Change any of those and you are not comparing two measurements of the same thing. You are comparing two different instruments.
Think of it the way a market index works: the basket of stocks is fixed so that price changes reflect the market, not a change in what you are measuring. Your prompt panel is that basket. If you add a prompt between run one and run two, the delta is contaminated. If one engine goes dark mid-comparison, the delta is contaminated. If you ran five repetitions before and three after, the delta is contaminated.
This matters practically. If you are tracking five engines and one stops responding during your post-fix probe, you cannot pool the partial results and call it a score. That engine’s absence must be flagged, not averaged away. See what to do if one AI engine goes dark mid-scan for how to handle that case without discarding the run.
The frozen panel is not a nice-to-have. It is the measurement instrument. Without it, you have an impression, not a result.
How wide is the noise floor, really?
On Collimer’s own 100-prompt, five-engine panel, a single run’s score carries a confidence interval of roughly plus or minus 7 points, so a two-point move is indistinguishable from no move at all (Collimer’s panel re-probes, August 2026). Across four August re-probes of that identical panel, the point score itself drifted less than one point, which is exactly what a stable brand should look like: the interval is wide, the trend is flat, and no single reading means much on its own.
Panel size narrows the floor. An earlier 285-prompt version of the same panel carried an interval closer to plus or minus 3.5 points. More prompts and more repetitions buy you a tighter instrument; a 10-prompt spot check buys you almost nothing.
What does that imply for a two-point move? It implies nothing. A move has to be larger than the interval itself before the pre-fix and post-fix readings stop overlapping, and only at that point is the delta real. If you call every two-point move a win, you will also call every two-point drop a crisis, and you will spend your week reacting to variance instead of shipping fixes.
One more number worth holding: in one 548,000-page study, only 15% of the pages ChatGPT retrieved were ultimately cited (as of June 2026). “They crawled it” is not “they cite it.” A fetch is not a mention. A mention is not a citation. Each step has its own drop-off, and a score that conflates them will mislead you about where the real gap is.
What are the five honest outcomes of a verification?
After you re-probe with a frozen panel and enough repetitions to carry an interval, there are exactly five things you can honestly report:
- Not shipped. The change is not yet live or not yet crawlable. The re-probe is premature. Wait and re-probe once the page is confirmed indexed.
- Pending. The change is live but the confidence intervals still overlap. The delta exists but is not yet distinguishable from noise. Re-probe in one week.
- Verified moved. The post-fix interval sits entirely above the pre-fix interval. The delta is real. This is the outcome you are working toward.
- No measurable change. The intervals overlap completely and the point estimate has not shifted. This is a finding, not a failure. It tells you the fix did not move this panel, which is information: either the fix was wrong, the panel does not cover the prompts where the fix should show up, or the lag window was too short.
- Moved against. The post-fix interval sits entirely below the pre-fix interval. The fix made things worse. This happens. Publish it, diagnose it, and ship a correction.
One receipt from our own design-partner work: after shipping a README-and-docs wave, a Series B devtools company’s citation rate moved from 1.25% to 5.83% in six weeks on its frozen panel. The intervals separated. Outcome: verified moved. One brand, one panel; a receipt, not a promise.
“No measurable change” is the outcome most founders resist publishing. Resist the resistance. A null result on a well-run panel is evidence. It rules out the fix as the cause of any future move, which narrows the search space for what to try next.
How long after shipping should I re-probe?
Two to four weeks is the practical window for most fixes, but the right answer depends on the engine and the type of change.
Engines publish little about their re-crawl schedules. In our own re-probes, a shipped page change typically takes one to three weeks to surface across engines, and the spread between the fastest and slowest engine is wide. If you re-probe at 72 hours, you are likely measuring the pre-fix state with a post-fix label on it.
Per-engine differences set the clock. If your panel spans multiple engines, the earliest valid re-probe date is set by the slowest engine in the set, not the fastest. Probing before the slowest engine has re-indexed means your panel is internally inconsistent, which puts you back in the contaminated-delta problem.
The practical rule: wait at least two weeks, confirm the page is indexed in each engine before you run the post-fix probe, and do not start the clock from when you shipped. Start it from when the last engine confirmed the crawl.
For a fuller treatment of how to measure across engines without conflating their different behaviors, see how to measure AI visibility properly, and the methodology page for how we compute the interval itself.
What can I do this week without a tool?
Pick ten buyer prompts you want to win. Ask each of three engines three times. Record cited, mentioned, or absent per run in a spreadsheet. That is your pre-fix baseline. Freeze it. Do not add prompts. Do not swap engines.
Ship your fix. Wait two weeks. Run the identical sheet again: same prompts, same engines, same three repetitions each. Count the outcomes per prompt. Anything that changes in fewer than two-thirds of runs is noise. Anything that changes in two-thirds or more of runs is a candidate result, not a confirmed one, but worth a second look.
That is 90 data points per wave, collected in under an hour, with no tool required. It will not give you a confidence interval in the statistical sense, but it will give you a directional read that is honest about its own limits. The two-thirds threshold is a practical proxy for “this is not random.”
The one thing you cannot do manually at scale is freeze the panel automatically, track engine availability, and compute overlapping intervals across dozens of prompts. That is where a tool earns its place. But the manual version is better than a single re-check, and you can start it today.
A score change should only be treated as real once the confidence intervals of the two measurements stop overlapping. A move smaller than the noise floor is not a result. Publish your nulls. Name your outcomes. And before you ship the next fix, run a free scan; it freezes the panel, so the re-check two weeks later means something.
Provider behavior changes quickly. This guide reflects what we know as of August 2026; we update it when the evidence shifts. The interval figures are from Collimer’s own panel re-probes; external numbers carry their sources because we believe in showing the work.
Measure where you stand.
Run a free scanRelated guides
-
How to explain AI visibility to your CEO in one slide
One slide, four lines: where you are cited today per buyer question, the three fixes shipping this month, what you honestly expect to move and by when, and how you will know it worked. No pooled score, no competitor rank, no revenue forecast.
-
Check your own AI visibility from Claude or Cursor: the MCP workflow
You can run an AI-visibility scan on any domain from inside Claude Code, Claude Desktop, Cursor, Windsurf, or VS Code. Collimer ships an MCP server with exactly one tool, installed with a five-line config block. Below: the setup, three prompts that work verbatim, and an unedited transcript of a real session, including our own unflattering score.
-
What GEO costs in 2026: DIY, tool, agency, or sprint, and how to choose
GEO costs time before it costs money. The honest ranges: $0 and 2 to 3 hours a week for DIY, $20 to $800 a month for a tracker, $1,500 to $25,000 a month for an agency retainer, and a fixed-scope sprint that ends cleanly in between. The route that fits depends on one question: can you ship a site change yourself this week?