The run most teams never take, and the only one that tells you whether the second one meant anything.
What the third run is for
The second run gives you one post-change observation. The third gives you a second one, with nothing changed in between — which lets you see how much the number moves on its own.
If run three sits close to run two and both are clearly different from the baseline, you have a result worth acting on. If run three swings back toward the baseline, what you saw in run two was substantially variance, and the honest conclusion is that the change did not demonstrably work.
This is the run that gets skipped, because by run two the team already has the story it wants. Taking it is the single practice that most separates an AEO programme that compounds from one that accumulates unverified folklore.
Recording the outcome honestly
Write down which of three things happened: the change worked, it did not, or the result was inconclusive. Inconclusive is a legitimate and common outcome, and recording it as such is what stops it from being remembered as a success six months later.
Record the effect size too, not just the direction. A change that moved one prompt from absent to mentioned is a different fact from one that moved you from mentioned to recommended across four prompts, and treating them as the same is how a tactic gets over-generalised.
What to do with each outcome
If it worked: do the same class of change on the next-highest-value gap. You have evidence about your specific category, which is worth more than any generic best practice.
If it did not: revisit the diagnosis rather than the tactic. The usual cause is a correct fix applied to the wrong failure mode.
If it was inconclusive: either the effect is real but smaller than your variance, or there is no effect. Both point to the same next step — pick a higher-leverage change rather than repeating this one harder.
Lab: Run 3 of 3
Needs a free SolCrys account (no credit card). Spends run 3 of 3.
Change nothing since run two.
Trigger the run.
Compare all three: baseline, post-change, and confirmation.
Write one paragraph stating worked, did not work, or inconclusive — and the effect size.
Keep that paragraph. It is the first entry in your own evidence base.
If the honest answer is inconclusive, write inconclusive. That record is worth more than a flattering one.
A single AI check is one sample, not a measurement. How many repeated runs and prompts you need before an AI-visibility number is trustworthy, with the published evidence on sample size and confidence intervals.
AI answers are non-deterministic, so a naive GEO test lies to you confidently. The four test-design rules that decide whether your AI-visibility result is signal or noise: measure a rate, test at the buyer's specificity level, test per language, and sort each cited source by the move.
Your CEO wants AI visibility on the dashboard and your gut says vanity metric. It is - if you track raw rank on vanity prompts. Here are the three conditions that turn it into a leading indicator, a CEO-ready reporting template, and an illustrative 100%-vs-0% scenario that proves the point.
AEO measurement is more like survey design than search analytics. The four sources of disagreement between AEO platforms, the five questions that surface them, and how SolCrys answers each one — with explicit uncertainty.
Turn AI answer gaps into governed marketing execution.
Start free with a ChatGPT visibility read, then add multi-engine tracking, Corporate Context governance, and the action-to-result loop when you are ready.