Public experiments
Most claims about how unstable AI answers are get passed around second-hand. We measured it on our own production data instead — the method, the sample sizes, and the numbers we don't yet consider conclusive are all below.
How this was measured: 263 prompts · 8 AI engines · 35 sampling days (most recent Thu Aug 06). All of it comes from our own nightly production runs, not a one-off experiment staged for this page. No brand names or tenant information appear — only aggregates.
Experiment 1 · How much engines disagree on the same question
n = 2,490Taking samples where the same prompt on the same day returned results from at least 4 engines, we check whether those engines agree about a given brand. Only the 2,490 samples where at least one engine mentioned the brand are counted — cases where nobody mentioned it aren't disagreement.
When at least one AI engine mentions a brand, this is how often the others disagree.
n = 2,490
This explains something buyers in this space run into constantly: different tools return different answers for the same prompt. Digiday (May 2026) quoted one directly — “run the same prompt through three tools and you get three different answers.” The number above suggests much of that divergence isn't in the tooling: **the engines themselves disagree**. A single manual search gives you one engine at one moment, not your visibility.
Experiment 2 · Does yesterday's mention survive to today
n = 50,647Holding the prompt and the engine fixed and varying only the date, we compare consecutive days — 50,647 adjacent-day pairs in total.
among samples mentioned the previous day
across all 50,647 pairs
The number on the left is the most useful thing on this page: same prompt, same engine, one day later — 15.6% of the brands mentioned yesterday are gone today. That isn't “your ranking dropped”, it's **the baseline moving underneath you**. So no single measurement is a conclusion; judging whether an optimisation worked requires a long enough window and a control group that wasn't touched — otherwise you're most likely crediting noise.
What we are deliberately not concluding
We can't say which engine is “more accurate”. Disagreement only shows they differ; none of the answers is a reference answer, and we hold no ground truth.
We can't extrapolate to the whole internet. The sample comes from prompts our clients actually monitor; the industry and language mix is not random.
We can't attribute every flip to the engines. Our own detection has error too; we only count cells where detection completed, but we have not hand-audited them.
The numbers above update automatically with each nightly run — nothing on this page is hand-written; changing the method means changing the published SQL function. Read the full methodology →