🚀 OrcaScope private beta — first 10 design partners get 3 months free Join now!
OrcaScope

Public experiments

Most claims about how unstable AI answers are get passed around second-hand. We measured it on our own production data instead — the method, the sample sizes, and the numbers we don't yet consider conclusive are all below.

How this was measured: 263 prompts · 8 AI engines · 35 sampling days (most recent Thu Aug 06). All of it comes from our own nightly production runs, not a one-off experiment staged for this page. No brand names or tenant information appear — only aggregates.

Experiment 1 · How much engines disagree on the same question

n = 2,490

Taking samples where the same prompt on the same day returned results from at least 4 engines, we check whether those engines agree about a given brand. Only the 2,490 samples where at least one engine mentioned the brand are counted — cases where nobody mentioned it aren't disagreement.

63.7%

When at least one AI engine mentions a brand, this is how often the others disagree.

n = 2,490

This explains something buyers in this space run into constantly: different tools return different answers for the same prompt. Digiday (May 2026) quoted one directly — “run the same prompt through three tools and you get three different answers.” The number above suggests much of that divergence isn't in the tooling: **the engines themselves disagree**. A single manual search gives you one engine at one moment, not your visibility.

Experiment 2 · Does yesterday's mention survive to today

n = 50,647

Holding the prompt and the engine fixed and varying only the date, we compare consecutive days — 50,647 adjacent-day pairs in total.

15.6%
Mentioned yesterday, gone today

among samples mentioned the previous day

5.2%
Verdict flipped day-over-day (either direction)

across all 50,647 pairs

The number on the left is the most useful thing on this page: same prompt, same engine, one day later — 15.6% of the brands mentioned yesterday are gone today. That isn't “your ranking dropped”, it's **the baseline moving underneath you**. So no single measurement is a conclusion; judging whether an optimisation worked requires a long enough window and a control group that wasn't touched — otherwise you're most likely crediting noise.

What we are deliberately not concluding

We can't say which engine is “more accurate”. Disagreement only shows they differ; none of the answers is a reference answer, and we hold no ground truth.

We can't extrapolate to the whole internet. The sample comes from prompts our clients actually monitor; the industry and language mix is not random.

We can't attribute every flip to the engines. Our own detection has error too; we only count cells where detection completed, but we have not hand-audited them.

The numbers above update automatically with each nightly run — nothing on this page is hand-written; changing the method means changing the published SQL function. Read the full methodology