A beautifully designed, sixty-page "GEO analysis report" lands on your desk. Before you sign anything: how do you tell whether it deserves to drive real decisions?
The question matters more this year than ever, because AI has pushed the cost of producing a long report toward zero. Sixty pages no longer means ten times the work of six. Page count and rigor have officially parted ways โ not a bad thing in itself (we draft with AI too), but it does mean "looks professional" is no longer a usable signal. You need harder tests.
Here are five. None of them require a technical background, all of them can be asked to a vendor's face, and good answers are easy to tell from bad ones.
Question 1: Is the data measured, or compiled?
A report about your brand's visibility in AI answers should be built around actual measurements: real buyer questions, sent to real AI engines, with a record of whether your brand came up, whether competitors did, and which sources the AI cited.
Ask the vendor for three things: the prompt list (which questions were asked), the engine list (which AI engines were asked), and the sampling window (when). Anyone who actually ran the measurement can produce all three without hesitation.
Bad sign: the report is thick, but nowhere in it is a single real AI answer. "AI visibility" appears in the title and the tagline, while the body is market sizing, channel playbooks and platform tips โ a digest of publicly available material.
Compiled research isn't worthless. It just cannot answer the one question you are actually paying to have answered: where does my brand stand, right now, inside AI answers?
Question 2: If someone else re-ran it, would the conclusions survive?
Reproducibility is the line between measurement and rhetoric.
A measurement with an open methodology can be re-run by a third party โ same prompts, same engines โ and land in roughly the same place (AI answers fluctuate, but conclusions shouldn't flip). Which suggests a wonderfully efficient question:
"If I asked an independent party to re-run this with your methodology, would you mind?"
A vendor who doesn't mind, and can hand over the methodology, is at least measuring in good faith. A vendor who hedges, or claims the algorithm is a trade secret, is asking you to take the numbers on faith.
For what an open methodology looks like in practice, see how we score โ not because everyone must do it our way, but because "open enough to be re-run" is a standard any serious vendor can meet. Not meeting it is a choice.
Question 3: Is the scoring methodology public?
Reports love composite scores: a fraction out of five, star ratings next to each competitor. Scores are fine. Scores without a methodology are not.
Three concrete follow-ups: How is a perfect score defined? What separates one band from the next? What is the denominator?
If those have answers, the score is a measurement. If they don't, the score is rhetoric โ and the job of rhetoric is not to inform you but to make the problem feel severe enough that the attached solution sells itself.
A related trick: point at any percentage in the report and ask what its denominator is. A percentage with no denominator should be read as an adjective.
Question 4: Who verifies the targets?
Reports usually end with targets: top-tier within so many months, followers multiplied, leads multiplied. Before asking "is this achievable", ask three more basic things:
- Who measures it when the time comes โ the vendor doing the work, or someone else?
- Against what methodology โ the same prompt list and engine list from Question 1?
- What happens if it's missed โ is that in the contract?
A target that fails all three isn't a commitment. It's atmosphere.
The healthier pattern: a measured baseline first, then a re-measurement schedule, with the verification standard written into the contract. Whether things improved โ and why โ gets decided by the same ruler, applied twice. That's how we deliver reports ourselves: we never promise what a score will reach; we promise you'll know clearly whether it moved and why.
Question 5: Did a human read this before it shipped?
Drafting with AI is fine; we do it too. Shipping without a single human read-through is not.
Three signals you can check in minutes:
- Template placeholders: blanks like "XX stores" or "XX followers" still sitting in the body text.
- Self-contradicting numbers: the market size in the executive summary disagreeing with the body โ sometimes by orders of magnitude.
- Unverified key facts: a certification used as a selling point throughout, with its registration number marked "to be confirmed".
This isn't pedantry. If nobody proofread the deliverable, nobody stress-tested the strategy inside it either. The check costs five minutes; a vendor who didn't spend them has told you what the deliverable was worth internally.
Hiring help is fine โ self-graded homework is the problem
Let's be clear about one thing: needing a team to handle content, channels and publishing end-to-end is a legitimate need, and full-service execution is a legitimate business. There are execution teams that do serious work.
What deserves your caution is not "full-service" but a specific structure: the people doing the work grading their own results. When player and referee are the same firm, every report converges on the same conclusion โ your situation is dire, fortunately our service addresses it, please renew next quarter.
The healthier structure separates execution from verification: execution goes to a team that's good at executing, verification goes to an independent party with no stake in the execution contract. This isn't distrust of the executor โ quite the opposite. Once the referee is independent, the player's good results finally mean something. For an execution team that does real work, an independently issued scorecard is worth more than any self-written summary.
The five questions, to go
- Is the data measured? โ ask for the prompt list, engine list and sampling window.
- Would the conclusions survive an independent re-run? โ is the methodology open enough to re-run.
- Is the scoring methodology public? โ how is full marks defined; what's the denominator.
- Who verifies the targets? โ who measures, against what, and what if it's missed.
- Did a human read it before it shipped? โ placeholders, contradicting numbers, unverified facts.
No technical background required โ and these five filter out most reports that have thickness but no measurement.
We're OrcaScope, a ruler and nothing else: we measure and verify AI visibility, we don't run campaigns and we don't publish content, and our scoring methodology is public in full on the methodology page. To see where your brand stands in AI answers today, the free checkup takes thirty seconds, no signup.