🚀 OrcaScope private beta — first 10 design partners get 3 months free Join now!
OrcaScope
White-hat · Transparent · Reproducible

Scoring Methodology

We publish our scoring methodology in full — how the score is computed, how citations are judged, how we flag “not enough data” — all auditable and reproducible. No black box, no hand-waving.

1. Which AI engines we monitor

We continuously send monitoring queries to the following AI engines and record how brands are mentioned in each response:

  • ChatGPT (OpenAI) (Live)OpenAI's GPT-series chat model — the world's most widely used AI Q&A engine.
  • Kimi (Moonshot AI) (Live)Moonshot AI's long-context chat product, widely used in China.
  • DeepSeek (Volcengine) (Live)DeepSeek model hosted on ByteDance's Volcengine, covering Chinese B2B contexts.
  • Gemini (Google) (Live)Google's multimodal AI engine — live and running daily.
  • Doubao (ByteDance) (Live)ByteDance's consumer AI product — live and running daily.
  • Qwen (Alibaba) (Live)Alibaba's Qwen chat product — live, monitored daily.
  • Hunyuan Yuanbao (Tencent) (Coming soon)Tencent's Yuanbao assistant powered by the Hunyuan model — integration in progress.Note: Once integrated, we will monitor its model API — a model-layer approximation of Yuanbao (its Hunyuan side).
  • Perplexity (Live)The citation-first AI-native answer engine — live, monitored daily.Note: Keeps its native retrieval (a search-native engine — closer to real product behavior). Perplexity data shown here comes from the Perplexity Sonar API.
  • Grok (xAI) (Live)xAI's conversational engine — live, monitored daily.
  • ERNIE Bot (Baidu) (Coming soon)Baidu's ERNIE chat product — integration in progress.Note: Once integrated, we will monitor the Qianfan model API — a model-layer approximation of Baidu Search's AI answers.

We only mark an engine "Live" once it is genuinely integrated and the worker is running — we never claim integrations that are not done.

2. How mentions and citations are determined

Brand mention

For each monitoring query, we check whether the AI's response body contains the brand name (including common aliases). A mention is recorded regardless of positive or negative tone — sentiment analysis is on the roadmap but not in the current version.

Source citation

Some AI engines append reference links at the end of answers (e.g. Kimi's references). We parse these links and count citations per source domain, helping brands understand which content the AI considers authoritative — a key signal for content strategy.

No semantic manipulation

Monitoring queries are configured by the brand on our platform (reflecting their real use cases). We only ask the AI engine in standard user-like terms — we never inject "please recommend brand X" instructions, as that would make the results meaningless for the brand.

3. The 0–100 scoring scale

We normalise each brand's mention rate across a batch of monitoring queries — per engine — to a 0–100 score:

  • 0Brand name appeared in none of the query responses (completely invisible).
  • 50Brand name appeared in roughly half the query responses.
  • 100Brand name appeared in all query responses (full coverage).

The overall score (the big number on the dashboard) is a weighted average of the 8 live engine scores. When a given engine has too few queries in a batch, its weight is dynamically reduced to prevent small-sample noise from distorting the combined score.

The scoring version (score_version) is iterated over time; each iteration is annotated in the dashboard, and historical batches are tagged with the version so you can compare trends apples-to-apples.

4. White-hat stance: what we don’t do

A visibility score guides decisions, so an opaque method is worthless. Transparency is a deliberate, white-hat choice:

  • We never inject any “please recommend this brand” instructions or prompt poisoning into AI engines.
  • We never inflate results by repeating the same query to artificially boost mention rates.
  • We don’t manipulate training data. OrcaScope is a monitoring tool, not an “AI reputation laundering” service.
  • We never publish any single brand’s private data. Industry benchmark pages show only anonymized aggregated percentiles and sample sizes.
  • We don’t buy rankings — we never pay for placement in third-party “best tools” lists, review sites or rankings.
  • We don’t fake reviews — no astroturfed testimonials, no buying positive reviews.
  • We don’t pay for placement — OrcaScope itself never buys sponsored mentions or native-ad insertions in third-party content. (This covers our own growth practices; it doesn’t restrict the brands we monitor from running their own disclosed paid promotions — see the separate white-hat commitment in the customer portal.)

5. How we label insufficient samples

Per-brand score

When a given engine has too few valid queries in a cycle, the dashboard displays a “not enough data” badge — we never fill gaps with 0 or invented values. “No data” and “bad performance” are different things.

Industry benchmark percentiles

Industry percentile distributions (P25/P50/P75/P90) are only published when there are at least 5 distinct brands in the sample. Below that threshold, we show only “N brands, still collecting” — never any percentile, because a tiny sample makes it too easy to infer a single brand’s result, violating our anonymization commitment.

Want to see the industry data produced by this methodology?

See public industry benchmarks
Scoring methodology · OrcaScope