How we measure

Methodology version 2026-07-21-v2

Cite42 measures how AI models rank brands in your category when asked buyer questions. We query each provider through its official API with that provider’s native web search enabled, and we read Google AI Overviews from live search results. We do not scrape consumer chat interfaces.

The surfaces

SurfaceAPI idHow we query it
OpenAIchatgptThe provider's official API, with its own native web-search tool enabled and retrieval left on automatic. The model decides per prompt whether to search, so evergreen questions are answered from its own knowledge and only time-sensitive ones trigger a live search. That is how a normal chat session behaves too.
AnthropicclaudeThe provider's official API, with its own native web-search tool enabled and retrieval left on automatic. The model decides per prompt whether to search, so evergreen questions are answered from its own knowledge and only time-sensitive ones trigger a live search. That is how a normal chat session behaves too.
PerplexityperplexityAn engine that searches the web on every call by design.
Google GeminigeminiThe provider's official API, with its own native web-search tool enabled and retrieval left on automatic. The model decides per prompt whether to search, so evergreen questions are answered from its own knowledge and only time-sensitive ones trigger a live search. That is how a normal chat session behaves too.
Google AI Overviewsgoogle_ai_overviewA live read of Google's search results, which is the production surface itself.

Every response includes a measurement block naming the provider, the access method, and the methodology version behind the numbers, so a report can always show where its data came from. We do not publish the specific model version behind each surface. That is an operational choice we change for cost and latency, and the methodology version above tells you whether two runs are comparable.

What each measurement claims

  • Rankings: which of the brands you name get mentioned, how often, and in what order. This is the one we are most confident in.
  • Compare: your brand’s standing against the specific competitors you name, computed the same way as rankings.
  • Sentiment: how positively or negatively the answers describe one brand.
  • Citations: which sources a provider’s web search surfaces for a prompt. Use it to find which domains own your category in AI retrieval, and where your content is absent. It is not a record of the citation links a consumer chat UI shows a logged-in user. The limits below explain why.

One measurement per tracker, on purpose. A brand mention is not a recommendation, a cited source is not a shortlist placement, and positive language is not visibility. Folding all of it into a single score would make the number harder to explain and easier to misread.

What this does not tell you

  • No tool can show you what a specific person saw. AI answers vary with login state, stored memory, custom instructions, country, and reasoning mode. Published analysis of ChatGPT’s own modes found only about a quarter of cited sources in common between its fast and thinking modes for the same prompts. A scraped anonymous browser session is one unpersonalized configuration, not your buyer’s screen. What carries over between runs is relative standing, measured the same way each time.
  • Citations diverge most between API and interface. Independent comparisons put source-level overlap between API responses and chat-interface answers in the single digits, and a large share of what a model asserts comes from its training data with no search result behind it at all. That is why our citations measurement claims the retrieval corpus and not the interface. If URL-level fidelity to a chat UI is your hard requirement, Cite42 is the wrong tool for it.
  • Answers are non-deterministic. The same prompt can return different brands on different runs. We query each prompt once per surface per run, so a single result is one sample rather than a verdict. Reliability comes from breadth, meaning more prompts and more surfaces, which is why trackers hold up to 30 prompts and run across five surfaces instead of repeating one prompt.
  • We cache answers for 24 hours, shared across customers. An identical prompt on the same surface can replay a stored answer from up to a day earlier. This keeps per-call pricing low. It also means two runs less than 24 hours apart are not independent observations. Scheduled trackers run weekly or monthly, which is always outside the cache window.
  • Coverage is not a ranking. Cite42 does not publish an official or permanent AI rank. It reports what the selected surfaces returned for your saved prompts on the runs you paid for.

Why weekly and monthly, not daily

Published variance analysis of brand answers finds that repeated sampling of one prompt is among the largest sources of movement in a score, while the identity of the model answering it is among the smallest. Day-over-day deltas on a single sample are therefore mostly noise, at roughly seven times the cost of a weekly run. Trackers run weekly or monthly, and a one-off check is available on demand when you have changed something and want to see it now.

Changes to this methodology

The methodology version changes whenever a provider model, retrieval policy, or scoring input changes. Stored results record the version they were produced under, so a change never silently rewrites your history, and you can see when a comparison crosses a version boundary instead of guessing. The docs go into more detail on each measurement.