No competitor explains their formula. We do. Every component of the visibility score, why we give a range instead of a single number, and the limits of the measurement — it's all here, we hide nothing.
The score measures more than whether a brand "appears" in AI answers: it also weighs where it appears, how much space it takes up relative to competitors, and how it is described.
0.45 × mention_rate + 0.25 × position + 0.20 × share of voice + 0.05 × sentiment + 0.05 × recommendation strengthHow often the brand appears in the answer across tracked runs. This is the heaviest signal: if the AI never mentions you, everything else is secondary. "I appear in half the answers" means a rate of 0.5.
How high up you are mentioned in the answer. Being at the top of a list is not the same as appearing at the end — we score position as 1/rank; appearing first earns full marks, appearing lower proportionally less.
Your share among all brands mentioned in the same answer. If an answer contains 10 brands and one is you, your share of voice is low; if it's you and one competitor, it's high. It contextualizes your visibility by competitor density.
How positively the brand is framed where it is mentioned. Appearing is good, but appearing as "not recommended" is different. Its weight is kept deliberately low: perception is an important signal but not as decisive as visibility itself.
Added so that merely being "listed" and being "clearly recommended" in the answer aren't counted the same. A direct recommendation scores highest, being in a list is medium, hesitant/conditional phrasing lower, and negative phrasing lowest.
The formula lives in one place and is versioned: when it changes, past days' scores can be recalculated with the same formula, so your trend comparisons don't break. We calibrate the weights with real data and update this page if there's a significant change.
AI answers are stochastic — the same question can be answered slightly differently today and tomorrow. A score based on a single run gives a misleading sense of precision.
We base the score not on a single day but on a sliding window formed by the recent runs for each prompt × platform. We measure the general trend of the recent period, not one day's random result. This naturally smooths out daily noise.
We show a lower and upper bound alongside the score. We produce it with bootstrapping: we resample the runs on hand a thousand (1000) times, derive the distribution of our score estimate, and present the 10% and 90% percentiles as a band. The band conveys how settled the measurement we have is: if it's narrow the estimate is reliable, if it's wide more runs are needed.
It's important to read this band honestly: the band shows how settled the current measurement is — it is not a prediction interval that guarantees tomorrow's score. The band itself does not fully capture the intra-day randomness of AI answers; we reduce that separately by accumulating more runs in the sliding window. So a narrow band means "the measurement has settled, the estimate is reliable", and a wide band means "more data is needed". When deciding, you look not at a single number but at the band and the trend; when the band is meaningfully wide we show the score together with it — for an unexaggerated, defensible measurement.
We take ChatGPT (OpenAI), Gemini (Google), Claude (Anthropic) and Grok (xAI) measurements through these providers' official APIs; Google AI Overviews is measured via DataForSEO (third party). This is not identical to the interface a consumer sees in the chat screen — it is a strong proxy. There can be differences in model version, system prompt and personalization between the API and the interface.
Google AI Overviews is a different category — and we say so openly. Google does not offer an official/billed API for AI Overviews answers. We measure this engine through a third-party search-data service called DataForSEO — that is, indirectly; DataForSEO gathers this data from Google search results. So of the five engines, four (ChatGPT, Gemini, Claude, Grok) are official APIs and the fifth is a reliable third-party source.
Instead of hiding this difference we state it openly throughout the product. We also say openly why proxy measurement is still valuable: trend, comparison and perception signals are consistent and comparable because they're collected the same way every day, under the same conditions. So you can confidently read not the last digit of the absolute number, but the change over time and the position relative to competitors.
In short: we don't tell you "you'll see exactly this on ChatGPT's screen". We say "when we ask the same question the same way every day, your brand looks like this and changes like this" — and that is what you need to take action.
If two different projects track the same question (for example "what is the best accounting software?"), we ask the AI that question once that day and use the resulting answer for both projects. We call this the intersection pool.
There are two reasons for this. First, cost: running the same question over and over is a needless API expense, and it would show up in our prices. Second and more importantly, consistency: everyone tracking the same question is measured against exactly the same raw answer, so cross-project comparisons stay fair and comparable. We re-analyze the brands within the answer from your own perspective for your brand — only the raw run is shared, your score is unique to you.
A good measurement is one that knows its limits. What you should keep in mind:
LLMs are stochastic. The same question can be answered differently even on the same day. That's why we look at the sliding window and confidence band rather than a single run — but there's no such thing as zero fluctuation; small daily shifts are normal.
Proxy measurement is not identical to your interface. Because we measure through the API, there can be deviations from the screen a consumer sees. Trust the trend and comparison more than the absolute number.
P2/P3 prompts don't run every day. Lower-priority prompts are scanned weekly or monthly; their score is based on the most recent runs, so it updates more slowly.
Brand recognition isn't flawless. We resolve the brand names in an answer with fuzzy matching; we tolerate spelling variations but in rare cases may miss a name or over-count. We continuously tighten the rules to systematically reduce such deviations.
Providers answer at different levels of detail. For example, Gemini on average counts noticeably more brands/entities in an answer than the other providers. This can be misleading when comparing "share of voice" (SoV) directly across providers — your share may look low on a more "talkative" provider, which doesn't mean your brand has weakened. Read share of voice not across providers but by its change over time within the same provider.
Google AI Overviews is not an official API. As explained above, we measure this engine through DataForSEO, via a third-party service — it is not a direct, contracted API like the other four engines. If that source has an outage, AIO measurement may briefly gap; the other engines are unaffected. See Proxy measurement.
You know the methodology — now see it work on your own brand. 3 days free, no credit card required.
ChatGPT, Gemini, Claude and Grok measurements are through official APIs; Google AI Overviews is done via DataForSEO (third party).