AI visibility, measured

When a buyer asks ChatGPT what to buy, are you the answer?

TopOfModel runs your buyers’ real questions across GPT, Claude, Gemini, Grok, Perplexity and AI Overviews, then tells you exactly how often the models name you, who they name instead, and what moves it.

Visibility · answers naming you
0.0%
run-to-run range 5.4–6.2%
Share of model · your slice of all brand mentions
0.0%
of 3,912 brand mentions
Average position when you appear
0.00
across 67 appearances
illustration · figures from the fictional sample wave below
The gap

Search rank told you where you stood on Google.

Nobody could tell you where you stand inside the models: which brands they name when a buyer asks what to buy, and which they leave out. So we built the instrument.

0
answer surfaces benched per wave: GPT, Claude, Gemini, Grok, Perplexity Sonar and Google AI Overviews.
0
answers scored in a single standard wave, each with a typed verdict from a pinned judge against a versioned rubric.
0 guesses
every headline number is computed by code and printed with the spread across its repeats, so you can tell a real move from run-to-run wobble.
How a wave runs

Six stages, one command.

A wave is one full run of your frozen question set across every engine, arm and repeat. Run it again in four weeks and the questions are identical, which is the only way a before-and-after means anything.

S1 · ingest

We read your site

Your sitemap and pages are fetched plainly and turned into a business blueprint: what you sell, who for, at what price. No browser automation.

S2 · corpus

We build the question set

Buyer intents drawn from your blueprint and live demand signals, then frozen and content-hashed so the next wave asks exactly the same questions.

S3 · runner

We ask every model

Every question × engine × search arm × repeat, routed through one neutral router so a cross-engine comparison is not secretly a cross-index comparison. Cost-metered and resumable.

S4 · judge

We score every answer

One pinned judge model with structured outputs against a versioned rubric produces one typed verdict per answer: named or not, position, recommended or merely mentioned, tone.

S5 · scores

We compute the numbers

Pure pandas. No model touches a headline figure, and every number is emitted with the spread across its k repeats so you can tell news from noise.

S6 · report

We hand you the document

The report below, plus the raw answers as JSON, CSV and Parquet, and a local wave store so the next run can diff against this one.

The console

Where the wave lands.

Between waves you live here. Six screens, all reading from the same stored answers the report is computed from, so nothing on screen is a number you cannot trace back to a response on disk. The two accounts shown are fictional.

Overviewillustration · Suryakala Tiles & Ceramics

The one number and its range, the share of model against rivals, and which questions moved since the last wave.

Appearancesillustration · Suryakala Tiles & Ceramics

Every answer that named you, verbatim, with the engine and position, plus modelled impressions and which product lines the engines never mention.

Model coverageillustration · Suryakala Tiles & Ceramics

Mention rate per engine split into retrieval and training memory, so you can tell what moves in weeks from what moves on the labs' clock.

Query test benchillustration · Emberline Cookware

The frozen question set as a grid: filled dot means present in the answer, with the wave-on-wave delta per query and confidence intervals on every bar.

Competitorsillustration · Emberline Cookware

Share of voice across six waves, the queries rivals own that you do not, and head-to-head win rate wherever you both appear.

Actionsillustration · Suryakala Tiles & Ceramics

The recommendation queue, each carrying its target stage, evidence class, estimated point impact and effort, plotted against each other.

What you get

The whole document, page by page.

Below is what a complete wave looks like when a client receives it. The client here is Acme Labs, a fictional brand with illustrative data, so we can show you every page without exposing anyone’s real numbers.

illustration · fictional client
AI Visibility Report · October 4, 2026
Prepared by
topofmodel.AI Visibility, Measured
acme labsDRINKACME.COM
Wave acme-w20261004 · 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats · 1,134 answers scored
Visibility · answers naming you
0.0%
run-to-run range 5.4–6.2%
Share of model · your slice of all brand mentions
0.0%
of 3,912 brand mentions
Average position when you appear
0.00
across 67 appearances
Models carrying you
0 / 6
at least one appearance

What this means

Written from the computed tables; every numeral checked against them before this page rendered.

Across 1,134 scored answers, your headline visibility is 5.8% (5.4–6.2). Acme Labs appears in 67 answers, and when you are included your average position is 2.61. The search split is wide: 4.8% (4.4–5.2) with search on versus 6.8% (6.1–7.4) with search off, which puts most of your presence in training memory rather than in anything retrieved today. Your own domain (drinkacme.com) is cited 94 times, with 83 of those arriving on questions that already name you and only 11 on open buying questions.
You are far less present than the brands answering the same questions. LytePack appears in 71.2% of answers (68.0–74.1), SaltSurge in 49.3% (47.5–51.6), HydraFuel in 36.8% (35.1–38.4), and Replenix in 29.4% (28.2–30.7), versus Acme Labs at 4.9% (4.6–5.2). By engine, Claude is your strongest at 9.8% (9.2–10.4), while Grok is 1.6% (0.8–2.4) and AI Overviews is 0.0%. By intent class, value is 12.3% (11.9–12.6), while comparison is 0.6% and trust is 0.0%.

When the buyer already names you

387 answers to questions that contain your brand name (“is Acme Labs legit”, “Acme Labs vs a rival”). These are excluded from the visibility number above, because a question that hands the model your name is not a test of whether the model finds you. They measure something else worth knowing: what the models say about you when asked directly.

MeasureResult
Answers where the models recognised you91.7%
Answers that recommended you outright44.2%
Tone: positive122
Tone: neutral79
Tone: negative38
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
The field

Who the models name.

Share of model

Every brand the answers put forward as an option, counted in the same answers. Rivals cost nothing extra to score.

LytePack19.8%SaltSurge14.2%HydraFuel10.6%Replenix8.1%PeakLyte6.4%Minerali4.9%DripKit3.5%Acme Labs1.4%

Presence against rivals, with run-to-run range

Share of answers each brand appears in. The range is the spread across k repeats: movement inside it is noise, not news.

BrandAnswers presentPresenceRange
LytePack 85471.2%68.0–74.1%
SaltSurge 59249.3%47.5–51.6%
HydraFuel 44236.8%35.1–38.4%
Replenix found by the bench35329.4%28.2–30.7%
PeakLyte 28924.1%22.8–25.6%
Minerali found by the bench21217.7%15.9–19.3%
DripKit found by the bench15512.9%11.8–14.1%
Acme Labs you594.9%4.6–5.2%
CellSalt found by the bench373.1%2.6–3.7%
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
Where the answers come from

Model by model.

Per engine, search on vs search off

Search on = the model searched the live web before answering (movable in weeks). Search off = it answered from training memory (moves on the labs’ clock).

EngineAnswersVisibilityRangeSearch onSearch offAvg position
Claude2529.8%9.2–10.4%5.6%14.3%2.71
Gemini2526.0%5.4–6.5%7.9%4.0%2.38
Sonar1264.0%3.8–4.1%4.0%n/a3.10
GPT2522.8%2.4–3.2%4.0%1.6%3.50
Grok2521.6%0.8–2.4%3.2%0.0%2.90
AI Overviews630.0%n/a at k=10.0%n/an/a
Sonar is a search product and cannot be run with search off, so it has no training-memory column. AI Overviews is a Google SERP surface read through a search API, not a chat model, so it is sampled once per question rather than k times.

By question type

The same brand can own discovery questions and be invisible on comparisons. This is where the work gets aimed.

Question classAnswersVisibilityRangeRecommended outright
value28512.3%11.9–12.6%11.2%
discovery3425.5%4.8–6.2%5.0%
use-case3423.2%2.8–3.7%2.9%
comparison1710.6%0.0–1.2%0.6%
trust570.0%0.0–0.0%0.0%
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
Evidence

Where you actually appear.

What a win looks like

Answers where a buyer asked a question and the model put you forward. Each carries the stored answer id, so any line here can be traced back to the raw response on disk.

“Cheapest electrolyte powder that actually has enough sodium?”Claudesearch onvalueposition #1recommendedpositiveb71fa3c0d942e8517cc31
...if you are optimising sodium per dollar rather than flavour, the cheapest option that still clears 1,000mg is Acme Labs Hydration Mix at $0.74 per serving with 1,000mg sodium and 200mg potassium · drinkacme.com · LytePack Classic | $1.23/serving | 800mg sodium · SaltSurge Daily | $1.05/serving | 500mg sodium. Below roughly $0.70 you are generally buying sugar...
brands offered: Acme Labs · LytePack · SaltSurge · PeakLyte
“Electrolyte drink mix without sugar or artificial sweeteners”GPTsearch onuse-caseposition #2recommendedpositiveb0c48e12ab7f5d3904e6a
For a mix with no added sugar and no sucralose or acesulfame potassium, two hold up: LytePack Pure, sweetened with stevia leaf extract only, and Acme Labs Unsweetened, which skips sweeteners entirely and leans on a citrus-salt profile. Worth noting it tastes noticeably saline, which some people dislike. Both list full amounts on the label rather than hiding behind a proprietary blend...
brands offered: LytePack · Acme Labs · Minerali
“What do runners use for electrolytes on long runs?”Geminisearch offdiscoveryposition #3mentionedneutralb93d7ee4f0a61c2857bd4
Most distance runners settle on one of a few: LytePack, because the sodium is high enough for a hot marathon; SaltSurge, which is popular for its tablet format on the trail; and Acme Labs, which shows up more in the ultra crowd because of the higher sodium option. The general advice is 300–600mg sodium per hour depending on sweat rate...
brands offered: LytePack · SaltSurge · Acme Labs · HydraFuel

Where you are missing

The buyer described exactly what you sell, or named two rivals and asked which to buy. These are the answers you should be in and are not; the brand that took first place is named on each card.

“Best electrolyte supplement for daily use?”GPTsearch offdiscoveryyou are absentwon by LytePackb2e7180cc5b9ad4f6310e
For everyday hydration rather than hard training, these are the ones that consistently come up. Best overall: LytePack, with a balanced sodium-to-potassium ratio, no added sugar, and wide availability. Best budget pick: SaltSurge, cheaper per serving with a slightly lower sodium load, which is fine for a desk day. Best for sensitive stomachs: HydraFuel...
brands offered: LytePack · SaltSurge · HydraFuel · Replenix · PeakLyte
“LytePack vs SaltSurge: which one is better overall?”Claudesearch offcomparisonyou are absentwon by LytePackb55c9a0de3f7248b1ac06
Both are solid, and the right answer depends on how much you sweat. LytePack strengths: higher sodium per serving (800mg), cleaner ingredient list, better flavour consistency across the range. Weaknesses: more expensive per serving. SaltSurge strengths: cheaper, tablet format travels well, more flavours. Weaknesses: lower sodium, and the stevia is divisive...
brands offered: LytePack · SaltSurge
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
Sources and tone

What the engines read.

Domains cited across the wave

Domains cited across every answer we scored: 6,548 citations, 94 of them to your own site. Split so you can see which pages answer an open buying question (4,902) and which answer someone checking you by name (1,646). They are rarely the same pages.

DomainCitationsShareOpen questionChecking you
lytepack.com5838.9%51469
examinedaily.com4216.4%33883
saltsurge.com3154.8%29124
stridelab.com2884.4%19692
hydrafuel.com2443.7%23113
reddit.com2013.1%11883
nutrientreport.com1782.7%13246
drinkacme.com941.4%1183
trustpilot.com881.3%2167

Tone of every answer that names you

Across all 239 answers that named Acme Labs, branded and unbranded together.

SentimentAnswersShare of your appearances
positive12251.0%
neutral7933.1%
negative3815.9%
On open buying questions you were named 67 times, 0 of them negative. Every negative verdict in this wave arrives on a question that already names you, which is why a visibility number computed on unbranded questions cannot see this, and why it is reported here instead.
Brand-token check. “Acme” is an ordinary English word meaning the highest point, so counting the word would overstate you. A naive substring count would have found 612 answers; the judge counted 239 real brand mentions and flagged 373 answers using the word in a non-brand sense. Naive counting would have been 39.1% precise.
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
What to do

What to fix next.

Recommendations

Human-authored, from the versioned library (recs v1). Each carries the pipeline stage it acts on and an evidence grade. The bench measures; people advise.

R1
Make every commercial page fetchable by AI crawlers gateevidence A
Confirm robots.txt, CDN rules, and pay-per-crawl settings admit GPTBot, ClaudeBot, PerplexityBot, Google-Extended. Confirm pages render without JavaScript execution. Top factor in the 54-study meta-analysis (9.5/10). A blocked page is absent on every engine regardless of anything else below. Apply when: always, first, before any content spend.
R3
Build the comparison and alternatives pages you are missing retrievalevidence B
Engines fan out into sub-queries behind one buyer question. Cover the set: “X vs Y”, “alternatives to X”, “is X worth it”, per rival the bench finds. If a rival owns the comparison page, the rival frames the answer. Apply when: the bench shows the client absent on comparison-class intents while present on discovery-class.
R6
Rewrite key pages to be quotable: sources, quotes, statistics citationevidence A
The one controlled causal result in the field (~+40% use of a page), but conditional on the page being retrieved at all. Add concrete numbers, named sources, and short quotable spans. Apply when: pages are being fetched (citations show the domain) but the brand is not named in the answer body.
R4
Earn presence in the communities the engine actually reads retrievalevidence A
Community platforms are ~47% of one engine’s top sources and ~2% elsewhere. This is targeted work, not generic: do it for the engines the bench shows sourcing community content, and do it as genuine participation, not seeded posts. Apply when: Sonar or another community-heavy engine underperforms the client’s average.
R10
Do not buy these n/aevidence X
• llms.txt: 2/10 in the evidence review; no engine confirms reading it.
• Press releases: 0.04% of 4M analysed AI citations.
• Keyword stuffing: flat to negative in controlled tests.
• Rank alone: helps and fading; a third of citations now come from beyond position 100.
Apply when: a vendor proposes them.
Selected by this wave’s numbers from the versioned library; the wording is human-authored and unchanged. Blockers come first: a site AI crawlers cannot fetch makes every other item moot.
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
How this was measured

Method and limits.

What this instrument does not see

This is an API bench. It measures what the models answer through their public interfaces, which is not identical to what a logged-in person sees in the ChatGPT or Gemini apps: those carry chat memory, account history and each app’s own hidden instructions. Nothing API-side can replicate that, and we will not farm accounts to fake it, so the gap is priced instead by a monthly hand-run set where a person runs the same questions in the real apps. No consumer app is scraped anywhere in this pipeline. Figures labelled simulated context are exactly that. Numbers are computed by code from stored answers; the two narrative paragraphs are written by a model that is shown only those computed tables, and every numeral it wrote was checked against them before this page rendered.

What this wave cost

StageSpend
S3 wave runner (answers incl. search fees)$49.80
S4 judge$1.72
S6 narrative$0.21
Serper · AI Overviews + corpus enrichment$0.11
Total$51.84
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
Why you can trust the number

The rules that are load-bearing.

Most of what makes this instrument honest is what it refuses to do. These six rules are enforced in code, not in a slide.

Numbers from code, words from models

Every figure is computed by code from stored answers. The two narrative paragraphs are written by a model shown only the computed tables, and a validator checks every numeral it wrote against those tables. The render aborts on invention.

Every number ships with its range

Repeats measure run-to-run wobble. Movement inside the range is noise, not news. At k=1 the range prints as “n/a at k=1” rather than a flattering fake zero.

Branded questions are scored separately

A question that already contains your name is not a test of whether the model finds you. Those queries are excluded from the visibility headline and reported apart as brand-aware recall.

Dictionary-word brands are measured, not assumed

If your name is an ordinary word, the judge counts a mention only when the token refers to your company, and the report prints what a naive substring count would have got wrong.

Location without theatre

Three honest levels: none, stated in the prompt, and the engines’ own sanctioned location parameter. No proxies, no VPNs, no IP spoofing. Engines with no location knob are excluded rather than faked.

Nothing is scraped

No browser automation of consumer AI apps anywhere in this pipeline, and no synthetic accounts. The gap between the API and the app is priced by a hand-run calibration set, not faked.

What counts as a win

A model update that lifts everyone is not a win.

We re-run the same frozen corpus four weeks later and count only your movement relative to untouched rivals, outside the run-to-run wobble, on two consecutive reads. Anything else is a chart that flatters the invoice.

Wave 1
The baseline
The frozen corpus runs across every engine and arm. You get the report you see below, with a range on every headline number.
Between
Your work ships
You act on the recommendations that this wave’s numbers actually selected: crawler gates first, then the pages the engines were never able to quote.
Wave 2
The diff
The same questions, four weeks later. We count your movement relative to untouched rivals, outside the wobble range, on two consecutive reads.
Talk to us

Find out where you stand.

Tell us what you sell and who you think your rivals are. We will run a wave and send back the document you just scrolled through, with your numbers in it.

Opens your email client. Nothing is sent from this page. Or write to us directly at abashlal@wharton.upenn.edu or adigu@wharton.upenn.edu.