When a buyer asks ChatGPT what to buy, are you the answer?
TopOfModel runs your buyers’ real questions across GPT, Claude, Gemini, Grok, Perplexity and AI Overviews, then tells you exactly how often the models name you, who they name instead, and what moves it.
illustration · figures from the fictional sample wave below
The gap
Search rank told you where you stood on Google.
Nobody could tell you where you stand inside the models: which brands they name when a buyer asks what to buy, and which they leave out. So we built the instrument.
0
answer surfaces benched per wave: GPT, Claude, Gemini, Grok, Perplexity Sonar and Google AI Overviews.
0
answers scored in a single standard wave, each with a typed verdict from a pinned judge against a versioned rubric.
0 guesses
every headline number is computed by code and printed with the spread across its repeats, so you can tell a real move from run-to-run wobble.
How a wave runs
Six stages, one command.
A wave is one full run of your frozen question set across every engine, arm and repeat. Run it again in four weeks and the questions are identical, which is the only way a before-and-after means anything.
S1 · ingest
We read your site
Your sitemap and pages are fetched plainly and turned into a business blueprint: what you sell, who for, at what price. No browser automation.
S2 · corpus
We build the question set
Buyer intents drawn from your blueprint and live demand signals, then frozen and content-hashed so the next wave asks exactly the same questions.
S3 · runner
We ask every model
Every question × engine × search arm × repeat, routed through one neutral router so a cross-engine comparison is not secretly a cross-index comparison. Cost-metered and resumable.
S4 · judge
We score every answer
One pinned judge model with structured outputs against a versioned rubric produces one typed verdict per answer: named or not, position, recommended or merely mentioned, tone.
S5 · scores
We compute the numbers
Pure pandas. No model touches a headline figure, and every number is emitted with the spread across its k repeats so you can tell news from noise.
S6 · report
We hand you the document
The report below, plus the raw answers as JSON, CSV and Parquet, and a local wave store so the next run can diff against this one.
The console
Where the wave lands.
Between waves you live here. Six screens, all reading from the same stored answers the report is computed from, so nothing on screen is a number you cannot trace back to a response on disk. The two accounts shown are fictional.
Suryakala Tiles & CeramicsBuilding materials · India
AppearancesSuryakala Tiles & Ceramics
Wave · 6 · Sep 11 ▾
Table view
Run new wave
Answers containing you · Wave 6
214/ 600
▲ 41 vs Wave 5 · intent × model checks
Est. monthly impressions modeled
1.6M
query volume × platform share × mention rate
Models carrying you
5 of 6
Grok still parametric-blind to you
Product lines appearing
3 of 5
▼ bathroom + outdoor lines absent
Latest appearances
Verbatim answer excerpts from the bench · your brand highlighted · full log in Table view
"best vitrified tiles for living room india"Geminiposition #2PositiveSep 11
…for living rooms, contractors most often shortlist Kalinga Ceramics and Suryakala; Suryakala's double-charged vitrified range is noted for holding polish in high-traffic rooms…
"floor tiles that don't stain"Sonarposition #1PositiveSep 11
…Suryakala's nano-sealed vitrified tiles are repeatedly recommended on housing forums for stain resistance, ahead of pricier imported options…
Suryakala Tiles & CeramicsBuilding materials · India
Model coverageSuryakala Tiles & Ceramics
Wave · 6 · Sep 11 ▾
Table view
Run new wave
Coverage by model · Wave 6
Mention rate on the bench · 3 phrasings per intent · k = 5 runs each · ± 4 pts
MODEL
MENTION RATE
AVG POS
SENTIMENT
RETRIEVAL / PARAMETRIC
6 WAVES
Gemini · Google
58%
2.4
Positive
65/35
Sonar · Perplexity
52%
2.9
Positive
100/0
GPT · OpenAI
44%
2.8
Positive
70/30
Claude · Anthropic
37%
3.2
Neutral
55/45
DeepSeek · DeepSeek
31%
3.6
Neutral
45/55
Grok · xAI
12%
4.5
Neutral
80/20
Retrieval visibility · moves in weeksParametric visibility · moves with training
Read
What the split says
Gemini and Sonar carry us. Both lean on live retrieval, where the Aug interventions land fastest. Sonar is 100% retrieval: every gain there is earned by pages, not priors.
Grok is a parametric problem. 80% of its answers come from priors that predate us. No quick fix: it moves with editorial footprint over quarters, not weeks. Priced into the plan, not promised.
DeepSeek splits the difference. Watch it as the control: if retrieval-led work moves Gemini but not DeepSeek, the lift is real, not drift.
Where our mentions come from
Source types behind answers that cite Suryakala · Wave 6
Editorial + blogs
54%
Community + forums
19%
Marketplaces
11%
Own site
9%
News
7%
Press releases: 0% of cited sources, matching the evidence base. Community share is why the Sonar number leads.
Method
Judge model pinned for this engagement (v2026-08-01). Browse-on and browse-off runs are scored separately; the split column reports where each model got its answer. Mention rates carry ± 4 pts at k = 5 with 3 phrasings per intent. Every number on this screen is also in the Table view export.
Model coverageillustration · Suryakala Tiles & Ceramics
Mention rate per engine split into retrieval and training memory, so you can tell what moves in weeks from what moves on the labs' clock.
Filled dot = present in answer · Wave 6, k = 5 majority across 3 phrasings · Δ = models gained vs Wave 3
QUERY
GPT
GEMINI
CLAUDE
GROK
DEEPSEEK
SONAR
Δ W3
best non-toxic cookware 2026
▲ +2
ceramic vs stainless cookware which lasts
▲ +3
cookware without PFAS that sears well
▲ +2
emberline vs halewood pans
·
best cookware set under $400
▲ +1
induction-safe ceramic pans
▲ +2
cookware brands made in usa
▲ +1
replace nonstick pan how often
·
wedding registry cookware picks
▲ +2
is ceramic coating safe at high heat
▼ -1
Coverage by query class · Wave 3 vs Wave 6
Whiskers are k = 5 confidence intervals · interventions Aug 12 + Aug 21
Wave 3 · Aug 1Wave 6 · Sep 11
Reading this screen
Wave 3 is the pre-intervention baseline. Aug 12: comparison hub live. Aug 21: editorial campaign first placements. The 19-point coverage move clears both intervals, and the gains concentrate in the two query classes the interventions targeted. Full 100-query table exports from the Table view.
Query test benchillustration · Emberline Cookware
The frozen question set as a grid: filled dot means present in the answer, with the wave-on-wave delta per query and confidence intervals on every bar.
Emberline at the baseline for readability: 11% → 19% across six waves. Halewood flat, Stonebriar drifting down.
Gap queries
They appear, we do not · ranked by category volume
QUERY
WHO OWNS IT
US
dutch oven worth it which brand
Halewood, Stonebriar
cookware that lasts 20 years
Halewood
best pan for glass top stove
Stonebriar, Calder + Main
non stick alternatives that work
Halewood, Vessel & Vine
Each gap query feeds an action card on the Actions screen.
Head to head · win rate
When we both appear, how often we outrank them
vs Halewood
31%
vs Stonebriar
44%
vs Calder + Main
58%
vs Vessel & Vine
66%
Read
Where the share is moving
The gain is real and attributable. +8 pts of share over six waves, with the two Aug interventions marked on the bench. Halewood held; the share came out of Other and Stonebriar.
Halewood owns durability queries. Three of four gap queries are durability-shaped. That is the next editorial front: the statistics rewrite on the Actions screen targets it.
Competitorsillustration · Emberline Cookware
Share of voice across six waves, the queries rivals own that you do not, and head-to-head win rate wherever you both appear.
Durability claims rewritten with statistics + sources
Generation · GEO study, up to +40%
+1 to 3 pts
S
✓ Done
P3
llms.txt request from vendor
No measured effect
0 pts
S
✓ Declined
8 of 14 recommendations shown · full list with owners and due dates in the Table view · impact estimates re-fit after every wave.
Done
3
+4 pts landed
In progress
2
land by Wave 8
Queued
3
est. +7 to 12 pts
Impact vs effort
Estimates from the evidence base · sized to this account
Sequence: everything upper-left first. llms.txt plotted where the evidence puts it.
Actionsillustration · Suryakala Tiles & Ceramics
The recommendation queue, each carrying its target stage, evidence class, estimated point impact and effort, plotted against each other.
What you get
The whole document, page by page.
Below is what a complete wave looks like when a client receives it. The client here is Acme Labs, a fictional brand with illustrative data, so we can show you every page without exposing anyone’s real numbers.
illustration · fictional client
AI Visibility Report · October 4, 2026
Prepared by
topofmodel.AI Visibility, Measured
Wave acme-w20261004 · 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats · 1,134 answers scored
Visibility · answers naming you
0.0%
run-to-run range 5.4–6.2%
Share of model · your slice of all brand mentions
0.0%
of 3,912 brand mentions
Average position when you appear
0.00
across 67 appearances
Models carrying you
0 / 6
at least one appearance
What this means
Written from the computed tables; every numeral checked against them before this page rendered.
Across 1,134 scored answers, your headline visibility is 5.8% (5.4–6.2). Acme Labs appears in 67 answers, and when you are included your average position is 2.61. The search split is wide: 4.8% (4.4–5.2) with search on versus 6.8% (6.1–7.4) with search off, which puts most of your presence in training memory rather than in anything retrieved today. Your own domain (drinkacme.com) is cited 94 times, with 83 of those arriving on questions that already name you and only 11 on open buying questions.
You are far less present than the brands answering the same questions. LytePack appears in 71.2% of answers (68.0–74.1), SaltSurge in 49.3% (47.5–51.6), HydraFuel in 36.8% (35.1–38.4), and Replenix in 29.4% (28.2–30.7), versus Acme Labs at 4.9% (4.6–5.2). By engine, Claude is your strongest at 9.8% (9.2–10.4), while Grok is 1.6% (0.8–2.4) and AI Overviews is 0.0%. By intent class, value is 12.3% (11.9–12.6), while comparison is 0.6% and trust is 0.0%.
When the buyer already names you
387 answers to questions that contain your brand name (“is Acme Labs legit”, “Acme Labs vs a rival”). These are excluded from the visibility number above, because a question that hands the model your name is not a test of whether the model finds you. They measure something else worth knowing: what the models say about you when asked directly.
Measure
Result
Answers where the models recognised you
91.7%
Answers that recommended you outright
44.2%
Tone: positive
122
Tone: neutral
79
Tone: negative
38
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
The field
Who the models name.
Share of model
Every brand the answers put forward as an option, counted in the same answers. Rivals cost nothing extra to score.
Presence against rivals, with run-to-run range
Share of answers each brand appears in. The range is the spread across k repeats: movement inside it is noise, not news.
Brand
Answers present
Presence
Range
LytePack
854
71.2%
68.0–74.1%
SaltSurge
592
49.3%
47.5–51.6%
HydraFuel
442
36.8%
35.1–38.4%
Replenix found by the bench
353
29.4%
28.2–30.7%
PeakLyte
289
24.1%
22.8–25.6%
Minerali found by the bench
212
17.7%
15.9–19.3%
DripKit found by the bench
155
12.9%
11.8–14.1%
Acme Labs you
59
4.9%
4.6–5.2%
CellSalt found by the bench
37
3.1%
2.6–3.7%
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
Where the answers come from
Model by model.
Per engine, search on vs search off
Search on = the model searched the live web before answering (movable in weeks). Search off = it answered from training memory (moves on the labs’ clock).
Engine
Answers
Visibility
Range
Search on
Search off
Avg position
Claude
252
9.8%
9.2–10.4%
5.6%
14.3%
2.71
Gemini
252
6.0%
5.4–6.5%
7.9%
4.0%
2.38
Sonar
126
4.0%
3.8–4.1%
4.0%
n/a
3.10
GPT
252
2.8%
2.4–3.2%
4.0%
1.6%
3.50
Grok
252
1.6%
0.8–2.4%
3.2%
0.0%
2.90
AI Overviews
63
0.0%
n/a at k=1
0.0%
n/a
n/a
Sonar is a search product and cannot be run with search off, so it has no training-memory column. AI Overviews is a Google SERP surface read through a search API, not a chat model, so it is sampled once per question rather than k times.
By question type
The same brand can own discovery questions and be invisible on comparisons. This is where the work gets aimed.
Question class
Answers
Visibility
Range
Recommended outright
value
285
12.3%
11.9–12.6%
11.2%
discovery
342
5.5%
4.8–6.2%
5.0%
use-case
342
3.2%
2.8–3.7%
2.9%
comparison
171
0.6%
0.0–1.2%
0.6%
trust
57
0.0%
0.0–0.0%
0.0%
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
Evidence
Where you actually appear.
What a win looks like
Answers where a buyer asked a question and the model put you forward. Each carries the stored answer id, so any line here can be traced back to the raw response on disk.
“Cheapest electrolyte powder that actually has enough sodium?”Claudesearch onvalueposition #1recommendedpositiveb71fa3c0d942e8517cc31
...if you are optimising sodium per dollar rather than flavour, the cheapest option that still clears 1,000mg is Acme Labs Hydration Mix at $0.74 per serving with 1,000mg sodium and 200mg potassium · drinkacme.com · LytePack Classic | $1.23/serving | 800mg sodium · SaltSurge Daily | $1.05/serving | 500mg sodium. Below roughly $0.70 you are generally buying sugar...
“Electrolyte drink mix without sugar or artificial sweeteners”GPTsearch onuse-caseposition #2recommendedpositiveb0c48e12ab7f5d3904e6a
For a mix with no added sugar and no sucralose or acesulfame potassium, two hold up: LytePack Pure, sweetened with stevia leaf extract only, and Acme Labs Unsweetened, which skips sweeteners entirely and leans on a citrus-salt profile. Worth noting it tastes noticeably saline, which some people dislike. Both list full amounts on the label rather than hiding behind a proprietary blend...
brands offered: LytePack · Acme Labs · Minerali
“What do runners use for electrolytes on long runs?”Geminisearch offdiscoveryposition #3mentionedneutralb93d7ee4f0a61c2857bd4
Most distance runners settle on one of a few: LytePack, because the sodium is high enough for a hot marathon; SaltSurge, which is popular for its tablet format on the trail; and Acme Labs, which shows up more in the ultra crowd because of the higher sodium option. The general advice is 300–600mg sodium per hour depending on sweat rate...
The buyer described exactly what you sell, or named two rivals and asked which to buy. These are the answers you should be in and are not; the brand that took first place is named on each card.
“Best electrolyte supplement for daily use?”GPTsearch offdiscoveryyou are absentwon by LytePackb2e7180cc5b9ad4f6310e
For everyday hydration rather than hard training, these are the ones that consistently come up. Best overall: LytePack, with a balanced sodium-to-potassium ratio, no added sugar, and wide availability. Best budget pick: SaltSurge, cheaper per serving with a slightly lower sodium load, which is fine for a desk day. Best for sensitive stomachs: HydraFuel...
“LytePack vs SaltSurge: which one is better overall?”Claudesearch offcomparisonyou are absentwon by LytePackb55c9a0de3f7248b1ac06
Both are solid, and the right answer depends on how much you sweat. LytePack strengths: higher sodium per serving (800mg), cleaner ingredient list, better flavour consistency across the range. Weaknesses: more expensive per serving. SaltSurge strengths: cheaper, tablet format travels well, more flavours. Weaknesses: lower sodium, and the stevia is divisive...
brands offered: LytePack · SaltSurge
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
Sources and tone
What the engines read.
Domains cited across the wave
Domains cited across every answer we scored: 6,548 citations, 94 of them to your own site. Split so you can see which pages answer an open buying question (4,902) and which answer someone checking you by name (1,646). They are rarely the same pages.
Domain
Citations
Share
Open question
Checking you
lytepack.com
583
8.9%
514
69
examinedaily.com
421
6.4%
338
83
saltsurge.com
315
4.8%
291
24
stridelab.com
288
4.4%
196
92
hydrafuel.com
244
3.7%
231
13
reddit.com
201
3.1%
118
83
nutrientreport.com
178
2.7%
132
46
drinkacme.com
94
1.4%
11
83
trustpilot.com
88
1.3%
21
67
Tone of every answer that names you
Across all 239 answers that named Acme Labs, branded and unbranded together.
Sentiment
Answers
Share of your appearances
positive
122
51.0%
neutral
79
33.1%
negative
38
15.9%
On open buying questions you were named 67 times, 0 of them negative. Every negative verdict in this wave arrives on a question that already names you, which is why a visibility number computed on unbranded questions cannot see this, and why it is reported here instead.
Brand-token check. “Acme” is an ordinary English word meaning the highest point, so counting the word would overstate you. A naive substring count would have found 612 answers; the judge counted 239 real brand mentions and flagged 373 answers using the word in a non-brand sense. Naive counting would have been 39.1% precise.
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
What to do
What to fix next.
Recommendations
Human-authored, from the versioned library (recs v1). Each carries the pipeline stage it acts on and an evidence grade. The bench measures; people advise.
R1
Make every commercial page fetchable by AI crawlersgateevidence A
Confirm robots.txt, CDN rules, and pay-per-crawl settings admit GPTBot, ClaudeBot, PerplexityBot, Google-Extended. Confirm pages render without JavaScript execution. Top factor in the 54-study meta-analysis (9.5/10). A blocked page is absent on every engine regardless of anything else below. Apply when: always, first, before any content spend.
R3
Build the comparison and alternatives pages you are missingretrievalevidence B
Engines fan out into sub-queries behind one buyer question. Cover the set: “X vs Y”, “alternatives to X”, “is X worth it”, per rival the bench finds. If a rival owns the comparison page, the rival frames the answer. Apply when: the bench shows the client absent on comparison-class intents while present on discovery-class.
R6
Rewrite key pages to be quotable: sources, quotes, statisticscitationevidence A
The one controlled causal result in the field (~+40% use of a page), but conditional on the page being retrieved at all. Add concrete numbers, named sources, and short quotable spans. Apply when: pages are being fetched (citations show the domain) but the brand is not named in the answer body.
R4
Earn presence in the communities the engine actually readsretrievalevidence A
Community platforms are ~47% of one engine’s top sources and ~2% elsewhere. This is targeted work, not generic: do it for the engines the bench shows sourcing community content, and do it as genuine participation, not seeded posts. Apply when: Sonar or another community-heavy engine underperforms the client’s average.
R10
Do not buy thesen/aevidence X
• llms.txt: 2/10 in the evidence review; no engine confirms reading it.
• Press releases: 0.04% of 4M analysed AI citations.
• Keyword stuffing: flat to negative in controlled tests.
• Rank alone: helps and fading; a third of citations now come from beyond position 100.
Apply when: a vendor proposes them.
Selected by this wave’s numbers from the versioned library; the wording is human-authored and unchanged. Blockers come first: a site AI crawlers cannot fetch makes every other item moot.
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
illustration · fictional client
How this was measured
Method and limits.
What this instrument does not see
This is an API bench. It measures what the models answer through their public interfaces, which is not identical to what a logged-in person sees in the ChatGPT or Gemini apps: those carry chat memory, account history and each app’s own hidden instructions. Nothing API-side can replicate that, and we will not farm accounts to fake it, so the gap is priced instead by a monthly hand-run set where a person runs the same questions in the real apps. No consumer app is scraped anywhere in this pipeline. Figures labelled simulated context are exactly that. Numbers are computed by code from stored answers; the two narrative paragraphs are written by a model that is shown only those computed tables, and every numeral it wrote was checked against them before this page rendered.
What this wave cost
Stage
Spend
S3 wave runner (answers incl. search fees)
$49.80
S4 judge
$1.72
S6 narrative
$0.21
Serper · AI Overviews + corpus enrichment
$0.11
Total
$51.84
Illustration · fictional client, illustrative data · Method: 20 intents × 3 phrasings · 5 model families via one neutral router · k = 2 repeats per cell · search-on and search-off as separate passes · answers scored by a pinned judge · corpus frozen before the run · ranges are the k-repeat spread.
topofmodel.
Why you can trust the number
The rules that are load-bearing.
Most of what makes this instrument honest is what it refuses to do. These six rules are enforced in code, not in a slide.
Numbers from code, words from models
Every figure is computed by code from stored answers. The two narrative paragraphs are written by a model shown only the computed tables, and a validator checks every numeral it wrote against those tables. The render aborts on invention.
Every number ships with its range
Repeats measure run-to-run wobble. Movement inside the range is noise, not news. At k=1 the range prints as “n/a at k=1” rather than a flattering fake zero.
Branded questions are scored separately
A question that already contains your name is not a test of whether the model finds you. Those queries are excluded from the visibility headline and reported apart as brand-aware recall.
Dictionary-word brands are measured, not assumed
If your name is an ordinary word, the judge counts a mention only when the token refers to your company, and the report prints what a naive substring count would have got wrong.
Location without theatre
Three honest levels: none, stated in the prompt, and the engines’ own sanctioned location parameter. No proxies, no VPNs, no IP spoofing. Engines with no location knob are excluded rather than faked.
Nothing is scraped
No browser automation of consumer AI apps anywhere in this pipeline, and no synthetic accounts. The gap between the API and the app is priced by a hand-run calibration set, not faked.
What counts as a win
A model update that lifts everyone is not a win.
We re-run the same frozen corpus four weeks later and count only your movement relative to untouched rivals, outside the run-to-run wobble, on two consecutive reads. Anything else is a chart that flatters the invoice.
Wave 1
The baseline
The frozen corpus runs across every engine and arm. You get the report you see below, with a range on every headline number.
Between
Your work ships
You act on the recommendations that this wave’s numbers actually selected: crawler gates first, then the pages the engines were never able to quote.
Wave 2
The diff
The same questions, four weeks later. We count your movement relative to untouched rivals, outside the wobble range, on two consecutive reads.
Talk to us
Find out where you stand.
Tell us what you sell and who you think your rivals are. We will run a wave and send back the document you just scrolled through, with your numbers in it.