What AI visibility actually is
AI visibility is how often an AI engine names your brand in an answer, and which sources it cites when it does. That is the whole definition. It isn't a position, and it isn't something Google hands you in a report.
The reason it needs its own name is that the surface changed. A ranking tells you where your page sits in a list that a person still has to read. An AI answer replaces the list. If the engine names three tools in your category and you aren't one of them, your position on page one never enters the conversation.
Two very different things get filed under this one heading. The first is whether the model knows you at all, which lives in its training weights and updates when the model does. The second is whether the engine goes out and fetches your page while it's answering, and then cites it. Most dashboards fold both into a single percentage. They are separate problems with separate fixes, and the rest of this piece keeps them apart.
Why AI visibility has no rank number
Ask an engine the same question five times and you can get five different brand lists. Temperature, model routing, a silent version bump, the time of day. None of that is yours to control, and none of it holds still long enough to be called a rank.
So one run is a sample, not a measurement. Getting named in one answer out of one isn't full visibility. It's a coin landing heads once.
This is why Lumina's checker reports a frequency with a Wilson 95% confidence interval instead of a position. Wilson is the interval you want at small sample sizes, because it stays inside the scale where the textbook formula runs off the end of it. Three mentions in three runs comes back as 44% to 100%, which is an honest way of saying the sample is too small to be sure of anything. Five in five tightens it to 57% to 100%. Still not a rank.
Any product handing you a clean position for your brand in ChatGPT is doing one of two things. Either it runs far more repeats than its price suggests, or it's rounding a coin flip and printing it as a number.
Training memory and live retrieval answer different questions
Ask an engine a question with web search switched off and you learn what it remembers. Ask the same question with search switched on and you learn what it can find. Those are two different business problems.
Brand recall sits in the weights. It moves when the model is retrained, which isn't a timeline you influence, and it's why a company founded this year is invisible to the training layer no matter how good its content is. Retrieval is the opposite. It depends on whether a bot can reach your page today, whether the page answers the question, and whether the engine trusts the domain enough to quote it. That one you can fix this week.
Perplexity has no training layer to test. Sending it a no-search request is an error, not an empty result. It's a search product and never answers from weights alone. So the memory-versus-retrieval comparison has three engines in it, not four, and saying so on the page beats showing a fourth column with nothing in it.
The practical read: if you are invisible in the training layer but cited in the retrieval layer, your content is working and your brand is young. If it's the other way round, people know you and your pages aren't reachable or aren't quotable. The second one is a technical problem, and it's the cheaper of the two to fix.
Live audit: four engines, two markets, measured not modelled
Everything below came out of live calls through Lumina's own proxy on 11 and 12 September 2026, billed to a real DataForSEO account. The unedited responses for ChatGPT, Perplexity and the three mentions endpoints are kept in the repo under studies/ai-visibility/. The Claude and Gemini figures come from the run log rather than a saved file, so treat those two as recorded rather than re-checkable.
Four engines, two markets, and a citation graph that changes shape depending on who you ask and where you ask from.
Measured through the Lumina worker against the DataForSEO AI Optimization API. Prices are from the 12 September round (gpt-5.6-luna, claude-haiku-4-5, gemini-3.1-flash-lite, sonar); citation counts are from the 11 September round, which ran gpt-4.1-mini and gemini-2.5-flash-lite. Markets: United States and Germany.
seo tool, YouTube is the most-cited source by a factor of four. Seobility follows at 751, SE Ranking 576, OMR 503, Semrush 438. Reddit lands seventh.vertexaisearch.cloud.google.com redirect. The other three return the real URL. Anything counting domains has to resolve those first, or exclude Gemini and say which.What the engines actually cite
The headline finding is that there is no such thing as the AI citation graph. There is one per engine, and it changes again per market.
Take the archived German AI Overview answers that mention seo tool. The most-cited source across them is youtube.com, with 3,118 mentions. Seobility, the highest-placed vendor, sits at 751. If your plan for that category was a well-written blog post, the thing outranking you for the citation is a video.
Now the same domain across two engines. Ahrefs in the US market picks up 6,167 mentions in Google AI Overviews and 1,950 in ChatGPT. The AI search volume attributed to those answers is 4,056,660 and 56,483. That gap isn't a measure of how good Ahrefs is at ChatGPT. It's a measure of how much bigger the observed Google surface is, and it's why a single blended AI visibility score is close to meaningless.
These counts are from the 11 September comparison round, which is why the model names differ from the price table further down. The two ChatGPT figures are the two saved web-enabled calls.
| Engine | Model tested | Citations | Publisher in URL |
|---|---|---|---|
| Perplexity | sonar | 20 | Yes |
| Gemini | gemini-2.5-flash-lite | 8 | No, redirect |
| Claude | claude-haiku-4-5 | 7 | Yes |
| ChatGPT | gpt-4.1-mini | 4 and 3 | Yes |
Perplexity is the outlier worth knowing about. It returned nearly three times as many sources as Claude, and the 20 annotations on the saved call spread across 17 separate domains instead of clustering on two or three. That breadth, at the lowest grounded price of the four, is what makes it the most useful retrieval probe.
One more thing belongs in any honest write-up of this. The observed-mentions data for ChatGPT covers the United States in English and nothing else. A German location code doesn't return a thin result, it is rejected outright as an invalid parameter. So when a vendor shows you a ChatGPT visibility figure for a German brand, that figure is either modelled or produced by live prompting. It isn't observation.
The four metrics worth tracking
Four numbers cover almost everything people are actually asking when they say they want to track AI visibility.
| Metric | What it answers | How to read it |
|---|---|---|
| Mention rate | How often the engine names you at all | A share of repeated runs, always with an interval. One run tells you nothing. |
| Citation rate | How often it links you as a source | Lower than mention rate, and the one that sends traffic. A brand can be named constantly and cited never. |
| Share of voice | Your mentions against the competitors named beside you | The only one that survives a model update, because everyone moves together. |
| Source profile | Which domains the engine leans on in your category | Tells you where to be published if it isn't going to be your own domain. |
Mention rate and citation rate get confused constantly, and the difference decides your tactics. A model naming you from memory costs you nothing and sends you nothing. A model citing your page puts a link in front of a reader. If you only ever measure the first, you will conclude things are going well while nobody arrives.
Share of voice is the metric I would report to anyone who has to justify a budget. Absolute mention counts wobble every time a model ships. Your position relative to the four competitors the engine names in the same breath is far steadier, and it's the comparison a reader is making anyway.
Why AI visibility tools cost what they cost
The pricing in this category looks arbitrary until you measure the calls. Then it stops looking arbitrary. Prices below are per request, measured on 12 September against the models the checker now pins.
| Engine | Model | Without web search | With web search |
|---|---|---|---|
| ChatGPT | gpt-5.6-luna | $0.0015 | $0.0129 |
| Claude | claude-haiku-4-5 | $0.0024 | $0.0238 |
| Gemini | gemini-3.1-flash-lite | $0.0018 | $0.0439 |
| Perplexity | sonar | Not possible | $0.0062 |
Reading a model's memory is close to free. Making it go and search costs roughly nine to twenty-four times as much, and the money goes on the search rather than the tokens, so reaching for a smaller model saves you nothing. Multiply that by a handful of repeats, several prompts and four engines, and you arrive at why list prices in this space start around $25 a month and climb past $250.
There is a second cost nobody warns you about, and I only know it because I shipped the bug. A parameter a model refuses fails the request at zero charge. When I pinned a newer ChatGPT model, it turned out to reject a token-limit field the older one accepted. Every ChatGPT call in the run failed. There was no bill to spot it on and no error to catch, and the score rendered happily with one engine silently contributing nothing.
So the field set belongs to the model, not to the engine. Two models from the same vendor took different parameters. Probe a new model before you pin it, because a rejected field is billed at nothing and the probe is therefore free.
What actually moves the number
Five things, roughly in order of how quickly they pay off.
Crawler access comes first, and it's the only binary item on this list. A retrieval bot that gets a 403 can't cite the page it was blocked from. Everything else here is a matter of degree. This one either works or it doesn't, and plenty of sites block the search-side bots by accident while meaning to block the training-side ones. Check your robots.txt against the bots before you spend a day on content.
Decide which layer you are fighting for. If the training layer already knows you, retrieval work is about being the source it quotes. If it doesn't, retrieval is the only layer available to you for the next year or so, and the whole strategy becomes being fetchable and quotable today.
Make the page easy to quote. Engines lift short factual statements. A clear definition in the first two sentences under a heading that matches the question gets pulled far more often than the same fact buried in paragraph nine.
Structured data helps with accuracy rather than volume. It doesn't buy you a citation, but it decides whether the thing quoted about you is correct, which matters when the quote is the only impression a reader gets.
And be where the engine already looks. The German result above is the blunt version of this. If YouTube is the most-cited domain in your category and you have no video, you are absent from the largest single source the engine draws on, however good the article is.
What these numbers can't tell you
The observed-mentions archives are the part most likely to be oversold, so here is what they actually contain. Sistrix appears in 1,409 archived German AI Overview answers. The recorded questions behind those answers are strings like optimierung seo, seo optimierung and http s. Those are search queries, not conversational prompts. Anyone describing that dataset as real user prompts is describing something else.
Three more limits worth stating out loud:
- No engine publishes how it picks sources. Every causal claim in this field, mine included, is inference from outputs.
- A free run is a handful of repeats. That is enough to separate never-mentioned from mentioned-often, and not enough to resolve 40% from 55%.
- Personalization and chat memory make your own screen the worst possible measuring instrument. Checking whether ChatGPT names you by asking it yourself tells you about your account, not your brand.
None of that makes the measurement worthless. It makes the interval around it the honest part of the report.
FAQ
Where to start
Confirm the retrieval bots can reach you. OAI-SearchBot, Claude-SearchBot and PerplexityBot are the search-side agents, and blocking them by accident while aiming at the training crawlers is the most common own goal in this whole field.
Crawler Access Checker →Not your brand name. The questions a buyer asks before they know you exist, phrased the way they would type them. Keep the list fixed, because changing the prompts between runs makes the comparison worthless.
Query Fan-Out →Run each prompt several times per engine and record how often you are named. Read the confidence interval rather than the headline percentage, and ignore any single run entirely.
AI Visibility Checker →The source profile is more actionable than your own score. If the engine keeps quoting one aggregator or one video channel in your category, getting onto that surface beats another post on your own domain.
What the engines cite ↑Absolute numbers move every time a model ships. Your standing against the competitors named beside you is the line worth putting in a report, because it survives the version bumps.
The four metrics ↑See how often the engines name you
Lumina's free AI Visibility Checker runs your prompt across ChatGPT, Claude and Gemini, repeats it, and reports the frequency with a Wilson 95% confidence interval instead of a made-up rank. No signup.
Run the AI Visibility Checker →