Ask two tools how visible your company is in AI answers and you will get two different numbers for the same month. Neither is lying. They are counting different things, and none of them publishes the definition on the same screen as the score.
The short answer
An AI visibility score estimates how often AI engines name your company when buyers ask about your category. Most bundle three separate results into a single figure: whether you were named, whether you were named with a link, and where in the answer you landed. They move independently, which is why two tools rarely give you the same number.
The three things a single score bundles together
Every vendor that sells you a number is compressing at least three results into it. Pulling them apart is most of the work, and it is the part that tells you what to fix.
| Result | What it means | What it cannot tell you |
|---|---|---|
| Mention | An answer named your company in its text | Whether the engine trusts your pages, or whether anyone arrived |
| Citation | An answer named you and linked one of your pages | Whether the buyer read that part of the answer |
| Prominence | Where in the answer you landed, opening line or closing list | Whether the question was one that matters to your margin |
Mention: you were named
A mention is an AI answer that names your company in its text, with or without a link. It tells you the engine associates your name with the question that was asked. It says nothing about whether the engine trusts your pages, and nothing about whether anybody arrived on your site because of it.
Mentions are the easiest of the three to move and the easiest to misread. ChatGPT and Gemini will both name a company they recognise, whether or not the question was one that could ever produce an order. Being named in an answer about your category is not the same as being named in an answer about a part you actually stock. A score that averages the two drifts upward while the answers that matter stay unchanged.
Citation: you were named with a link
A citation is an answer that names you and attaches a link back to one of your pages. It is a stronger result than a mention, because the engine is pointing at your page as the source of something it just said. It is also the only one of the three that can put a visitor on your site.
Citation and ranking have come apart, which is why this is worth measuring on its own terms. A spec page can be quoted by Perplexity or cited inside Google's AI Overviews without ranking for the query at all. It can be quoted by an engine without ever reaching the first page of ordinary search results, and the reverse happens too.
Prominence: where in the answer you landed
Prominence is your position inside the answer itself, whether that answer came from Copilot, Gemini or Google's AI Mode. Named in the opening sentence is a different outcome from named in a closing list of alternatives, and a buyer reading on a phone may never reach the second one. Most scores fold this in silently, weighted by a rule the vendor chose.
This is the one people skip, and it is usually where the interesting news is. A quarter where you moved from the bottom of answers to the top can register as almost no change in a blended score. The count of answers naming you never moved.
Why two tools give you two different numbers
Three choices sit behind every score, and no two products make them the same way. None of these is a flaw. They are design decisions, and they are why the number on one dashboard cannot be set beside the number on another.
- The questions. Every score is produced by running a list of questions and counting what comes back. Change the list and you change the score. A tool weighted toward category questions flatters a brand with good marketing pages. One weighted toward part-level questions flatters a brand with a deep catalogue. The question set is the instrument, and it is rarely published.
- The engines. Some tools sample ChatGPT alone. Others sample ChatGPT, Gemini, Perplexity and Copilot together, and weight them by a formula of their own. Google's own reporting shows how much that choice matters. Analytics groups assistant traffic into a channel that excludes Google's AI Overviews and AI Mode entirely, while Search Console reports on those two and nothing else.
- The counting rule. Even the definition of one appearance varies. In Search Console's generative AI reporting, if two of your pages turn up in the same answer, Google counts that as one. A tool that counted it as two would report a higher number for an identical month, and would not be wrong. It would be answering a different question.
Put those three together and the conclusion is narrow but firm. A score is comparable to itself, measured the same way, over time. Set two vendors' scores side by side and you are comparing two definitions, not two months.
The same month, three different numbers
Here is what that looks like in practice. The numbers below are invented to make the arithmetic visible. They are not measurements of anyone.
Say a distributor runs 40 buyer questions through the engines in one month. Across those 40 answers, the company is named 12 times. Of those 12, five carry a link back to one of its pages. Of those five, three put the company in the opening sentence rather than a closing list.
Twelve, five and three. One month, one company, one set of questions, and three defensible answers to "how visible are we." A vendor leading on mentions reports the first. One leading on citations reports the second. One weighting prominence heavily reports something closer to the third. Every one of them can call its figure an AI visibility score.
The gap between 12 and five is the part worth acting on, and a blended score hides it. It says the engines know the company exists and mostly do not treat its pages as the source. That is a content and catalogue problem with a clear shape, and you only see it when the numbers stay in separate columns.
What to watch instead
You do not need a score. You need an answer to a narrower question, and it is one you can put in your own words.
When a buyer asks an engine who stocks a part you carry, does the answer name you.
That question survives everything above. It does not depend on a vendor's weighting, and it does not change definition between quarters. It maps onto revenue you can recognise, because it is the moment a quote either comes to you or goes to somebody else. Keep it to the parts that carry your margin, ask it the same way every month, and log the three results separately.
The mechanics of running that are the same whether you buy a tool or not, and they are worth understanding before you buy one. We walked through them in how to tell whether AI search is sending you anything, including which free reports cover which half. If you want a definition of the term on its own, AI visibility is in the glossary, and the practice of sampling it over time sits beside it.
What to ask before you buy one
If you are being sold a visibility number, three questions will tell you what you are actually buying. None of them is hostile, and a vendor worth using will answer all three without much prompting.
- Show me the question list. Not a sample, the whole thing. If it is generic marketing questions rather than the parts and cross-references your buyers ask about, the number will move without your revenue moving.
- Which engines, and how often? ChatGPT, Gemini, Perplexity and Copilot do not answer the same question the same way, and a weekly sample and a monthly one produce different trend lines from identical reality.
- How do you count one appearance? Two of my pages in one answer: is that one or two? Named without a link: does that count at all? The answers define the number more than the sampling does.
You are not testing the vendor. You are finding out which of the three results the number leans on, and whether it is measuring the practice or just packaging it. Then you can read it correctly for the next two years.
If it were my call
I would buy a tracker for the logging and ignore the headline number it puts on the dashboard.
The instrument underneath is genuinely useful. Running a fixed list of questions on a schedule and recording what came back is tedious, and software is good at tedious. The score on top of it is a packaging decision, made by someone who needed a single figure for a sales page. Take the three columns and keep your own record. Treat the score as what it is: one vendor's opinion about how to weight three things you are better off reading separately.



