Open ChatGPT, ask it who sells a part you stock, and see whether your name comes up. That is the whole test, and it takes about ten minutes. Run it once and you have an anecdote, so the rest of this is how to turn those ten minutes into a number you can watch.

What you're actually measuring

Two different things happen when ChatGPT answers a buyer, and it helps to keep them apart from the start.

The first is getting named. A maintenance engineer asks where to find a seal kit, and the answer says your company is one place to look. Your name is in the text. The second is getting named with a link, where the answer attaches your page as the source it drew from. That second one is a citation, and it is the more valuable of the two, because the buyer can act on it without typing anything else.

You can have one without the other, in both directions. Plenty of answers name a distributor and link somewhere else entirely. Some link a page without ever naming the company that owns it. Tracking them as one number hides the gap, which is why mention rate and citation rate are worth keeping in separate columns.

Neither one is a ranking. There is no position 1 to hold in a ChatGPT answer, so what you are watching is frequency, not placement. Ranking and getting cited have genuinely come apart. Ahrefs studied 15,000 long-tail queries across four assistants and found that only 12% of cited URLs ranked in Google's top 10 for the same prompt.

That cuts both ways, and mostly in your favor. It means a spec page can get quoted without ever reaching Google's first page. Whether any of it is reaching you in the first place is a separate question. We went through the two free reports, and what each one misses, in how to tell whether AI search is sending you anything.

Why one run tells you nothing

Here is the part most people skip, and skipping it is how you end up confidently wrong.

Ask ChatGPT the same question twice and you can get two different answers. Ask it from a different account and the gap widens. Turn web browsing on and it may reach for pages it ignored a minute earlier with browsing off. Model versions change under you. None of that is a malfunction; it is how these systems work, and it means a single check is close to worthless as evidence.

So when you run one query, see a competitor's name, and conclude that the competitor owns that question, you have not learned that. You have learned what one answer looked like once.

Three things convert the anecdote into a measurement, and none of them requires software. This is the whole of what AI visibility tracking amounts to:

  • The same questions every time. Change the wording and you have started a new baseline without meaning to.
  • The same schedule every time. A monthly check on roughly the same day is enough. Sporadic checks tell you nothing about direction.
  • The answer stored word for word. Paste the whole thing. Your summary of it will drift toward what you expected to see.

You will hear the argument that manual checking is not real tracking, usually from companies selling tracking. There is something to it at scale. But the argument proves less than it looks like. A fixed set of questions, run on a fixed schedule and logged word for word, is a measurement by any fair definition. It is a small one. Small and real beats large and imaginary.

The starter set: ten questions to run

Almost every guide on this subject tells you to build a set of questions and then leaves you to invent them. So here are ten, written the way industrial buyers actually type. Swap the bracketed parts for your own parts, brands and region, and keep the wording fixed after that.

  1. What is a [PART NUMBER] and who sells it?
  2. I need a replacement for [DISCONTINUED PART NUMBER]. What are my options?
  3. What is the [BRAND B] equivalent of [BRAND A PART NUMBER]?
  4. Is [PART NUMBER A] interchangeable with [PART NUMBER B]?
  5. I need a [COMPONENT TYPE] rated for [PRESSURE OR TEMPERATURE] in [MATERIAL]. What should I look at?
  6. What size [COMPONENT] do I need for [APPLICATION]?
  7. Who stocks [PART NUMBER] in [REGION]?
  8. Who are the best distributors for [BRAND] [PRODUCT CATEGORY] in [REGION]?
  9. Which suppliers carry [PRODUCT CATEGORY] with same-day shipping?
  10. What should I look for when buying [PRODUCT CATEGORY] for [INDUSTRY APPLICATION]?
Four question shapes to cover: a part number on its own, a cross-reference between two brands, a specification with no part number, and a sourcing question about who stocks it

Four shapes to cover. Miss one and the set has a blind spot.

Those ten are not arbitrary. They cover the four shapes an industrial question tends to take. A part number on its own. A cross-reference between two brands. A specification with no part number attached. And a sourcing question about who actually has the thing. Questions one through four are where your catalog either does its job or does not. Seven through nine are where you find out whether the engines think of you as a supplier at all.

Ten is a starting point, not a ceiling. A full set built out across a catalog runs to several hundred questions, sorted by product line and buying stage. We will publish ours under an open license, so you can take it rather than rebuild it. For now, ten that match your actual line card will tell you more than three hundred generic ones. These are what we mean by benchmark prompts: a fixed list you rerun rather than a list you improvise.

One note on choosing them. Pick questions you would be genuinely pleased to be named in, not questions you already know you win. The point is to find the gap, and the gaps are rarely where you would guess. This is also the difference between guessing at demand and reading it, which is the idea behind prompt-shaped demand.

Running it the same way every time

The measurement is only as stable as the conditions around it. Four of those conditions matter enough to write down and hold.

A four-step loop: fix the question set, run on a fixed schedule, log each answer word for word with the date, then count named against named-with-a-link, and repeat

The loop runs monthly. Step one only happens once.

Start a new chat for every question. An existing conversation carries context forward, and that context changes the answer. One question, one fresh chat.

Hold the browsing setting steady. Whether the model searches the web while answering changes which pages it can reach, and which of your pages an AI crawler got to in the first place. Either always on or always off is fine. Switching between runs is not, because you will not know which change caused which result.

Keep the account consistent. Answers can differ between accounts. Use the same one, or use a signed-out session every time, and note which you chose.

Record the date and the model. When something moves, the first question is always whether the model changed underneath you. You cannot answer that later if you did not write it down.

None of this is difficult. It is just easy to skip, and skipping it is what turns a baseline into a pile of screenshots.

What to write down

A spreadsheet is enough. Seven columns will carry you a long way. The date, the question as you asked it, whether you were named, and whether you were named with a link. Then roughly where in the answer you appeared, and which other companies showed up.

That last column is the one people leave out and then wish they had. Knowing you were absent is useful. Knowing who was there instead tells you where the answer is coming from. It is also how you notice a competitor starting to get named for your own part numbers.

We go through the record itself in more depth separately, including why each column earns its place and what changes the right schedule. The short version is that the log is the asset, not any single run of it. The glossary entry on AI citation tracking covers the definition if you want it in one paragraph.

Want to see what the numbers imply before you start logging? Our AI visibility calculator turns a mention rate and a link rate into what they mean across a whole question set. Useful for sanity-checking a first month against something.

When a spreadsheet stops being enough

Ten questions once a month is about twenty minutes of work. Two hundred questions across four engines, several times a month, is a job, and at that point a spreadsheet becomes the bottleneck rather than the tool.

Software exists for this. Several platforms run question sets on a schedule and report how often you were named and linked. Ahrefs and Semrush both ship something here, alongside a crowd of smaller specialists. We are not going to tell you which to buy. The category moves monthly, and the right answer depends on how many questions and engines you need.

What is worth saying is that the order matters. Run the manual version first, even briefly. It teaches you which questions are worth paying to monitor, and you will pick a tool far better after a month of doing it by hand than before. Buying first tends to produce a dashboard nobody reads.

ChatGPT is the sensible place to start, because it is the one with the broadest reach. Pew Research found that 34% of US adults had used ChatGPT by early 2025, about double the 2023 share. Your buyers sit inside that population, even if nobody has surveyed them specifically.

Perplexity, Gemini and Copilot behave differently enough to deserve their own treatment, and we will get to that comparison rather than hand-waving at it here. If you want the wider picture of how these answers get assembled, we covered that ground separately.