A baseline is a dated record of what the AI engines say when someone asks about what you sell. You spend an afternoon asking the questions a buyer would ask, writing down every answer, and counting how often you get named. By dinner you've got a number, and something to measure next quarter against.

One afternoon, a spreadsheet, and the questions your buyers actually ask.

What an AI-search baseline actually is

A baseline is one dated page of notes: the questions you asked, the answers you got back, and how many of them named your company. That's the whole idea. It isn't a score a tool hands you, and there's no industry number to compare it against.

Doing it by hand the first time is worth the afternoon. You find out which buyer questions hand the answer to somebody else, and that turns out to be more useful than the count. The ongoing practice this becomes has a name, AI visibility tracking, and it's the same work on a repeat schedule.

One thing to settle before you start. Trying to work out whether AI answers send you any traffic at all is a different question, with a different method. We wrote that one up separately: how to tell whether AI search is sending you anything. This piece assumes you've decided to go and look directly.

Six-step timeline of an AI-search baseline afternoon: build the prompt list, fix the rules, run the list, mark up the answers, work out two rates, date and file the snapshot.

The afternoon, end to end. Durations are our estimates, not measurements.

Step 1: write down the questions your buyers actually ask

About forty-five minutes. Open a spreadsheet and write one buyer question per row. Not keywords. Whole questions, phrased the way somebody outside your company would type them.

The good ones are already in your building. Pull from last month's quote requests, the notes your counter staff keep, and whatever people type into the search box on your own site. You're looking for the shapes that repeat:

  • A part number on its own.
  • A cross-reference from a competitor's number.
  • "Who stocks" plus a brand you carry.
  • A spec question that ends in a purchase.
  • The broad category question nobody in your market owns yet.

Aim for twenty to thirty. That's our default, and here's the reasoning rather than a rule. Below about twenty, a single strange answer moves the result enough to fool you. Above thirty, you won't finish this afternoon, and an unfinished baseline is worth nothing.

A fixed list like this is what the trade calls a set of benchmark prompts, and the word "fixed" is doing the work. You'll reuse this exact list next quarter. Write your own, too. A generic list of industrial questions would cost you the part that matters, which is that these are your buyers asking about the brands you actually stock.

Step 2: pick your engines and fix the rules

Two engines are enough for a first baseline, and the rules you set now matter more than which two you pick. Start with ChatGPT. Pew Research Center found that 44% of US adults reported using it in February 2026, so it's the one your buyers have most likely touched. Add Google's AI Overviews, the answer block that sits above the blue links. It turns up in front of people who never went looking for an AI at all.

Then write your rules at the top of the sheet, and don't change them mid-afternoon:

  • Signed in or signed out. Pick one.
  • Browsing or web access on or off. Pick one and write down which.
  • One run per question. No re-rolling an answer you didn't like.
  • Same day, same sitting.

None of those choices is the right one in any absolute sense. Writing them down is what makes the afternoon repeatable, and repeatable is the only property that turns this into a measurement instead of a story.

Step 3: run the list and paste every answer in

This is the mechanical hour. Paste question one, copy the entire answer, paste it into the row. Repeat. Don't summarise as you go and don't clean anything up, because the wording you delete is usually the wording you'd want in January.

Two habits make the hour cleaner. Paste the whole answer, including the part where the engine names other companies, since that's half of what you're measuring. And run the awkward questions, the ones you suspect will go badly. Skipping those is the quickest way to a baseline that flatters you and tells you nothing.

If you want the procedure for one engine in more detail, with the account and browsing settings spelled out, we covered tracking brand mentions in ChatGPT on its own.

Step 4: mark up what came back

Now you turn a pile of answers into a record you can compare against. Go back through each row and fill in seven fields. Seven is a proposal, not a standard, and each one earns its place by being something you'd want to know a quarter from now.

The seven-field log: date, engine, prompt verbatim, named, linked, position in the answer, and competing sources named, shown with one filled example row.

The seven fields, with one row filled in. Copy the header row into your own sheet.

The field that carries the most weight is the split between named and linked. An engine can describe your company accurately and send the reader somewhere else entirely. Those are two different outcomes, and collapsing them into one tick hides the more useful of the two.

The longer treatment of what to record, and how often, is in what to log, and how often. The term itself is defined in the glossary entry on AI citation tracking. Log the competing sources by name, as well. Finding out that the same three suppliers get named for your part numbers is the most actionable thing this afternoon produces.

Step 5: turn the log into two rates

Two numbers come out of the sheet. Your mention rate is how many answers named you, out of the questions you asked. Your citation rate is how many named you and linked to you. Both are counts, not opinions, which is what makes them worth writing down.

Say you ran twenty-five questions and you were named in four answers, two of which linked to your site. That's your baseline. It's a small number and it's supposed to be. A low mention rate on a first pass is the ordinary result for a company that has never done this.

The arithmetic is trivial, and easy to get subtly wrong at the end of a long afternoon. We built a free AI visibility calculator that does it in the browser and saves nothing. Paste your counts in and it gives you both rates. If you'd rather understand the distinction first, the difference between mention rate and citation rate is written up on its own.

Step 6: date it and file it

Put today's date on the sheet, write your two rates at the top, and save it where you'll actually find it in three months. That last part sounds like filing advice. It's the step people skip, and skipping it costs you the entire exercise.

Write one line under the rates about what you noticed. Which questions went badly, and who kept getting named instead of you. That note is what you'll read first next quarter, before you look at either number. Everything else about tracking this over time is a repeat of the afternoon you just spent.

What this baseline can't tell you

A one-afternoon baseline is a snapshot, and snapshots have real limits. Being straight about them is what keeps the number useful instead of decorative.

It has sampling error. Ask the same engine the same question on the same afternoon, and you can get a different answer. These systems don't return one fixed result the way a search ranking does. Your twenty-five questions are a sample of every question a buyer might ask, and a small one.

It doesn't tell you why. You'll see that a competitor gets named and you don't, and the sheet is silent on the reason. That's a separate investigation, and we've written about how AI search picks its sources if you want to start it.

And a single baseline, on its own, is worth almost nothing. Its whole value is that you run the same list again on the same rules next quarter, and get a second number to compare. One number is a fact about one afternoon. Two numbers are a direction.

If your site is young or lightly linked to, set your expectations for the shape of the change rather than its size. The first movement usually shows up on the narrow questions, not the broad ones.

A specific cross-reference or an unusual part number is a question with few good answers available, so a thin, accurate page can win it. The broad question, the one that names an entire category, tends to go to whoever the web already links to most. You'll likely see the specific answers turn over first, while the broad ones sit still for a long time.

How long is genuinely not something this afternoon can measure, and we're not going to put a date on it. What you can check cheaply is whether the engines can read you at all, which is a different problem from whether they like you. An llms.txt file, a plain-text map of your site written for AI crawlers, is one of the things worth looking at there.

If it were my call

Twenty-five questions. ChatGPT and AI Overviews, nothing else on the first pass. Signed out, browsing on, one run each. Repeat it the first Monday of the quarter, and put the next one in the calendar before you close the sheet.

We'd skip the other engines until that's actually running for two quarters. Adding Perplexity and Gemini to a baseline nobody has repeated yet gives you four times the work and no more direction than you had. Get one honest comparison first. Then widen it, using the same fixed question list you already wrote.

The one-screen checklist

Everything above, in the order you'd actually do it. If you want the vocabulary behind step four, it's in the citation tracking entry.

  1. Write twenty to thirty real buyer questions in a spreadsheet, one per row.
  2. Pick two engines. Write your rules at the top: signed in or out, browsing on or off, one run each.
  3. Run every question. Paste the full answer into the row. Don't skip the awkward ones.
  4. Mark up each row across the seven fields, keeping named and linked apart.
  5. Count your mention rate and your citation rate.
  6. Date the sheet, write one line about what you noticed, and diary the next run.

The header row, to copy into your own sheet: date · engine · prompt verbatim · named · linked · position in the answer · competing sources named