More and more clients ask me the same thing: “does my brand show up in AI answers — and how do I even check?” The good news: you can measure it. The less good news: none of the available methods is perfect, and you have to treat the numbers with care, without blind faith in any single “score.” Let me show you how I actually measure a brand's presence in AI — on two levels: quantitative (whether and how often you appear, and how much real traffic you get from it) and qualitative (how the assistant describes you, and whether the facts about you check out).
Why measuring visibility in AI plays by different rules
In classic SEO we had a comfortable point of reference for years: a position in Google, a click count, impressions for a phrase. In AI answers that comfort disappears, and you have to make peace with it before you measure anything.
First, the click stops being necessary. The assistant often answers in full and the user never visits a site. The traffic that would have reached you in the classic model stays on the assistant's side. That doesn't mean you get nothing from it — I'll come back to that under the limits.
Second, there is no single “position”. AI answers are heavily personalized: they depend on the user's history, the context of the conversation, the country, the language, even the way the question is phrased. That's nothing new either — classic Google results have been personalized for a long time. So we're not hunting for one “true” position; we're after a general picture and a comparison: how I do against competitors on typical, repeatable queries.
Third, the same question asked twice can give different answers. Models are non-deterministic, their sources shift, and providers update them without warning. Measurement is therefore more like a weather forecast than a meter reading — it shows a tendency, not a number carved in stone.
Two layers of measurement: quantity and quality
I split the whole thing into two layers, because in conversations about “presence in AI” they keep getting blurred.
The quantitative layer answers: do I appear, how often, for which queries, and how much real traffic I get from it. It's numbers — impressions, traffic, share against competitors.
The qualitative layer answers a different question: what exactly the assistant says about me. Whether the description is positive, whether the facts are right, whether it pins someone else's flaws or outdated information on me.
The two layers matter equally and neither replaces the other. You can appear in every other answer (great quantitative layer) and be described there as “a cheap option for beginners” (weak qualitative layer). A citation counter won't show that — which is why I measure both at once.
The quantitative stack — what I actually measure with
No single tool shows everything. So I combine several sources, and only from putting them side by side do I build a picture.
- Google Search Console — the Generative AI tab. It shows impressions and clicks from Google's generative answers. That's data straight from the source, so I treat it as the starting point on the Google side.
- Bing Webmaster Tools. It has a similar tab and, in practice, tends to give a bit more information than GSC. Worth it, because Copilot and some assistants lean on the Bing index.
- GA4 — real traffic from assistants. Traffic coming in from ChatGPT, Perplexity, Copilot or Gemini (with the referrer preserved) can be filtered out. It's the closest to a “hard” number of actual visits — but with an important caveat: it only catches the clicks that carry a referrer. You won't see a large part of AI's influence here, because it reaches you as branded or direct traffic, driven by an earlier recommendation (more on that under the limits).
- Prompt and competitor-comparison tools. Here I use chatbeat.com (by Brand24), but the category is crowded and the tools are similar — they differ in nuances and individual features. Alternatives worth knowing: Share of Model, Ahrefs Brand Radar, Profound, Peec AI, Otterly, Semrush (the AI module). They all do essentially the same thing: they ask the models a set of phrases from your industry and show whether and how often your brand appears against competitors.
In short: I count real traffic (GA4), I count impressions in Google and Bing, I study prompts in a chatbeat-class tool, and I compare against competitors. Only together do these sources give a sensible picture — on its own, each one overstates or understates part of the story.
The qualitative stack — what the numbers won't show
The quantitative layer tells you whether and how often you appear. It won't tell you how you're described — and that often decides the outcome. So on top of the numbers I examine the content of the answers themselves, and I look at two things.
The first is image: whether the assistant talks about the brand positively, neutrally or negatively, and in what context it recommends you (or advises against you). The second is factual accuracy: whether the information about the brand — offer, prices, locations, differentiators — is current and true, whether the model isn't confusing you with a similarly named competitor, and isn't pinning someone else's flaws or outdated stories on you.
That's a topic of its own, and I develop it in the piece on brand perception in AI answers. Here it's enough to remember one thing: measurement without the qualitative layer is blind to what your customer actually reads.
Where AI gets its information about you — the measurement that points the way
The most practical part of measurement isn't “whether” and “how often,” but where the model draws what it says about you and your competitors from. AI answers don't come from nowhere — the assistant builds them from specific sources: your site, opinions and reviews, Reddit threads, articles, directories, industry rankings and comparison sites.
So when I measure, I check not just the result but its foundation: what exactly the model assembles your brand's description from, and what it assembles your competitors' descriptions from. That comparison can say more than the “position” alone. If the model cites a rival from an industry ranking, a “best of” roundup or a discussion where you simply aren't present — you have a ready-made direction: those are the places worth showing up in to improve your own result.
At that point measurement stops being a report for the drawer and becomes a list of places to win. And that's the natural bridge to visibility work — because now you know not only how you're doing, but where to add signals so you do better.
Need this kind of measurement?
Measuring a brand's visibility in AI — both quantitatively and qualitatively — is done by Tomek Sikora, an independent SEO and AI Search (GEO) specialist with 12+ years of experience across 30+ markets. I combine data from GSC, Bing Webmaster Tools and GA4 with prompt tools, examine the sources and the competitive comparison, and at the end you get not a chart but a list of concrete actions. Month-to-month work with no lock-in, a sprint model for smaller projects, and a free 20-minute consultation to start.
See the service: GEO / AI Visibility →What's actually worth counting — and what's a vanity metric
The easiest things to count are the ones that matter least. A raw citation number with no context looks good on a chart, but it won't tell you whether it's working for you or against you. Same with a “position” treated as one objective truth — at this scale of personalization, that's an illusion.
What I actually watch: real traffic (GA4), the impression trend in Google and Bing (not a single reading, but the direction over time), presence in the recommendation for the queries people actually type before a decision — buying and comparison queries, not generic phrases — plus the comparison against competitors and the image and factual accuracy. One good trend-and-comparison metric is worth more than ten “green” counters sitting side by side.
If you'd rather someone put this together for you and turn it into a plan — that's exactly what I do.
The limits you need to know about
I'll say it plainly, because I'd rather you heard it from me than from disappointment: none of these methods is perfect.
Tools like chatbeat, Ahrefs Brand Radar or Share of Model don't account for the full personalization of answers — and in assistants that personalization is very high. That's not a flaw in any one tool; it's the nature of the medium, exactly like classic Google, which has personalized results for years. So you don't get a precise ranking, only a general comparison and a picture of the situation. And that's fine — as long as that's how you treat the numbers.
The second thing, the most important: AI “takes” the click, but the recommendation works anyway. The customer may not visit your site from the link in the answer — but if the assistant recommended you, they'll reach you later: by typing the brand into Google, or by coming in directly. You won't see that in referral traffic, and that's where most of it happens. So you must never judge presence in AI by clicks alone — you have to read it together with branded and direct traffic.
Measuring visibility in AI is more like a weather forecast than a meter reading. Watch the tendency and the comparison against competitors, not one number carved in stone.
And that's why you measure it continuously — the way you monitor rankings in SEO. Models, their sources and the way they describe brands all change over time, so a one-off audit is just a snapshot of a single moment; the value is in repetition and in watching the direction. The one case where measurement is premature: if the brand doesn't appear in AI answers at all, there's nothing to assess about the description — build visibility first, then measure how you're described.
FAQ
Frequently asked questions
Yes, but as a general picture and a trend — not as a single “position.” By combining several sources — GSC, Bing Webmaster Tools, GA4 and a prompt tool — you get a reliable comparison against competitors and a direction of change over time. There is no single precise ranking in this medium.
There is no single one that shows everything. I combine GSC (the Generative AI tab), Bing Webmaster Tools, GA4 and a tool like chatbeat.com. Alternatives — Share of Model, Ahrefs Brand Radar, Profound, Peec AI, Otterly, Semrush — do essentially the same thing and differ in nuances; the choice is more about budget and features than one “best” tool.
Not to start. GSC, Bing Webmaster Tools and GA4 are free and give you surprisingly much. Paid tools add prompt research, competitor comparison and continuous monitoring with alerts — and that's the point where they start to earn their keep.
Because the recommendation drives traffic you can't see directly. The customer reads about you in the assistant's answer, then reaches you later via a branded search or a direct visit. There was no AI click, but the decision was made under its influence.
Continuously, the way you track rankings in SEO. Models and their sources change without warning, so what matters is the trend and reacting to shifts, not a one-off audit once a year.
Both at once. You can be highly visible and, at the same time, poorly or inaccurately described. I dig into the qualitative layer in the piece on brand perception in AI answers.
The source context tells you the most. Once you check what the model builds its description of you and your competitors from, you see where a competitor is present and you're missing — a ranking, a roundup, reviews or a discussion. Those are ready-made places to win. That's how measurement stops at a chart and turns into a to-do list.