Every AI-visibility report is only as good as the list of questions behind it. I've seen tools show 70% visibility for a brand that nobody's AI assistant ever recommends — because the prompts were so close to the brand's own vocabulary that no competitor stood a chance. And I've seen the opposite: a solid brand scoring near zero because half the prompts asked for definitions, where no provider gets named at all.
This is the practical part the “how to measure AI visibility” articles skip: how to build the prompt set itself. It's the method I use when I set up monitoring for brands, from payment providers to global consumer goods, stripped of client names.
Why the prompt set decides everything
An AI-visibility tool doesn't measure “your visibility in AI”. It measures how often you appear in answers to the specific questions you gave it. Change the questions and the number changes, sometimes dramatically. So the prompt set is not a technical detail you delegate to whoever sets up the tool. It's the definition of what you're competing for.
Pick the prompts the way you'd pick keywords for a business-critical SEO strategy: by what your buyers actually ask and what's worth money, not by what makes the chart look good.
Five rules for a prompt set that tells the truth
- No brand names in the prompt. Not yours, not competitors'. The point is to see whether the model names you on its own. A prompt with your brand in it measures brand perception, which is a separate, useful test — but not visibility.
- Every prompt must be able to name a provider or product. “What is a payment gateway?” doesn't. “Which payment gateway should a small online shop choose?” does. Definitions and how-to questions belong in a separate informational set, if at all.
- Write as the buyer, not as the marketer. Real people describe their situation: “I run three restaurants and need one system for online orders and payments.” That's a far better test than a keyword with “best” in front of it.
- Map each prompt to a page. For every prompt, note which page on your site should answer it. If no page exists, you've just found a content gap.
- Freeze the set for at least a quarter. Changing prompts every month makes trends meaningless. Keep a reserve list and swap in new prompts only at the start of a new period.
Four types of prompts to cover
A good set covers the whole path from “I have a problem” to “which one should I buy”. I split prompts into four intent types, because each one tells you something different.
Prompt types · example category: online payments for small businesses
| Type | Example prompt | What it shows |
|---|---|---|
| Problem / awareness | “Clients pay me late — how can I get paid faster in a small business?” | Whether you're part of the answer before the buyer knows they need your category |
| Need / situation | “I sell online and in a shop. Can I have one provider for both?” | Whether AI connects your offer to a specific situation |
| Recommendation | “Recommend a payment gateway for a small online shop in Poland.” | Whether you're named when someone asks directly |
| Selection / shortlist | “I'm choosing between three payment providers — what should I compare?” | How you're described next to competitors |
Recommendation and selection prompts are closest to revenue, so they should carry most of the weight. Awareness prompts are worth tracking separately: they show whether you're building presence earlier in the journey, which pays off later.
How many prompts, and how to split them
For most companies 50–100 prompts is the right range. Fewer, and one lucky or unlucky answer swings the result. More, and nobody reads the report.
- Split by business priority, not evenly. If one service line brings most of the revenue, it gets most of the prompts.
- Mark a core of about 30. These are the prompts you report on every month. The rest adds depth.
- Keep a reserve of 20–25. Ready to swap in next quarter, so the set can evolve without breaking trends.
- Count language and market variants. Each language version is a separate prompt in most tools. Start with your main market and add others once the method works.
How to run the check and what to score
Once the set is ready, the check itself needs discipline — otherwise you measure noise.
- Baseline first. Run the full set before you change anything on your site. Without a baseline you'll never know what your work changed.
- Same conditions every time. Same models (ChatGPT, Gemini, Perplexity, Google AI Overviews), fresh sessions without personalisation, the same market and language.
- Repeat monthly. AI answers vary run to run, so trends over months matter more than any single result.
What to score depends on the prompt type:
- Recommendation and selection prompts: are you recommended, in which position, which competitors appear next to you, and how are you described.
- Informational prompts: is your site cited as a source, and does the answer reuse your content.
Keep those two scores separate. As I wrote in domain vs brand visibility, being cited as a source and being recommended as a brand are two different contests.
Common mistakes
Prompt set · what to avoid
Don't
- Put your brand or product names in prompts
- Fill the set with “what is…” questions
- Copy your keyword list word for word
- Change prompts every month
- Judge by one run of one model
Do
- Write the way your buyers describe their situation
- Weight prompts by business priority
- Map every prompt to a target page
- Freeze a core set per quarter
- Track trends across several models
A good prompt set takes a few hours to build and saves months of misleading reports. If you'd like a second pair of eyes, I build prompt sets and monitoring as part of GEO and AI visibility work.
Method based on prompt sets I build for brand monitoring in AI; examples are generic and not taken from any client.
FAQ
Frequently asked questions
For most companies 50–100 prompts, with a core of about 30 reported every month and a reserve of 20–25 for future quarters. Fewer prompts make results too random; many more make the report unreadable.
No. Visibility prompts should be generic, so you see whether the model names you on its own. Prompts with your brand name measure how AI describes you — brand perception — which is a separate test.
At least ChatGPT, Gemini, Perplexity and Google AI Overviews, in the market and language you sell in. Answers differ between models, so track each one separately.
Once as a baseline before you change anything, then monthly. AI answers vary between runs, so look at trends over several months rather than single results.
It is visibility measured by actually asking AI models a fixed set of questions and recording whether and how your brand appears, as opposed to estimating it from rankings or traffic.
As a starting point, yes, but not word for word. Turn keywords into the questions and situations buyers actually describe, and drop terms where no provider or product would be named in the answer.