If you want to measure AI visibility over time, start by making the test repeatable. Use the same buyer questions each month. Keep the same language, country, device type and prompt wording. Then track the same outputs from ChatGPT, Gemini, Perplexity, Claude and Grok in one scorecard. That gives you a baseline you can compare.

The mistake we see most often is treating AI visibility like a one-off check. Someone asks five random questions in English, on a laptop in London, then compares that with ten different questions in German on a phone a month later. The numbers look precise, but they do not mean much. A useful process is smaller and stricter. It should be easy enough to run every month, and clear enough that a business owner or agency can explain why the score moved.

Start with a baseline you can repeat - how to set a first measurement using the same buyer questions, language, country and device rules each time

Your first measurement is not there to impress anyone. It is there to create a fixed point.

We set the baseline around four controls:

  1. The same buyer intent
  2. The same language
  3. The same country or city context
  4. The same device rules

If your website is in Polish, measure in Polish. If you sell in Germany, test as a German buyer would ask. If most enquiries come from mobile, note that and keep it consistent. AI systems do not always expose device-specific differences in the same way search results do, but your testing process still needs one rule and one record.

A good baseline sheet includes:

  • Date of test
  • Model used, for example ChatGPT or Perplexity
  • Model version if visible
  • Logged-in or logged-out status
  • Language used in the prompt
  • Country target
  • City target if local intent matters
  • Device type
  • Prompt text
  • Full answer copied exactly
  • Whether your business was named
  • Rank or position of mention if the answer is list-based
  • Source domains cited, where the tool shows them
  • Notes on context

This matters because AI answers vary for real reasons. A query asked in Danish for buyers in Copenhagen can produce a different set of businesses from the same query asked in English for Denmark generally. That is not noise. That is the market.

For local businesses, define the geography properly. In the UK, a plumber in Leeds should not test only “best plumber UK”. They should test “emergency plumber in Leeds”, “boiler repair Leeds”, and similar buyer questions in the form local customers actually use. In Ireland, county and city terms often matter together. In the Baltics and Nordics, inflection, spelling and local phrase order matter more than many English-first tools assume.

Keep the prompt wording stable. Do not “improve” it every month. If you change a prompt, retire the old one and mark the new one as a separate series. Otherwise you cannot tell whether visibility changed or the test changed.

At Seonis, this is why we build measurement around the site’s own language rather than translated approximations. If your market buys in Lithuanian, German or Polish, that is the language that belongs in the baseline. If you want a broader process for multilingual sites, our guide to SEO tools that work when your site is not in English covers the practical issues.

Choose the metrics that show real visibility - which numbers to track, including whether the business is named, where it appears, how often and in what context

The simplest useful metric is named or not named. But on its own, that is too blunt.

We track visibility using six practical measures.

1. Mention rate

This is the percentage of prompts where the business is named at all.

If you test 20 prompts across five models, you have 100 prompt-model results. If your business appears in 28 of them, your mention rate is 28%.

This is the easiest trend line to explain month by month.

2. Position in the answer

If the model gives a ranked list, note where you appear.

  • First named
  • In the top three
  • Lower in the list
  • Mentioned only after follow-up
  • Not named

This matters because being fourth in a list of ten is not the same as being the first recommendation in a short answer.

3. Share of voice against named competitors

Record which other businesses are named in the same answer.

You do not need a complicated weighting model at the start. A simple monthly count works:

  • Our brand mentions
  • Competitor A mentions
  • Competitor B mentions
  • Generic non-brand answers

This helps you spot whether a drop is your problem or a market-wide shift in how the model answers.

4. Context of mention

Being named is not enough. You need to know why.

Classify the mention as:

  • Recommended provider
  • Example brand
  • Included in a comparison
  • Listed by location
  • Mentioned with caution or negative framing
  • Cited only as a source, not a recommendation

A brand that appears often but only as “one option among many” may need stronger category pages, clearer proof, or better citations.

5. Citation visibility

Where the tool shows sources, track whether your website is cited directly, whether third-party sources cite you, or whether the answer is built from other domains.

This is especially useful in Perplexity, and sometimes in other tools depending on the interface. If your site is rarely cited but review sites and directories are, that tells you where the model is finding confidence.

6. Intent coverage

Split prompts by buying stage:

  • Discovery
  • Comparison
  • Local intent
  • Problem-led
  • Brand validation

Then score mention rate within each group.

Many businesses are surprised by this. They may be visible for branded or near-branded checks, but absent for early discovery prompts such as “best accounting software for small manufacturers in Germany” in German. That is where growth usually sits.

Build a prompt set from real buying journeys - how to create a small list of practical prompts for discovery, comparison and local intent without guessing

A prompt set should come from your sales process, not from imagination.

We build small sets first. Usually 12 to 20 prompts are enough for a stable monthly scorecard. More than that is fine if you have the time, but most small teams will not maintain it.

Start with three sources.

Sales and enquiry data

Look at:

  • Emails from prospects
  • Contact form submissions
  • Sales call notes
  • Demo requests
  • Live chat transcripts
  • CRM reason codes if you use them

Pull out the actual wording buyers use when they describe the problem, compare options, or ask for local help.

Search data

Use:

  • Search Console queries
  • Paid search terms if you run ads
  • Internal site search
  • Search term reports from marketplaces or directories, if relevant

Search Console is especially useful because it shows the language and phrasing that already bring impressions. It will not tell you AI visibility directly, but it is a strong source for realistic prompts.

Customer-facing staff

Ask the people who hear objections every week:

  • What do buyers compare us with?
  • What do they ask before they are ready to buy?
  • What local wording do they use?
  • What category words do they avoid?

This is often where the best prompts come from.

Then group the prompts.

Discovery prompts

These are category-level. The buyer knows the problem, not the brand.

Examples:

  • Best payroll software for small firms in Poland
  • How to choose a tax adviser in Cork
  • Alternatives to spreadsheets for warehouse stock control in German

Comparison prompts

These are shortlist questions.

Examples:

  • Which is better for a small retailer, X or Y
  • Best alternatives to [competitor]
  • Compare local bookkeeping services for restaurants in Tallinn

Local intent prompts

These are for buyers who want a place, provider or specialist in a location.

Examples:

  • Immigration lawyer in Malmö for employers
  • Best physiotherapy clinic in Riga centre
  • Managed IT support for SMEs in Birmingham

Keep the wording close to natural speech. Do not stuff prompts with every possible modifier. AI systems respond to intent and phrasing, not just token matching.

For businesses operating in more than one country, keep separate prompt sets by market. A German-language set for Germany is not the same thing as a German-language set for Austria. The same applies across the EU more broadly. Consumer language, legal terms and local buying habits differ, even where the language overlaps.

If your content process is weak in local languages, fix that before expecting stable gains. Our guide on publishing local-language SEO articles without translation explains why translated filler rarely holds up in either search or AI answers.

Score results in a way a small team can maintain - how to turn answers from ChatGPT, Gemini, Perplexity, Claude and Grok into a simple monthly scorecard

A scorecard should be boring enough to survive.

We recommend a spreadsheet or dashboard with one row per prompt-model pair. That means each prompt is tested in ChatGPT, Gemini, Perplexity, Claude and Grok, then scored using the same fields.

A simple scoring system works well:

Core score per result

  • 3 points - business named first or strongly recommended
  • 2 points - business named in top three or clearly included
  • 1 point - business named, but weakly or low in the answer
  • 0 points - not named

Then add two optional flags:

  • +1 if your own website is cited directly
  • +1 if the answer includes strong commercial context, for example “good for small manufacturers” when that is your target segment

Do not let the score become too clever. If nobody can score the answer in under a minute, the system will not last.

Monthly roll-up

For each month, calculate:

  • Total prompts tested
  • Total prompt-model results
  • Mention rate
  • Average score
  • Mention rate by model
  • Mention rate by intent type
  • Direct citation count
  • Number of unique competitors named

This gives you both a headline and enough detail to diagnose movement.

Keep raw answers

Always save the original answer text or a screenshot. Models change. Interfaces change. A summary note like “mentioned second” is not enough when you need to explain a trend three months later.

Use one scoring guide

Write down your rules in plain English.

For example:

  • “Named first” means the first brand in a list or the first explicit recommendation in prose.
  • “Strongly recommended” means the answer gives a clear reason tied to buyer fit.
  • “Weak mention” means the brand appears in a long list without explanation.

This sounds obvious until two people score the same answer differently.

At Seonis, we automate the repetitive part of this work and report it in the owner’s language, because manual tracking across five models gets old quickly. If you want to see the method in a more focused form, our article on checking whether AI tools mention your business walks through the basics.

A rise or drop is only useful if you can tie it to a cause.

When the score moves, check five areas.

1. Content changes

Did you publish new category pages, service pages or articles that match the prompt set?

Look for:

  • Better coverage of buyer questions
  • Clearer product or service descriptions
  • More local detail
  • Better comparison content
  • Stronger proof, such as case studies or credentials

If the content is in the wrong language, too generic, or obviously translated, do not expect durable gains.

2. Links and citations

AI systems often rely on the same web signals that shape search visibility, plus citation patterns from trusted pages.

Check:

  • New backlinks from relevant sites
  • New directory or association listings
  • Press mentions
  • Reviews on platforms the models seem to cite
  • Consistency of business name, address and phone details for local firms

For local businesses in the UK and Ireland, citation consistency across major directories still matters. Across the EU, the exact mix of directories and trade portals varies by country, so the useful check is not “have we listed everywhere” but “are we present on the sources buyers and models actually use in this market”.

If authority is thin, it often shows up first in AI comparisons. That is one reason we built a members’ system around credit-based link exchange in practice, so smaller businesses can improve authority without pretending links do not matter.

3. Technical issues

If your pages cannot be crawled, rendered or understood properly, visibility can slip.

Check for:

  • Robots rules blocking key pages
  • Noindex on commercial pages
  • Broken canonicals
  • Slow or failing server responses
  • Structured data errors where relevant
  • JavaScript rendering problems on platforms such as WordPress, Shopify, Webflow or Ghost

Do not overstate technical SEO for AI answers, but do not ignore it either. If the models or their source pipelines cannot access your pages cleanly, your content will not do its job.

4. Search demand shifts

Sometimes visibility changes because buyer language changed with the season, regulation or market conditions.

This is common in tax, legal, travel, energy and B2B software. Refresh the prompt set only when the market clearly changed, and keep old and new versions marked separately.

5. Model behaviour

Sometimes the model changed.

One month, a tool may favour concise list answers. Next month, it may give fewer brand names and more generic advice. That is why you track each model separately. If your mention rate falls in Claude only, while ChatGPT, Gemini, Perplexity and Grok stay stable, the likely cause is not your site alone.

AI visibility is not the end result. It is a leading indicator.

The monthly report should compare visibility with three business signals.

Enquiries and leads

Track:

  • Contact form submissions
  • Calls
  • Demo requests
  • Quote requests
  • Qualified leads if you have a clear definition

Look for lag, not just same-month correlation. A rise in AI mentions may show up in enquiries later, especially in B2B or higher-consideration services.

Branded demand

Watch whether more people search your brand name and branded product terms. You can see this in Search Console and in paid search data if you run brand campaigns.

If AI tools start naming you more often, branded search often follows before direct conversions do. People hear the name in an answer, then search for the business separately.

Search Console trends

Compare AI visibility with:

  • Impressions on key non-brand queries
  • Clicks to commercial pages
  • Average position on core topics
  • Growth in long-tail queries in the local language

Do not expect a perfect match. Search and AI answer systems are not the same product. But if both move up on the same topic cluster, that is usually a real improvement in market visibility rather than a random fluctuation.

A simple monthly report can be one page:

  • Headline AI visibility score
  • Mention rate by model
  • Best and worst prompt groups
  • Main changes since last month
  • Likely causes
  • Impact on enquiries, branded demand and Search Console
  • Next actions for content, authority and technical fixes

That is enough for an owner to review, and enough for an agency to act on.

The main point is this. Measuring AI visibility over time is not about chasing a perfect universal score. It is about running the same buyer-centred test every month, in the right language and market, then comparing the result with what changed on the site and in the business. If the method is stable, the trend becomes useful. If the method keeps changing, the score is just theatre.

That is the standard we use at Seonis: SEO and AI visibility in your website’s language. We track whether ChatGPT, Gemini, Perplexity, Claude and Grok name the business when buyers ask, and we tie that back to the content, authority and technical work that can actually move the number. For small teams, that is what makes the report worth reading.