If you want to know whether AI tools mention your business, start with a simple rule. Do not ask one vague question in ChatGPT, Gemini or Perplexity and treat the answer as evidence. Use a fixed set of buyer-style prompts, run them on a schedule, save the outputs, and score whether your brand appears in a useful context.
For most small businesses and agencies, the right way to track brand mentions in ChatGPT, Gemini and Perplexity is a choice between three levels. Manual checks are fine when you only need a quick read. A spreadsheet process works when you need repeatable evidence without buying software. A dedicated platform makes sense when you need regular monitoring across several AI tools, in several languages, with proof of change over time.
The practical answer: what most teams should compare first
Before you compare tools, decide what you are trying to prove.
Usually it is one of these:
- Are we mentioned at all?
- Are we mentioned for the commercial questions buyers actually ask?
- Is our visibility improving month by month?
- Can we show this clearly to a client, manager or owner?
Those are different jobs. A one-off manual check can answer the first. It cannot answer the third very well. A proper monthly report can answer the third, but only if the prompts, timing and scoring stay consistent.
For a small business, the practical split is usually this:
- Manual checks, if you want a quick baseline and have a short list of high-value prompts
- Spreadsheet tracking, if you need a cheap and repeatable process
- Specialist software, if you need broader coverage, team workflow, multilingual reporting and less manual work
For an agency, the threshold is lower for using software. Once you are checking multiple clients across several markets, the admin cost of screenshots, prompt libraries and scoring becomes real. Consistency also matters more. If two account managers test the same client in slightly different ways, the report stops being comparable.
Budget matters, but time matters more than many teams admit. If someone spends hours each month checking ChatGPT, Gemini, Perplexity, Claude and Grok manually, that is already a cost. The question is whether that time produces evidence you can trust.
If you want a simple starting point, read our guide on how to track AI mentions of your company. It is the same principle we use in Seonis, just without the nightly automation.
What counts as a brand mention in ChatGPT, Gemini and Perplexity
Not every mention is useful.
A proper tracking setup needs a definition of what counts. Otherwise one person marks a result as a win because the brand name appears once, while another marks it as a miss because the answer would never send a buyer your way.
We usually split mentions into three types.
First, a named mention. Your business name appears in the answer. That is the minimum threshold.
Second, a recommendation mention. Your business is not just named, it is suggested as an option, included in a shortlist, or described as suitable for a need.
Third, a buyer-relevant mention. The prompt reflects a real commercial search or question, and the answer places your business in a context that could influence a buying decision.
That last part matters most. If an AI tool mentions your company when asked “list software companies in Europe”, that may not help much. If it mentions you when asked “best accounting software for small manufacturers in Germany” or “who can help with payroll compliance in Poland”, that is far more valuable.
So define your scoring before you begin. A simple model works well:
- 0 = not mentioned
- 1 = mentioned by name only
- 2 = recommended or included in a relevant list
- 3 = strongly recommended in a buyer-relevant answer
You can add notes such as:
- Position in list
- Whether competitors were named
- Whether the answer linked to your site
- Whether the explanation was accurate
- Whether the answer was in the local language
Accuracy deserves its own check. A mention with the wrong service, wrong geography or wrong product category is not a clean win. In some sectors it is worse than no mention.
Also decide whether you care about citation behaviour. Perplexity often shows sources more explicitly. ChatGPT, Gemini and others may vary by mode, account state or interface. If referral traffic matters, track not just whether you are named, but whether the answer gives the user a path to your site.
Option 1 - manual checks in the AI tools themselves
Manual testing is enough when the business is small, the prompt set is short, and the main goal is a baseline.
That might mean:
- 10 to 20 prompts
- one country or language
- one product line
- one monthly check
- one person doing the work consistently
Done properly, manual testing is more disciplined than most teams expect.
Start with a prompt list. Do not improvise. Use categories such as:
- “best” and “top” comparison prompts
- “who offers” service prompts
- local prompts with city or country names
- problem-led prompts
- alternative-to-competitor prompts
- industry-specific use case prompts
Write them as buyers would. If your website is in Polish, German, Swedish or Lithuanian, test in that language first. If you sell across borders, test the local language and English separately. The answer patterns can differ a lot.
Then set the testing conditions as tightly as you can:
- Same prompts each time
- Same language each time
- Same AI tool and mode each time
- Same date window each month
- Fresh chat where possible
- Screenshots or exports saved immediately
Use a clean browser profile if you can. Personalisation is hard to remove completely, but you can reduce carry-over effects. Prompt drift is another problem. If you ask follow-up questions in the same thread, later answers are no longer comparable to a first-pass check.
The main limits of manual testing are these.
Personalisation. Some tools may reflect account history, location, device state or prior context.
Prompt drift. Tiny wording changes can alter the result.
Lack of history. If you do not store outputs properly, you cannot prove improvement or decline later.
Weak coverage. One person usually checks too few prompts to represent the market well.
Manual work is still useful. It is often the best first step before you invest more. If your business depends heavily on one tool, you can go deeper with a tool-specific checklist, such as our guides on how to check if Gemini mentions your business and how to check if Perplexity mentions your brand.
Option 2 - spreadsheets and in-house tracking
A spreadsheet process is the middle ground. It is often the right answer for small teams that need repeatable evidence but are not ready for specialist software.
The structure can be simple.
Create one row per prompt, per tool, per language, per date. Then track fields such as:
- Date
- Tool, for example ChatGPT, Gemini or Perplexity
- Market or country
- Language
- Prompt category
- Exact prompt text
- Brand mentioned, yes or no
- Mention score, 0 to 3
- Competitors named
- Source or citation shown
- Notes on accuracy
- Screenshot link
- Tester initials
Store screenshots in a dated folder structure. If you use Google Drive or SharePoint, keep the file naming consistent. For example:
2026-09-UK-ChatGPT-prompt-07.png
That sounds minor, but it saves time later.
A good in-house process also needs a prompt library. Keep one tab for active prompts and one for retired prompts. Do not keep changing the list casually. If you need to add new prompts, label them clearly so you do not compare a six-month trend line built on different questions.
Scoring should be written down. If you have more than one person doing checks, add examples for each score level. Otherwise reporting becomes subjective.
You can also add a weighted score. For example, a prompt like “best CRM for estate agents in Ireland” may be worth more than a broad educational prompt. Weighting is useful for agencies reporting to clients with a clear commercial niche.
There are still limits.
Spreadsheets do not automate the checking itself. They organise the process, but someone still has to run the prompts and collect the evidence.
They also struggle once you add complexity:
- multiple clients
- several languages
- five AI tools
- weekly checks
- trend reporting
- publishing and SEO work tied to mention changes
That is where many teams hit the wall. They can measure mentions, but cannot connect the movement to what changed on the site, what content was published, or what authority signals improved.
Option 3 - dedicated AI visibility platforms
Specialist platforms exist because manual and spreadsheet methods break down at scale.
A dedicated AI visibility platform can automate parts of this work across ChatGPT, Gemini, Perplexity, Claude and Grok. Depending on the product, that may include:
- scheduled prompt runs
- prompt libraries by market and language
- mention detection
- competitor comparison
- scoring and trend charts
- screenshots or stored outputs
- reporting by client or country
- alerts when visibility changes
Before paying, check what the platform actually measures.
Some tools only monitor whether your domain appears in cited links. That is useful, but it is not the same as being named in the answer. Others focus only on English-language prompts, which is a weak fit for businesses selling in Danish, German, Polish, Latvian or other local languages.
For multilingual European businesses, ask these questions first:
- Can it run prompts in our actual site language?
- Can it report in that language?
- Does it support the AI tools we care about, not just one or two?
- Does it store historical evidence we can review later?
- Can we see the exact prompts used?
- Can we separate branded prompts from non-branded buyer prompts?
- Can we compare markets cleanly?
- Does it help us connect mention changes to content and authority work?
This last point matters. Tracking alone does not fix anything. If you are paying for software, it should either help you act on the findings or fit into a workflow that does.
That is the gap we built Seonis to cover. We do not just check whether ChatGPT, Gemini, Perplexity, Claude and Grok mention a business. We also research keywords in the site’s own language, write and publish articles natively, run each piece through a language gate, build authority through our member link exchange, and report the results in the owner’s language. The tracking matters because it is tied to the work that can change the outcome.
If you are an agency comparing whether to build this yourself or offer it through a platform, our partner programme for agencies shows how we structure that workflow.
How to compare tools for multilingual European businesses
If your website is not in English, tool comparison gets stricter.
A lot of AI visibility products are built around English prompts, English reporting and English-first content assumptions. That creates two problems.
First, the monitoring misses the real buyer queries in your market.
Second, the improvement work often defaults to translated content, which is where many businesses have already been burnt.
For businesses in the UK and Ireland, English may still be the main language, but local phrasing matters. “Accountants for contractors” and “accountancy for limited companies” do not behave the same. In the Baltics, Poland, Germany and the Nordics, the gap is usually larger. A translated prompt set often misses how buyers really ask.
So when comparing tools, use these criteria.
Local-language prompts
The platform should let you build and save prompts in the market’s own language. Not as a translated afterthought. As the base case.
This affects both tracking quality and the work that follows. If an AI tool never sees your brand in the language your buyers use, the report may look fine in English and weak where your revenue actually comes from.
Reporting in the owner’s language
Reporting should be understandable without a translation step. For owner-led businesses, this matters more than vendors often realise. If the monthly report arrives in a language the decision-maker does not use day to day, it gets skimmed or ignored.
Publishing workflow
Tracking software is only part of the job. Ask how the business will act on the findings.
If the tool says you are absent from buyer-relevant prompts in Finnish or German, what happens next? Can it connect the issue to content gaps, weak topical coverage, thin service pages or missing authority signals? Can it publish to WordPress, Shopify, Webflow or Ghost without turning the process into another project for your team?
Proof of change over time
This is essential. You need dated evidence.
A useful system should show:
- what prompts were checked
- what the answer looked like then
- what changed later
- what work happened in between
Without that, every report becomes anecdotal. You remember that “it seems better than before”, but you cannot prove it.
For UK businesses, this is mostly a process question. For EU businesses operating across several member states, it becomes a governance question too. Teams often need a clear record of who approved content, what language version went live, and when changes were published. If agencies are involved, that audit trail matters even more.
Coverage across tools
Do not assume one AI tool is enough. Buyer behaviour is spreading across ChatGPT, Gemini, Perplexity, Claude and Grok. Coverage does not need to be identical across all five, but your reporting should reflect where your audience is likely to ask.
Evidence, not just scores
A chart is useful. A screenshot is better. The best setup has both.
When a client or owner asks, “Were we actually mentioned?”, you should be able to show the answer, not just the percentage.
What most teams should do next
If you are early on, start manually. Pick 15 buyer-relevant prompts in your main language, run them in ChatGPT, Gemini and Perplexity, and save the outputs. That gives you a baseline.
If you need a monthly process, move to a spreadsheet before you buy software. It will force you to define prompts, scoring and evidence properly.
If you already manage several markets, several clients or several languages, skip straight to a platform. At that point, the real question is not whether software costs money. It is whether manual work is already costing more, while giving you weaker proof.
That is especially true if your site is not in English. You need local-language prompts, local-language reporting, and content that is written natively rather than translated. Otherwise you may end up tracking a problem your workflow cannot solve.
The important part is not the label on the method. It is discipline. To track brand mentions in ChatGPT, Gemini and Perplexity properly, you need fixed prompts, buyer relevance, stored evidence and a way to compare results over time. Everything else is just a nicer wrapper around that core process.