How to Measure AI Visibility: Tools and the Manual Method
AI visibility measures something entirely different from the traffic reaching your site. Measuring AI traffic in GA4 looks at how many people came to your site from ChatGPT or Perplexity; AI visibility doesn’t measure traffic — it measures mentions: how often, in what context (positive, negative, neutral), and with what accuracy your brand shows up in the answers these tools generate. There are two ways to measure this: newly emerging tracking tools, and a simple manual method you can apply yourself. The second one remains the most reliable starting point while the field is still this new.
Why AI Visibility Matters, What It Gets You
Users now ask a lot of questions directly to ChatGPT, Perplexity, or Gemini instead of Google — and these tools mention some brands while generating the answer and skip others entirely. A prospective customer who never visits your site but asks “who’s the best X expert” and gets your name in the answer — that’s a win; if your name doesn’t appear, that’s a loss — but neither shows up in any analytics report.
This measurement is the real outcome indicator of GEO work: the only concrete proof of whether the technical and content work described in the GEO Guide is actually paying off is how often your brand shows up in AI answers. For the mechanics behind this kind of mention, read my piece on Brand Mentions: Why They’re Stronger Than a Backlink — it explains why a mention with no link is still a strong signal. Here I’ll focus on the actual question: how do you measure this mention?
The Manual Method: How to Test It Yourself
The manual method isn’t complicated, but it takes discipline. Four steps:
- List 10-15 typical questions in your industry. These should be general questions a real customer might actually ask, phrased like “how do I choose the best X,” “how much budget should I set for X,” “who’s a recommended provider for X service” — and they should not contain your own brand name. Two examples from my own list: “agency or freelancer for Google Ads management,” “how much does AI consulting cost.”
- Ask these questions monthly to ChatGPT, Perplexity, and Gemini, one by one. Ask the same question to all three — one might mention your brand and another might not, and that gap alone is useful information.
- Note three things for every answer: whether your brand appeared (yes/no), in what context it appeared (positive / negative / neutral — was it just listed by name, or characterized with something like “X is reliable for this”), and which competitors appeared in the same answer.
- Log the results in a simple table. No complex tool needed, a Google Sheets tab is enough.
Your tracking table might look like this:
| Question | ChatGPT | Perplexity | Gemini | Context | Competitors mentioned |
|---|---|---|---|---|---|
| “which expert/agency is recommended for…” | Yes | No | Yes | Neutral — just listed by name | 2 competitor names appeared |
Repeated monthly, this table turns into a time series: you can see which month you showed up in which tool, and how the context changed. It’s not a fancy dashboard, but it’s real, and you verified it with your own eyes.
Newly Emerging Tools: How Reliable Are They
Over the past year, dozens of new tools have appeared under names like “AI visibility tracking” or “brand mention tracking” — some free, some paid SaaS subscriptions. These tools generally auto-submit your chosen queries to multiple AI models on your behalf, scan the answers, and produce a “visibility score” or “mention count.”
To be honest: this space is so new that no tool’s coverage or reliability has settled yet. Asking the same query twice on the same day can produce a different answer (AI models aren’t deterministic), how often the tools actually access which model isn’t transparent, and the methodology behind whatever “score” they report usually stays inside the product’s own black box. It wouldn’t be surprising to see one tool say “40% visibility” for a brand while another says “12%” for the same brand.
This doesn’t mean these tools are useless — if manually tracking a large volume of queries isn’t practical, they can serve as a starting point. But none of them currently gives you a “definitive” number; they’re all an estimate, a sample. The manual method is slower, but since you verify everything you see with your own eyes, it remains the most reliable starting point — especially for a small business, half an hour every few months gets you the same information without a subscription fee.
How Often Should You Test
Monthly checks are enough for most small businesses. AI models’ training data and real-time crawling layers don’t change weekly; thirty days is a reasonable interval to see whether a meaningful shift has occurred.
You should tighten this interval in two cases: if you’ve launched a new product or service (it takes time for AI to “learn” this and reflect it in answers, checking more often lets you track that), or if your prices change frequently (AI mentioning you with an old price is worse than not mentioning you at all — visibility with wrong information doesn’t count as a win). In these cases, checking every two weeks is more accurate.
Increasing test frequency doesn’t automatically improve data quality — what matters most is asking the same questions, in the same format, every time. Consistency matters more than frequency.
Something I Noticed Doing the Manual Test
I’ve been running this manual test regularly for a few clients for a few months now (I don’t name them here, that’s my policy). I noticed something odd: asking the same question twice on the same day, back to back, one tool would mention the brand once and not the other time — with nothing actually having changed. The first time I saw it, I thought it was a bug; after repeating it a few times, I realized this is an inherent inconsistency in the models: the same prompt, on the same day, can produce a different answer.
The practical implication: asking a single question once and concluding “I’m not appearing, there’s a problem” or “I’m appearing, great” is misleading. Don’t draw a conclusion from a single sample without repeating the same question a few times — which is why, when filling out the table above, I now ask each question at least twice looking for a consistent result.
You Can Ask Me to Do This
You can ask me to do this: work that typically takes 2 hours to 2 weeks, remote, billed hourly — I’ll work out 10-15 industry-specific questions with you, run the first manual scan, and hand you a table you can track monthly. See my AI Search Visibility service page for details, or write to me via the contact page.
Frequently Asked Questions
Is AI visibility the same thing as measuring AI traffic in GA4?
No. Measuring AI traffic in GA4 shows how many visitors came to your site from ChatGPT or Perplexity — meaning someone clicked through. AI visibility measures whether your brand appears in the answer that tool generates, even if nobody clicked at all. One measures traffic, the other measures mentions; they’re complementary but answer different questions.
Is there a free, reliable AI visibility tracking tool?
Not currently — most tools on the market are paid subscriptions, and none gives a “definitive” number (as explained above). For a small business, the cheapest and most transparent method is running the manual test above yourself, monthly; the cost is half an hour of your time.
How many questions should I ask, and should I use ChatGPT/Perplexity/Gemini all three?
10-15 questions are enough if they reasonably cover your industry — more makes the monthly check unsustainable. I’d recommend using all three tools, because they can give different results for the same question; looking at just one gives you an incomplete picture.
How do I find out if my competitors show up in AI answers?
While asking the same questions, read the full answer and note which brand names appear — AI answers usually list several options side by side. Adding a “competitors mentioned” column to the tracking table above lets you see over time who’s rising and who’s disappearing.