AI Search

How to Measure Generative Engine Optimization

The four GEO metrics that actually exist: AI Overview citations, ChatGPT and Perplexity mentions, AI referral traffic and crawler logs, tracked for free.

Published 7/31/2026By Maxence Vanderswalmen
Analytics dashboard used to track AI search visibility

How does generative engine optimization work?

Every AI answer engine runs the same three-step loop: it crawls and indexes content, retrieves candidate passages when someone asks a question, then generates an answer that cites a handful of sources. Google AI Overviews, ChatGPT Search, Perplexity and Microsoft Copilot differ in the details, but retrieval plus citation is the mechanic underneath all of them. Generative engine optimization is the work of becoming one of those cited sources more often.

Measurement maps onto each step of that loop. You can verify that engines crawl you (server logs), verify that they cite you (manual answer checks), and verify that citations send visitors who buy (analytics). If you want the full background before the metrics, start with our primer on what generative engine optimization is and whether you should care. This article assumes you have decided to care and now want proof that the work moves anything.

GEO measurement and tracking: the four metrics that exist

Strip away vendor language and there are exactly four things you can observe from outside the engines: citations of your pages in Google AI Overviews, mentions of your brand in ChatGPT and Perplexity answers, referral traffic from AI engines in your analytics, and AI crawler activity in your server logs.

Nothing else is measurable today. There is no impression report from OpenAI, no ranking API from Perplexity, no citation dashboard inside Search Console. Every tool on the market, free or paid, is built on top of those four raw signals. Understanding each one before you spend anything is the cheapest insurance against buying a dashboard that reformats data you already had.

Metric 1: citations in Google AI Overviews

Search Console counts impressions and clicks from AI Overviews inside the regular Performance report, but offers no filter to isolate them. A page cited in an overview and a page ranking in the classic results look identical in the export. The only reliable way to observe citations is to run your priority queries yourself and record what comes back.

Do it in a clean state: logged out or in a fresh profile, from your target country, on desktop and mobile. Record three things per query: does an AI Overview appear at all, is your domain cited, and which page. Repeat weekly on the same day. Overviews change between runs, sometimes between refreshes, which is exactly why a one-off check tells you nothing and a 12-week series tells you a lot.

Metric 2: mentions in ChatGPT and Perplexity

Same protocol, different engines. Ask each engine the buying-intent questions your customers actually ask, such as "best insulated running vest for winter" or "is this brand legit", then log whether your brand is named and whether your pages appear as sources. Perplexity shows numbered citations on every answer, which makes logging fast. ChatGPT Search lists sources when it browses the web; when it answers from model memory instead, log the brand mention on its own, because a recommendation without a link still shapes the shortlist.

Phrase the queries the way shoppers do, not the way marketers do. And keep the list stable from week to week: the value of this tracker comes from asking the same questions repeatedly, not from covering every conceivable question once.

Metric 3: AI referral traffic in your analytics

GA4 already records visitors who click through from AI engines. Filter session source for chatgpt.com, perplexity.ai, copilot.microsoft.com and gemini.google.com, then group them into a custom channel so the trend is visible in standard reports. One caveat matters here: clicks from AI Overviews arrive as ordinary google organic traffic, so this segment only covers the standalone engines, not Google's own AI layer.

Expect small absolute numbers. For most ecommerce stores this segment is a sliver of what organic search delivers. Small does not mean worthless: watch conversion rate and revenue per session on the segment, because a visitor arriving from an AI answer tends to land late in their research, with the comparison work already done. Judge the segment on quality first and volume second.

Metric 4: GPTBot, ClaudeBot and PerplexityBot in your logs

Your server logs answer a prior question: can AI engines read you at all? Grep your access logs, or your CDN analytics, for the documented user agents: GPTBot and OAI-SearchBot from OpenAI, ClaudeBot from Anthropic, PerplexityBot from Perplexity. On-demand agents such as ChatGPT-User and Perplexity-User show up when a real user's question triggers a live fetch of your page, which makes them the most interesting lines in the file: each hit is a human asking something your page might answer.

Check robots.txt and your CDN bot protection before reading too much into low counts. Plenty of stores block AI crawlers by accident through an old catch-all rule or an aggressive firewall default. Zero hits across every AI crawler almost always points to blocked access rather than absent demand, and restoring access comes before any content work.

A simple weekly tracking setup with no paid tool

One spreadsheet, four tabs, about thirty minutes a week. Tab one: 20 priority queries checked weekly, with columns for AI Overview present, domain cited, page cited, plus brand mentioned and page cited for ChatGPT Search and Perplexity. Tab two: a monthly GA4 export of the AI referral segment with sessions, conversion rate and revenue. Tab three: monthly crawler hit counts per bot, pulled from logs or CDN analytics. Tab four: a change log of everything you shipped on the site, so that when a metric moves you have candidate explanations instead of guesses.

Track citation rate across the whole query set rather than any per-query rank. Cited or not cited is the signal. After 90 days you have a baseline, a trend line, and a defensible answer to whether GEO deserves more of your roadmap, which is more evidence than most of your competitors will ever collect.

Do you need an AI visibility product like Athena?

A wave of AI visibility products has formed around GEO, with Athena among the names that surface in searches for the category. The pitch is broadly the same across vendors: run your prompt list at scale across engines, every day, from multiple locations, and log citations, mentions and share of voice automatically. That is real leverage once your query list outgrows half an hour of manual checking, or once you need competitor benchmarks across markets.

Buy the automation only after the manual tracker has proven the channel matters for your store. A paid dashboard sitting on top of a segment that sends a handful of visits a month is reporting theater. The spreadsheet costs nothing and answers the only question that matters at the start: are we visible at all, and is it trending anywhere?

The honest limits of GEO measurement

Every metric above is a sample, not a census. Generated answers vary by session, account history, geography and time of day, so your weekly check observes one draw from a distribution, not a fixed ranking. Demand data does not exist either: no AI engine publishes how often your priority questions get asked, so there is no equivalent of search volume to size the opportunity before you invest.

Attribution leaks too. A shopper who reads about your product in a ChatGPT answer and buys three days later through a brand search is recorded as organic or direct, never as AI. Some of your best GEO outcomes will be invisible in the referral data, and some of your AI referral revenue would have arrived anyway.

So treat GEO metrics as directional evidence and keep them next to channels you can measure precisely. Paid search remains the cleanest revenue signal in most ecommerce accounts, and the right benchmark for any budget reallocation. If that side of your measurement needs the same rigor, our Google Ads agency page details how we run it.

FAQ: measuring GEO in practice

Can Search Console isolate AI Overview traffic?

No. Impressions and clicks from AI Overviews are counted inside the standard Performance report with no dedicated filter, so they cannot be separated from classic results. Manual query checks remain the only way to confirm your citations.

How long until GEO tracking shows a real trend?

Plan for 90 days of weekly checks before reading anything into the data. Generated answers are volatile from week to week, so the trend across a stable query set is the signal and any single snapshot is noise.

Do I need a paid tool to measure generative engine optimization?

Not to start. A spreadsheet tracker, a GA4 referral segment and a monthly look at crawler logs cover all four observable metrics. Paid AI visibility products earn their fee when your query list, market coverage or competitor benchmarking outgrow manual checking.