mentrel LLM visibility & brand intelligence
Get an audit
Analytics  /  AI search

How to track the traffic AI sends you

AI discovery happens in three stages: an engine crawls your pages, may make you visible in an answer, and a person may then click through. Only the crawl and the click touch your servers, so those are the two you can measure, in your logs and in Google Analytics. Here is how to read each signal without fooling yourself.

Updated Sep 10, 2026 / 14 min read / By mentrel
The short version
  • AI discovery has three stages: an engine accesses your page, may make you visible in its answer, then a person may refer themselves by clicking. Only the access and the click touch your servers, so those are the two you can measure.
  • Crawlers like GPTBot, OAI-SearchBot and ChatGPT-User do not run JavaScript, so analytics never sees them. You find them in your server or CDN logs, by user-agent.
  • A human clicking a ChatGPT citation usually arrives tagged ?utm_source=chatgpt.com and lands in GA4 as chatgpt.com / referral.
  • In GA4: Traffic acquisition to count Sessions by Session source / medium, and User acquisition to count Active users by First user source / medium. Both are the right reports.
  • Whatever GA4 shows is a floor, not a ceiling. The free app, mobile, and stripped referrers push a big share of real AI visits into "Direct."
The model behind everything below Access is not visibility. Visibility is not traffic.
Stage 1 · access

AI access

Did an engine read your page?

Seen in server / CDN logs
→
Stage 2 · the blind spot

Answer visibility

Did the engine mention, cite, or recommend you in its answer?

Happens off your servers. Invisible to logs and GA4. Needs visibility monitoring.
→
Stage 3 · referral

AI referral

Did a person click through to you?

Seen in Google Analytics

Every week someone opens their analytics, sees a line that says ChatGPT, and asks the same two questions in the same breath: is AI reading my site, and is AI sending me visitors? They sound like one question. They are not. Conflating them is the single reason most "AI traffic" reports are wrong, and getting the difference straight is most of the work.

This guide walks through both signals from the ground up: what the bots are, how to catch them, how real people from AI answers show up in Google Analytics, the exact reports to open, and the one caveat that keeps everyone honest. It is written to be the page you send someone when they ask how this actually works.

Part 01

The distinction that clears up the confusion

Recall the funnel at the top of this guide: an engine accesses your page, may make you visible in an answer, and a person may then refer themselves by clicking. Only the first and last stages leave a mark on your own servers, which is exactly why those two are the ones you can measure today.

The middle stage is the one that trips people up. It never touches your infrastructure, so no log and no analytics tool can see it, yet it is the stage that decides whether the other two ever happen. A crawl does not prove the engine surfaced you, and a click cannot happen until the engine has surfaced a link to you. Hold onto that funnel: it is the thread running through this whole guide.

Zoom into the two stages that reach your servers and you find four distinct actors. Three are machines (all of Stage 1), and one is a person (Stage 3). Only that last one is "traffic" in the sense your analytics means it.

Machines reading your pages

Server / CDN logsInvisible to GA4
GPTBot
Crawls your content so it may be used to train a future model. This is the "reading to learn" bot.
OAI-SearchBot
Builds the search index ChatGPT pulls from when it cites and links sources. Being here is how you become quotable.
ChatGPT-User
Fetches your page live, in real time, because one person asked a question and your page looked relevant right then.

A person arriving from an answer

Google AnalyticsVisible to GA4
The referral click. A human reads an AI answer, sees your link cited, clicks it, and lands on your site. This is the visit that shows up in analytics, usually carrying ?utm_source=chatgpt.com.

It helps to give the two categories names. AI access is any automated request an engine makes against your infrastructure, and it lands in your logs. AI referral is a human who clicked through from an answer, and it lands in your analytics. Almost every "how do I track this" question is really asking which of these two you mean:

AI access AI crawler / search bot / agent → your server → your logs
AI referral AI answer → human click → browser → your server + GA4

So, to answer the question you came in with directly: yes, the distinction is real. The bot that reads your page to help train a model (GPTBot) is a different actor from the fetch that happens while ChatGPT answers a live question (ChatGPT-User), and there is a third, OAI-SearchBot, that keeps the index those citations come from. All three are automated, none of them execute the JavaScript your analytics depends on, and none of them will ever appear in a Google Analytics report. This is confirmed by OpenAI's own bot documentation.

The person clicking through is the only one of the four that analytics can count. Keep that map in your head and every "how do I track this" question answers itself: bots go to your logs, humans go to your analytics.

Part 02

Meet the crawlers, and how to actually see them

Google Analytics runs as a snippet of JavaScript inside a real browser. AI crawlers request your HTML straight from the server and do not run that script. That is not a bug you can fix. It is the whole reason crawler activity is invisible to your dashboard and visible only in the raw record of who asked your server for what: your access logs.

OpenAI's bots, by the book

OpenAI publishes distinct user-agents for distinct jobs, and treating them as one thing is how people accidentally make themselves invisible to AI search. Here is the full set:

User-agentWhat it doesObeys robots.txtWhere
GPTBot/1.4 Crawls content that may be used to train OpenAI's foundation models. Yes Logs
OAI-SearchBot/1.4 Indexes pages so they can be surfaced and cited in ChatGPT's search features. Yes Logs
ChatGPT-User/1.0 Fetches a page live when a user's question or a GPT action calls for it. Not always Logs
OAI-AdsBot/1.0 Validates landing pages for ad safety and relevance. Yes Logs

Two details worth internalizing. First, ChatGPT-User is user-initiated, so OpenAI notes that robots.txt rules may not apply to it the way they do to the automated crawlers. Second, blocking GPTBot stops training, but it does not stop OAI-SearchBot or ChatGPT-User, which are the two that decide whether ChatGPT can cite and open your page at all. Block the wrong one and you quietly remove yourself from AI answers while thinking you protected your content.

The other engines you will see in your logs

ChatGPT is the largest source, but it is not the only one. Grep your logs for these too:

  • Anthropic (Claude): ClaudeBot, Claude-User, Claude-SearchBot
  • Perplexity: PerplexityBot (indexing) and Perplexity-User (live fetch)
  • Google: Google-Extended is the token that governs Gemini training, and Googlebot feeds AI Overviews
  • Microsoft: Bingbot underpins Copilot
  • Others: Bytespider (ByteDance), Applebot-Extended, Meta-ExternalAgent, Amazonbot

In a raw access log, each request wears its user-agent openly. A GPTBot hit looks like this:

access.log
51.222.253.10 - - [09/Sep/2026:14:22:07 +0000] "GET /pricing HTTP/1.1" 200 18453 "-"
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"

How you read those logs depends on your stack. If you sit behind Cloudflare, its AI Crawl Control dashboard already classifies AI crawlers and shows request volume, the paths they hit, and status codes, and it can meter or block them. Its labels map cleanly onto the three OpenAI jobs above: Cloudflare files GPTBot under AI Crawler, OAI-SearchBot under AI Search, and ChatGPT-User under AI Assistant, so these are not just our conceptual categories, they are how the edge providers model it too. On a plain Nginx or Apache server, filter the access log for the user-agent tokens above. Any log tool (a hosted service or your own pipeline) works the same way: the signal is the user-agent string, not a tracking cookie.

Trust, but verify: a user-agent is only a header

One caution before you build reports on this. A user-agent is self-reported, and scrapers routinely spoof GPTBot to borrow its reputation, so a line that says GPTBot is a claim, not proof. To confirm a hit is genuinely OpenAI, cross-check the requesting IP against OpenAI's published ranges (it lists them separately at openai.com/gptbot.json, searchbot.json, chatgpt-user.json and adsbot.json), ideally with a reverse DNS lookup too. This is what "verified bot" features do for you: Cloudflare and similar tools confirm identity by published IP ranges, reverse DNS, and emerging cryptographic signatures rather than trusting the header. It is the main reason a managed edge view is more trustworthy than grepping a raw log.

Put plainly, a log line supports some claims far more strongly than others:

The claimCan you know it?
Something requested /pricingYes, reliably, from your logs
The request said it was GPTBotYes, but it is self-reported
It was genuinely OpenAI infrastructureOnly with IP or reverse DNS verification
The page was used in an AI answerNo, not from the crawl alone
You were cited or recommendedNo, invisible to logs and to GA4
A person clicked through from AIOften, through referral and UTM analytics
!

Crawling is not endorsement, and it is not a visit.Bots crawl far more than they ever send back. It is normal to see thousands of crawl requests per referred click. A page can be crawled heavily and never cited, or cited and never clicked. Volume in your logs proves only that a crawler requested your content. It does not prove the content was indexed, used in an answer, cited, or recommended.

Part 03

The visit that actually shows up: ChatGPT referrals

When ChatGPT cites a source and a person clicks it, that click is a normal browser visit, and your analytics can see it like any other. The tell is in the URL. ChatGPT appends a tracking parameter to many of its outbound links, so the address the visitor lands on looks like this:

the link a reader clicks
https://yoursite.com/guide?utm_source=chatgpt.com

That utm_source=chatgpt.com is what lets Google Analytics attribute the visit. GA4 reads it and records the source as chatgpt.com with the medium referral, so it lands in your reports as a tidy row:

GA4  ·  Session source / medium
chatgpt.com / referral ............ 1,204 sessions

One important caveat before the how-to: the parameter is common but not universal. It shows up reliably on citations from ChatGPT's search results, and it is missing on plenty of other click-throughs, especially from the mobile app and the free tier. That is the gap we deal with in Part 06. For now, know that where the parameter is present, tracking is easy.

Part 04

The two GA4 reports, step by step

You named the two reports that matter, and you named them correctly. They answer different questions, and you want both.

A. Sessions by Session source / medium (how many visits)

This is the Traffic acquisition report. It counts sessions, so a returning visitor who comes back through ChatGPT three times counts three times. Use it to answer "how much traffic is AI sending me right now."

  1. Open Reports → Acquisition → Traffic acquisition.
  2. Change the primary dimension dropdown to Session source / medium.
  3. In the search box under the table, type chatgpt. You will see chatgpt.com / referral.
  4. Read the Sessions column. That is your ChatGPT visit count for the date range.

B. Active users by First user source / medium (how many people)

This is the User acquisition report. "First user source" is sticky: it records how a person first found you, ever. Use it to answer "how many of my users did AI introduce to me in the first place," which is the better number for judging AI as a discovery channel.

  1. Open Reports → Acquisition → User acquisition.
  2. Change the primary dimension to First user source / medium.
  3. Search chatgpt and read the Active users (or New users) column.
★

Filter by source, not medium.ChatGPT reliably passes utm_source, but often no utm_medium, so filtering on medium will drop real visits. Always match on the source (chatgpt.com). And to see exactly which of your pages AI sends people to, add Landing page + query string as a secondary dimension: you will see the destination pages, and the ?utm_source=chatgpt.com string right there in the row.

Part 05

Beyond ChatGPT: capturing the whole AI picture

ChatGPT is the giant, but Gemini, Perplexity, Copilot and a growing tail of others send real visitors too. Each has its own referring domain:

EngineReferring source in GA4
ChatGPTchatgpt.com, chat.openai.com
Google Geminigemini.google.com
Perplexityperplexity.ai
Microsoft Copilotcopilot.microsoft.com
Claudeclaude.ai
Othersgrok.com, deepseek.com, you.com, meta.ai

The shortcut: GA4's native "AI Assistant" channel

In May 2026, Google added an AI Assistant channel to GA4's default channel grouping. When it recognizes an assistant referrer it tags the medium ai-assistant and buckets the session automatically, with no setup from you. Convenient, but know its limits. Google's documentation describes the channel as traffic from sources like ChatGPT, Gemini, Deepseek, Copilot and Grok. Perplexity is not on that documented list, the list is Google-maintained and shifts over time, and the channel never captures a session that arrives with no referrer. Treat it as a useful headline, not a complete count.

The durable fix: your own AI channel group

Build a custom channel so you catch every engine, including the ones Google's list forgets, and any that launch next year. In Admin → Data display → Channel groups, create a group called "AI" with a rule that matches the source by regular expression:

Source matches regex
Session source matches regex:
chatgpt|openai|perplexity|gemini|bard|claude|copilot|bing.*chat|grok|deepseek|you\.com|meta\.ai

Now everything AI-driven collapses into one line you can trend over time, compare against organic search, and set conversions against. And wherever you control the link (a newsletter, your docs, a partner placement an AI might quote), add your own UTM tags as a failsafe, so attribution never depends on the engine choosing to pass a referrer.

Part 06

Why your number is a floor, not a ceiling

This is the part most guides skip, and it is the part that will save you from a bad decision. A large share of genuine AI referrals never carry a referrer at all, so GA4 files them under Direct, right next to people who typed your URL by hand. It happens because:

  • The free tier and the mobile apps frequently send no referrer information.
  • Some links are rendered with rel="noreferrer", which strips the source on the way out.
  • Privacy settings and in-app browsers drop referrer data as a matter of course.

The practical consequence is simple and important: your real AI-driven traffic is always higher than any report shows. Never look at an undercounted chatgpt.com / referral row and conclude AI "is not worth it." You are looking at the visible tip.

You cannot recover the exact figure, but you can triangulate the gap:

  • Watch for unexplained rises in Direct traffic to deep, specific pages. Nobody types a long guide URL from memory; that pattern is often unattributed AI or dark social.
  • Correlate with your logs. A climb in OAI-SearchBot and ChatGPT-User activity that precedes a Direct-traffic bump is a strong tell.
  • Add a "How did you hear about us?" field to key forms. Self-reported attribution catches what the referrer header lost.
  • Track branded search lift. People who meet you inside an AI answer often go and search your name afterward.
Part 07

Your tracking playbook

Everything above, compressed into a checklist you can actually run this week.

  • Bots go to logs. Set up a view of your server or CDN logs filtered for AI user-agents, so you can see crawl coverage over time.
  • Humans go to GA4. Bookmark the Traffic acquisition report on Session source / medium, filtered to chatgpt.
  • Add the discovery view. Check User acquisition on First user source / medium to see who AI introduces.
  • Build the "AI" channel group so every engine trends as one line, not scattered referral rows.
  • Add Landing page + query string to learn which pages AI actually sends people to, then make more of those.
  • UTM-tag every link you control, so stripped referrers cannot hide the traffic.
  • Watch Direct traffic to deep pages as your proxy for the untracked AI share.
  • Review monthly. This channel is moving fast; a number from six months ago is a different channel.
Part 08

What tracking cannot tell you

Tracking is a rear-view mirror. It counts the clicks you already earned. It is genuinely useful, and it has a hard ceiling: it can tell you how many people arrived, but never why the engine chose you, or a competitor, in the first place.

The questions that decide your AI traffic all sit upstream of the click. When a buyer asks ChatGPT for "the best tool for X," does your brand come up, or does a competitor own the answer? Which of your pages get cited, and which get read and ignored? What would it take to become the recommendation instead of a footnote? None of that is in your analytics, because it happens before anyone ever clicks.

This is the funnel from the very top of the guide, now with the stakes clear. AI visibility is three layers, and each is measured somewhere different:

  • AI access, measured in your logs: is an engine reading your site at all?
  • Answer visibility, measured by visibility monitoring: does the engine mention, cite, or recommend you when someone asks?
  • AI referral, measured in GA4: do those answers ultimately send you humans?

This guide covers the first and third, because they are the two that touch your servers. The middle layer is the one that decides the other two, and it hides from both your logs and your analytics: a crawl does not prove the engine surfaced you, and referral traffic only exists after the engine has surfaced a link to you. To know whether you are the answer, you have to go and ask the engines directly.

Microsoft has started to report a slice of that middle layer for its own engines. Our guide to Bing's AI Performance report covers what its citation data shows about Copilot, and what it leaves out.

Where mentrel comes in

Tracking tells you the score. We tell you how to change it.

mentrel runs the questions your buyers actually ask through the major answer engines, shows you where you surface and where a competitor owns the answer, and maps exactly what to change so you get cited more often. If your analytics is finally showing AI traffic, this is how you grow it on purpose.

See what AI says about your brand
Part 09

Frequently asked questions

Does ChatGPT actually send real traffic to websites?

Yes. When ChatGPT cites a source and a person clicks the link, that is a real visit. Those clicks usually carry utm_source=chatgpt.com and appear in GA4 as chatgpt.com / referral. The volume varies by site and is undercounted, because many AI visits arrive without a referrer and get filed as Direct.

Is GPTBot the same bot that answers questions?

No, and the difference matters. GPTBot collects data that may train future models. ChatGPT-User fetches a page live while answering one person's question. OAI-SearchBot builds the search index that ChatGPT cites from. Three separate user-agents, three separate jobs.

Why does my AI traffic show up as "Direct" in Google Analytics?

Because the visit arrived without a referrer. The free ChatGPT tier, the mobile apps, and links marked rel="noreferrer" commonly strip the source, so GA4 has nothing to attribute and defaults to Direct. It is the main reason your true AI traffic is higher than your reports say.

Can I see AI crawlers in Google Analytics?

Not in Google Analytics, but yes in your logs. Crawlers do not execute the JavaScript GA4 relies on, so they never reach it. Their HTTP requests are observable, though, in your server or CDN access logs, or in an edge tool like Cloudflare's AI Crawl Control. Filter by user-agent for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot and the rest. Note the word: you observe that a request happened, not what the model later did with the response.

Can I trust a "GPTBot" user-agent in my logs?

Not on its own. A user-agent is self-reported and easy to fake, and scrapers spoof GPTBot to look legitimate. To be sure a request is really OpenAI, cross-check the source IP against OpenAI's published ranges (gptbot.json, searchbot.json, chatgpt-user.json) and a reverse DNS lookup. Verified-bot features in tools like Cloudflare run that check for you.

Should I block GPTBot?

It is a trade-off, so decide by goal. Blocking GPTBot opts you out of model training, but it does not stop ChatGPT from citing or opening your pages, which is governed by OAI-SearchBot and ChatGPT-User. Block those two and you can quietly disappear from AI answers, so block deliberately, not by reflex.

How do I see which pages ChatGPT sends people to?

In GA4's Traffic acquisition report, filter Session source to chatgpt.com, then add Landing page + query string as a secondary dimension. You will get a table of the exact destination pages and their session counts.

What about Perplexity, Gemini and Copilot?

Track them the same way, by their referring domains (perplexity.ai, gemini.google.com, copilot.microsoft.com). Google does not publish a complete, static list of every source its built-in AI Assistant channel covers (its documented examples are ChatGPT, Gemini, Deepseek, Copilot and Grok, and that list shifts), so do not treat it as an exhaustive count. Build a custom "AI" channel group where you explicitly include Perplexity, Claude, and any other engine you care about. Google's own custom-channel help even uses Perplexity and Claude as examples.