Clicktag docs

AI crawlers

Clicktag tracks AI crawlers like GPTBot, ClaudeBot and PerplexityBot through a read-only Cloudflare token: 30 days imported on connect, then a sync every 15 minutes, with no code on your site.

Why your tracker can't see them

GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot fetch the raw HTML of a page and never run JavaScript, so the Clicktag script never fires for them and they never show up as visits. Search engine crawlers that do run it are recognized as bots by their user agent and left out of your visits too. Cloudflare sits in front of your site and sees every request, crawlers included. Clicktag reads those requests from Cloudflare's analytics and shows them on the Crawlers tab of your dashboard.

Connect Cloudflare

Your site must be verified in Clicktag and its domain proxied through Cloudflare, since Cloudflare only sees traffic it proxies.

  1. Open your site's settings, go to Integrations and find Cloudflare.
  2. Click Create a read-only token. Cloudflare opens its token page with Zone Read and Analytics Read already selected for all your zones. Create the token and copy it.
  3. Paste it into Paste the token and click Connect.

Connecting starts tracking on every verified site whose domain is in the token's zones, as long as the token can read that zone's analytics. A site matches the zone named after its hostname or a parent domain, so blog.example.com matches the zone example.com. Sites you add later get a Start tracking button on their Crawlers tab.

Only requests to the site's own hostname and its www twin count: example.com also counts www.example.com, but not other subdomains in the same zone.

Clicktag checks the token before saving it and stores it encrypted. It only ever shows the last 4 characters. One token covers every site you own; Replace token swaps it and Disconnect removes it.

What counts as verified

A hit is verified when Cloudflare verified the bot, in its AI Crawler, AI Search, AI Assistant or Search Engine Crawler category, or when the request came from an IP range the bot's company publishes. Clicktag checks the lists of OpenAI, Perplexity, Google, Microsoft, Apple, Common Crawl and DuckDuckGo. The IP check happens during the sync and IP addresses are never stored.

Cloudflare no longer verifies PerplexityBot. On our own sites, about 90% of the PerplexityBot hits Cloudflare didn't verify came from Perplexity's published ranges, so those count as verified. Anthropic publishes no ranges, so Claude's bots only count as verified when Cloudflare verifies them.

A request that claims to be a known bot, like GPTBot or Googlebot, without being verified counts as unverified. Unverified claims of GoogleOther, ExaSearchBot or Baiduspider are left out. Those hits are kept apart from the verified numbers, because anyone can put GPTBot in a user agent: on our sites, unverified hits claiming OAI-SearchBot, GPTBot, ChatGPT-User or Googlebot came from IPs outside those companies' published ranges.

Bots in Cloudflare's other categories, like SEO tools, webhooks, link previews, accessibility, advertising and security bots, are left out.

Bots we recognize

Bot Company Type
GPTBot OpenAI AI training
OAI-SearchBot OpenAI AI search
ChatGPT-User OpenAI AI assistant
ClaudeBot Anthropic AI training
Claude-SearchBot Anthropic AI search
Claude-User Anthropic AI assistant
PerplexityBot Perplexity AI search
Perplexity-User Perplexity AI assistant
Googlebot Google Search engine
GoogleOther Google AI training
Bingbot Microsoft Search engine
Applebot Apple AI search
meta-externalagent Meta AI training
meta-externalfetcher Meta AI assistant
Amazonbot Amazon AI training
Bytespider ByteDance Search engine
CCBot Common Crawl AI training
DuckDuckBot DuckDuckGo Search engine
DuckAssistBot DuckDuckGo AI assistant
MistralAI-User Mistral AI assistant
ExaSearchBot Exa AI search
Baiduspider Baidu Search engine

Googlebot includes Googlebot-Image, Googlebot-Video and Googlebot-News. Other bots Cloudflare verifies, such as PetalBot or YandexUserproxy, appear under the name in their user agent, with the type of their Cloudflare category. A verified request with a generic user agent shows as Unidentified.

What the numbers mean

  • AI crawler hits: verified hits from AI search, AI assistant and AI training bots.
  • Search engine hits: verified hits from search engine crawlers.
  • Pages crawled: distinct paths with at least one verified hit.
  • Unverified hits: hits from requests that claim a known bot but aren't verified.

Each total is compared with the same clock times moved back by as many calendar days as your range covers: today until 12:35 against yesterday until 12:35, the last 7 days against the 7 days before. The chart splits verified hits by type; search engines start hidden, click them in the legend to show them.

The Bots panel lists every bot with its verified hits and when it was last seen, the latest hour with a verified hit. A blocked count shows verified hits your site answered with 401, 403 or 429, so a firewall rule or bot setting that turns crawlers away is easy to spot. An unverified count shows the hits that only claimed to be that bot. Pages lists the top 100 paths with how many different bots read each one, and Responses shows the status codes verified crawlers got. Click a bot or a page to filter the whole tab.

Hit counts are Cloudflare's sample-adjusted estimates.

Sync and history

Limit Value
Sync Every 15 minutes
History imported on connect Last 30 days
Cloudflare's own retention 31 days
Clicktag retention Kept after Cloudflare drops it
Granularity 1 hour
Bots and pages per report Top 100
Cloudflare plan Any, including Free

Connecting imports the last 30 days, a week at a time. The Crawlers tab shows the progress and the numbers appear once the import is done. After that, a sync every 15 minutes adds the newest hours and reads the previous hour again to catch late data. Sync now in site settings queues one right away.

Cloudflare keeps this data for 31 days. Clicktag stores it per hour, so your crawler history keeps growing after Cloudflare has dropped it.

If Cloudflare stops accepting the token, the site shows Needs attention with the reason. Replace the token and the sync picks up where it stopped.

Stop tracking and disconnect

Stop tracking in a site's settings stops crawler tracking on that site and deletes its crawler history. Disconnect removes your Cloudflare token, stops tracking on all your sites and deletes their crawler history. Both ask before they delete anything. Tracking a site again imports the last 30 days again.

API

All routes are owner-only. GET /api/sites/{hostname}/crawlers/status and GET /api/sites/{hostname}/crawlers/overview also accept a management key with read, and POST /api/sites/{hostname}/crawlers/sync one with manage. The other routes need the owner's Firebase ID token. Someone else's site answers 404, and the overview and starting tracking need a verified site (403 otherwise).

  • GET, PUT and DELETE /api/cloudflare/connection: read, save ({ "token" }) or remove the Cloudflare token. GET and PUT return { "connected": false } or connected: true with token_hint (the last 4 characters), zones (how many zones the token could see when it was saved), created_at and updated_at; DELETE returns { "disconnected": true }.
  • GET /api/sites/{hostname}/crawlers/status: connection, zone and sync state: connected, linked, zone ({ id, name }, also shown before tracking starts when the token sees one; that lookup is kept for up to 5 minutes, or until the token is saved again), sync_state (backfill, live or error), backfill_from, backfill_cursor, synced_until, last_sync_at and last_error.
  • POST and DELETE /api/sites/{hostname}/crawlers/link: start tracking a site (returns its status), or stop and delete its history ({ "unlinked": true }).
  • POST /api/sites/{hostname}/crawlers/sync: queue a sync (202); during the import it continues from where it stopped. A site in error gets 409 with "revoked": true until the token is replaced.
  • GET /api/sites/{hostname}/crawlers/overview?from&to&tz: totals, previous period, series, bots, pages and status codes. from and to are unix seconds, to exclusive; hits count per UTC hour, and every hour that overlaps the range counts in full. Filter with f=bot:GPTBot, f=vendor:OpenAI, f=purpose:ai_search or f=path:/pricing, one per key; purpose is ai_search, ai_assistant, ai_training or search.

The overview returns granularity (hour for ranges up to 4 days, else day), totals and previous (ai_hits, search_hits, pages, unverified), series (t with verified hits per ai_search, ai_assistant, ai_training and search), bots (the top 100 by verified hits, with bot, vendor, purpose, hits, pages, unverified, blocked and last_seen), pages (the top 100 paths by verified hits, with path, hits, bots and last_seen; pages_capped says there are more), statuses (status and hits) and synced_until. last_seen is the Unix start of the latest hour with a verified hit.

The full schemas are in the OpenAPI file.

Errors

Status Error
400 invalid range, invalid filter, That doesn't look like a Cloudflare API token., Cloudflare rejected this token. Create one with Zone Read and Analytics Read., This token can't see any zones., This token can list zones but can't read analytics. Add Analytics Read., Connect Cloudflare first., Your Cloudflare token can't see a zone for <hostname>., Cloudflare analytics is turned off for <zone>.
401 sign in required, invalid or revoked management key
403 Install your tracking script to verify this site first, management keys cannot call this endpoint, scope <scope> required
404 unknown site, Crawler tracking is not set up for this site., Cloudflare is not connected.
405 method not allowed on /api/cloudflare/connection
409 Cloudflare can no longer read this zone. Reconnect Cloudflare in site settings. from POST .../crawlers/sync on a site in error
502 Could not reach Cloudflare. Please try again. when Cloudflare fails or times out while saving a token or starting tracking
503 Sync queue unavailable. Please try again., Linked, but sync could not start. Use Sync now to retry.

FAQ

Do I need to add code to my site?

No. Clicktag reads Cloudflare's request analytics with a read-only token. Nothing changes on your site or in your Cloudflare settings.

Does it work on the Cloudflare Free plan?

Yes. The analytics Clicktag reads are available on Free zones, with 31 days of history.

Why does Connect reject my token?

The token needs Zone Read to list your zones and Analytics Read to read requests. Connect checks both and says what is wrong: Cloudflare rejected the token, the token can't see any zones, or it can't read analytics. If Cloudflare itself fails, it asks you to try again.

Why does my site have no Cloudflare zone?

The token can't see a zone for that domain. Create a token that covers it and use Replace token in the site's settings.

What happens to my data when I disconnect?

Clicktag deletes the crawler history of every site that used the token. Your web analytics are not touched.