How-to guide · technical
How to check whether ChatGPT, Perplexity and Gemini can actually read your website
Before you spend a penny on AI search optimisation, find out whether the AI crawlers can even get in. This is a 20-minute check you can do yourself, and in our audits it fails more often than any content problem.
If an AI engine cannot fetch your page, or fetches it and finds an empty shell, nothing else on this site matters. You can have the best content in your sector and never be cited. In the audits we run, access and rendering problems are the single most common reason a business is missing from AI answers, and they are usually accidental.
This guide walks through the four layers where sites fail: robots.txt, the firewall or CDN, the HTML itself, and the engines' own behaviour. You need no special tools beyond a browser and a terminal.
Step 1: Read your robots.txt with AI crawlers in mind
Open https://yourdomain.com/robots.txt. Look for the AI user agents by name, and for blanket rules that catch them by accident.
The important distinction is between three jobs a bot can do:
- Training crawlers collect content to train models: GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Google's opt-out token for Gemini training), Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider.
- Search and retrieval crawlers build the index that AI answers are drawn from: OAI-SearchBot (ChatGPT search), Claude-SearchBot, PerplexityBot, Bingbot (ChatGPT's live search leans heavily on Bing's index), Googlebot (AI Overviews and AI Mode use the normal Google index).
- User-triggered fetchers load a page live when someone asks about it: ChatGPT-User, Perplexity-User, Claude-User.
If you block a search or retrieval crawler, you have opted out of that engine's answers. If you block only a training crawler, you have made a different decision that does not affect citations. Plenty of robots.txt files block "GPTBot" thinking that means "OpenAI", when it only means training. Note also that a User-agent: * block with Disallow: / catches every AI crawler that does not have its own rule.
Step 2: Check the firewall and CDN layer
Many sites are open in robots.txt and closed at the edge. Cloudflare started blocking AI crawlers by default on new domains in mid-2025, and other CDNs and WordPress security plugins added similar toggles. The robots.txt looks fine, the bot gets a 403 anyway.
Where to look:
- Cloudflare: Security, then Bots, and check the AI crawler and "block AI bots" settings. Also review any WAF rules that match on user agent.
- Vercel, Netlify, Fastly, Akamai, AWS WAF: check managed bot rules and any rule sets you or an agency added.
- WordPress: Wordfence, Sucuri, All In One Security and similar plugins all have bot-blocking features. Search their settings for "AI" or "bot".
- Hosting panels: some hosts (and some managed WooCommerce or Shopify apps) block by user agent at the server level.
If you are not sure what is happening at the edge, Step 5 (logs) will tell you.
Step 3: Fetch a key page as a plain client and look at the raw HTML
AI crawlers fetch the raw HTML and read what is in it. They do not run your JavaScript. So the test is simple: does the fact you care about appear in the raw response?
In a terminal:
curl -s -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)" https://yourdomain.com/pricing | grep -i -c "£"
Swap pricing for your services page, your about page and your contact page, and swap the search term for something that should be there (a price, your phone number, a service name). A count of zero means the crawler would not see it.
If you would rather use the browser: right-click, View Page Source (not Inspect). View Source is the raw HTML. Inspect shows the page after JavaScript has run, which is what Google sees and what most AI crawlers do not. Search the source for the same facts.
A quicker version: open DevTools, press Ctrl+Shift+P (Cmd+Shift+P on Mac), type "Disable JavaScript", press Enter, then reload. Whatever disappears is invisible to ChatGPT, Perplexity and Claude.
Step 4: Ask the engines directly
The engines will tell you what they can see. Run these prompts:
- In ChatGPT with search enabled: "Read https://yourdomain.com/services and tell me what services this company offers, where it is based and any prices listed."
- In Perplexity: the same prompt.
- In Claude: the same prompt.
- In Gemini and Google AI Mode: "What does [your company name] do and where are they based?"
Compare the answers with your actual page. Three outcomes are common. The engine reads it correctly (good). The engine says it cannot access the page or that the page appears to have no content (access or rendering problem, go back to Steps 1 to 3). The engine answers from third-party sources such as a directory listing or a review site rather than from your page (your page is not being retrieved or not being trusted; the other guides on this site cover that).
Step 5: Check your server or platform logs for crawler hits
Logs are the ground truth. Search the last 30 days of access logs for the user agents listed in Step 1. On Vercel, use the Logs tab and filter by user agent; on WordPress hosts, download the raw access log from the control panel; on Cloudflare, use Security Analytics filtered by bot.
What you want to see: 200 responses for OAI-SearchBot, PerplexityBot, ClaudeBot or Claude-SearchBot and Bingbot on your key pages. What indicates a problem: 403 or 429 responses for those agents (blocked at the edge), or no requests at all across a month (nothing is discovering you, which usually means poor indexing in Bing, no external links, or a block upstream).
One warning: anyone can fake a user agent string. If you see odd volumes, verify the IP ranges against the published lists (OpenAI, Anthropic and Perplexity all publish theirs) before drawing conclusions.
Step 6: Run the free audit and record a baseline
Run your domain through the AI Search & SEO Audit on this site. It checks robots.txt, sitemap, structured data and how clearly an assistant can describe your business, and it gives you a score to measure against. Save the report. When you have fixed what this guide surfaced, run it again.
How to know it worked
You have a robots.txt that allows the search and retrieval crawlers you want. Your logs show 200 responses for those agents. Your key facts appear in raw HTML on your five most important pages. ChatGPT, Perplexity and Claude can summarise those pages accurately when given the URL.
Common mistakes
- Blocking GPTBot and assuming that is a "ChatGPT" decision. It is a training decision only. OAI-SearchBot controls ChatGPT search visibility.
- Trusting Search Console. Google renders JavaScript, so a site can look perfect in Search Console and be blank to every other engine.
- Fixing robots.txt and forgetting the firewall. The 403 at the edge wins.
- Testing the homepage only. Pricing, services and contact pages are the ones AI engines need most, and they are the ones most often built with JavaScript widgets.
Frequently asked questions
Do AI crawlers run JavaScript?
+
No. The Vercel and MERJ log study of hundreds of millions of crawler requests found no evidence that GPTBot, ClaudeBot or PerplexityBot execute JavaScript, and follow-up tests through 2026 report the same. Gemini is the exception because it uses Google's rendering infrastructure. If your key facts only appear after scripts run, most AI engines never see them.
Which user agents should I look for in my logs?
+
OAI-SearchBot, ChatGPT-User and GPTBot (OpenAI), PerplexityBot and Perplexity-User (Perplexity), ClaudeBot, Claude-SearchBot and Claude-User (Anthropic), Googlebot and Google-Extended (Google), Bingbot (Microsoft, which ChatGPT search leans on), plus Applebot, Amazonbot and Meta-ExternalAgent.
My robots.txt allows everything. Am I fine?
+
Not necessarily. Cloudflare and other CDNs can block AI bots at the firewall before robots.txt is ever read, and JavaScript-rendered pages can be technically reachable but empty. Check all three layers: robots.txt, the edge, and the raw HTML.
Sources
- Vercel: The rise of the AI crawler · vercel.com
- Anagram: AI crawler user-agent list 2026 · anagram.ai
- OpenAI: Overview of OpenAI crawlers · platform.openai.com
Richard Daniel
Automation and Delivery Lead, Emerging Group
Richard leads automation and delivery across the Emerging group, working with EDP on client websites and with ETT on enterprise AI and process automation. He is the person who turns an audit finding into a working fix: crawler access, rendering, tracking, structured data and the plumbing that most marketing teams never see. He writes the technical guides on this site.
Read next
How to get indexed by Bing (because ChatGPT search leans on it)
OpenAI names Bing among the search providers behind ChatGPT search, and independent analyses in 2026 found most of ChatGPT's cited pages also rank in Bing's top results. Copilot is Bing. If your Bing coverage is thin, so is your ChatGPT visibility. Here is the fix.
How to run a 30-minute AI searchability audit of your own website
You do not need an agency to find out why AI engines are not recommending you. This is the exact half-hour checklist we run before any deeper work: access, rendering, entity, answers, schema, freshness and footprint. Most sites fail two or three of the seven, and they are usually fixable in a week.
How to write an llms.txt file (and what it does and does not do)
Google says you can ignore llms.txt. Ahrefs found 97 percent of the files it studied were never requested. No major AI lab has committed to reading it. It is still a 30-minute job with no downside, as long as you know exactly what you are and are not getting. Here is how to do it properly.
Get in touch