What it is
AI assistants write their own answers. When someone asks one for "a good bakery in Leeds" or "the best tool for X", it searches, reads a few pages and names a handful of options. GEO is about being one of them.
That starts with access. Each AI company runs its own crawlers, and your robots.txt file, firewall or CDN can let them in or shut them out, often without anyone noticing. After that it's about making the site easy for a model to read and describe accurately.
Why it matters
People ask assistants for recommendations
A growing share of people ask an AI assistant instead of searching. The assistant names only a few businesses, so being one of them matters more than ranking tenth in a list of links.
A blocked crawler means you're invisible
If robots.txt disallows a bot like OAI-SearchBot, Claude-SearchBot or PerplexityBot, that company's search can't read your pages. Some CDNs and firewalls also block AI crawlers by default or with a single setting, so a site can be shut out without anyone deciding to.
Many AI crawlers don't run JavaScript
If your main content only appears after JavaScript runs in the browser, crawlers that read the raw HTML see a mostly empty page. Server-rendered content is readable by every crawler.
Being named builds trust
A recommendation from an assistant works much like one from a friend. If it describes you accurately and cites your site, people arrive already expecting to find what they need.
What Crawlable checks
- AI crawler access in robots.txt
- How your
robots.txttreats each AI search crawler, user-triggered fetcher and training crawler, fromOAI-SearchBotandPerplexityBottoGPTBot,ClaudeBotandGoogle-Extended. - Firewall and CDN blocking
- For the AI bots
robots.txtallows, a test fetch of your homepage identifying as that bot, to see whether a firewall or CDN turns it away or serves a challenge page. - Server-rendered content
- How much of the page's text is in the HTML the server sends, rather than added later by JavaScript.
- Sitemap and llms.txt
- Whether you have an XML sitemap (and list it in
robots.txt), and whether you offer anllms.txtsummary for AI tools. - Entity clarity
- Whether your brand name is consistent, and whether
Organizationschema links your profiles elsewhere withsameAs. - A live Perplexity check
- We ask Perplexity the kind of question your customers might ask and record whether it cites your site in its answer.
Common problems and how to fix them
AI search crawlers blocked in robots.txt
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
Blocking training-only crawlers such as GPTBot or ClaudeBot is a legitimate choice and separate from search; keep the search bots allowed either way.robots.txt allows a bot, but the firewall refuses it
Main content only appears after JavaScript runs
No sitemap, or one robots.txt doesn't mention
/sitemap.xml listing your important pages and add a line to robots.txt: Sitemap: https://www.example.com/sitemap.xml.Questions
Should I block GPTBot and ClaudeBot?
That's a business decision. GPTBot and ClaudeBot collect training data, and blocking them stops your content being used to train future models. It doesn't stop ChatGPT or Claude search from reading your site, which uses OAI-SearchBot and Claude-SearchBot. Most businesses that want to be recommended keep the search bots allowed.
What is llms.txt?
llms.txt is a proposed convention: a plain-text file at your site's root that summarises who you are and links your key pages for AI tools. Adoption is still limited, so it's a low-effort nice-to-have rather than a requirement.
Does robots.txt control ChatGPT-User and Perplexity-User?
Only loosely. These fetch a page when a person asks the assistant to. OpenAI and Perplexity say these user-triggered fetchers may not follow robots.txt, so blocking them there is a weak control.