Generative engine optimisation

GEO: being recommended by AI

Generative engine optimisation is about AI assistants such as ChatGPT, Claude and Perplexity being able to reach your site, understand it, and feel confident naming you when someone asks for a recommendation.

What it is

AI assistants write their own answers. When someone asks one for "a good bakery in Leeds" or "the best tool for X", it searches, reads a few pages and names a handful of options. GEO is about being one of them.

That starts with access. Each AI company runs its own crawlers, and your robots.txt file, firewall or CDN can let them in or shut them out, often without anyone noticing. After that it's about making the site easy for a model to read and describe accurately.

Why it matters

People ask assistants for recommendations

A growing share of people ask an AI assistant instead of searching. The assistant names only a few businesses, so being one of them matters more than ranking tenth in a list of links.

A blocked crawler means you're invisible

If robots.txt disallows a bot like OAI-SearchBot, Claude-SearchBot or PerplexityBot, that company's search can't read your pages. Some CDNs and firewalls also block AI crawlers by default or with a single setting, so a site can be shut out without anyone deciding to.

Many AI crawlers don't run JavaScript

If your main content only appears after JavaScript runs in the browser, crawlers that read the raw HTML see a mostly empty page. Server-rendered content is readable by every crawler.

Being named builds trust

A recommendation from an assistant works much like one from a friend. If it describes you accurately and cites your site, people arrive already expecting to find what they need.

What Crawlable checks

AI crawler access in robots.txt
How your robots.txt treats each AI search crawler, user-triggered fetcher and training crawler, from OAI-SearchBot and PerplexityBot to GPTBot, ClaudeBot and Google-Extended.
Firewall and CDN blocking
For the AI bots robots.txt allows, a test fetch of your homepage identifying as that bot, to see whether a firewall or CDN turns it away or serves a challenge page.
Server-rendered content
How much of the page's text is in the HTML the server sends, rather than added later by JavaScript.
Sitemap and llms.txt
Whether you have an XML sitemap (and list it in robots.txt), and whether you offer an llms.txt summary for AI tools.
Entity clarity
Whether your brand name is consistent, and whether Organization schema links your profiles elsewhere with sameAs.
A live Perplexity check
We ask Perplexity the kind of question your customers might ask and record whether it cites your site in its answer.

Common problems and how to fix them

AI search crawlers blocked in robots.txt

Allow the bots that surface sites in AI search. For example:
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /
Blocking training-only crawlers such as GPTBot or ClaudeBot is a legitimate choice and separate from search; keep the search bots allowed either way.

robots.txt allows a bot, but the firewall refuses it

Check your CDN or firewall's bot settings (for example a "block AI bots" switch or a bot-fight mode) and allow the AI search crawlers you want. Verified crawlers come from their operators' published IP ranges, so many firewalls can allow them specifically.

Main content only appears after JavaScript runs

Render the page's main text on the server, using server-side rendering or static generation in your framework, so it's in the HTML every crawler receives.

No sitemap, or one robots.txt doesn't mention

Publish /sitemap.xml listing your important pages and add a line to robots.txt: Sitemap: https://www.example.com/sitemap.xml.

Questions

Should I block GPTBot and ClaudeBot?

That's a business decision. GPTBot and ClaudeBot collect training data, and blocking them stops your content being used to train future models. It doesn't stop ChatGPT or Claude search from reading your site, which uses OAI-SearchBot and Claude-SearchBot. Most businesses that want to be recommended keep the search bots allowed.

What is llms.txt?

llms.txt is a proposed convention: a plain-text file at your site's root that summarises who you are and links your key pages for AI tools. Adoption is still limited, so it's a low-effort nice-to-have rather than a requirement.

Does robots.txt control ChatGPT-User and Perplexity-User?

Only loosely. These fetch a page when a person asks the assistant to. OpenAI and Perplexity say these user-triggered fetchers may not follow robots.txt, so blocking them there is a weak control.

See how your site does

Free, no email, and about a minute from typing your domain to reading the report.

Run your audit