robots.txt, llms.txt & AI crawler access

Blocking AI crawlers guarantees invisibility. Opening the door without structured entity data guarantees confusion. Do both layers.

All articles6 min readMay 22, 2026

Key takeaway

Blocking AI crawlers guarantees invisibility. Opening the door without structured entity data guarantees confusion. Do both layers.

robots.txt GPTBotllms.txtallow AI crawlersgenerative engine optimization

Baseline: allow the bots you want answering for you

newsusa.ai explicitly allows GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, anthropic-ai, and major search bots. Mirror that pattern on your domain before investing in content, otherwise models literally cannot fetch your fixes.

llms.txt as an entity card

Publish llms.txt at your site root with canonical URLs, product names, founder/leadership facts, and a page index with one-line citations. Models and retrieval pipelines use it as a low-token map of what to trust on your site, complementing sitemap.xml, not replacing editorial proof off-domain.

JSON-LD and FAQ parity

Match visible FAQ copy with FAQPage schema; add Organization, WebSite, and Service nodes. Per-route WebPage and Article graphs help SPAs like NewsUSA.ai expose page-level context on navigation. Technical access without off-site press corroboration still leaves competitive prompts weak, combine GEO with placement strategy.

Frequently asked questions

What is llms.txt and do I actually need one?

llms.txt is a plain-text file at your site root, proposed by the llms.txt standard, that gives AI models and retrieval pipelines a low-token map of your canonical URLs, product names, and key facts. It is not mandatory, but it reduces the guesswork models otherwise do when they crawl a large or JavaScript-heavy site.

Is it safe to allow GPTBot and ClaudeBot in robots.txt?

Yes, for any site that wants AI assistants to be able to cite it. Blocking these crawlers in robots.txt does not protect content, it simply makes the site invisible to the assistants your buyers already use, while doing nothing to stop other forms of scraping.

Sources: robots.txt (Wikipedia)

Written by

← Back to all guides

Audit your brand in AI answers

See citation authority, accuracy, and competitor leakage on the prompts that matter.

No credit card · Personalized report on your strategy callWith Todd Simon · No credit card