Chapter 3 of 6 · updated 2026-10-02
Technical foundations: can engines read you?
Crawler access, JavaScript rendering, structured data, llms.txt and the other checks that decide whether an engine can read a page at all.
Let the right crawlers in
Each vendor runs separate crawlers for training, search indexing and user-triggered fetches. Blocking the training bot (GPTBot, ClaudeBot) does not affect citations; blocking the search bot (OAI-SearchBot, Claude-SearchBot, PerplexityBot) or the fetcher (ChatGPT-User, Claude-User) removes you from answers. Google's AI surfaces use the normal Googlebot crawl; Google-Extended only governs Gemini training and grounding.
Check the live robots.txt against the full crawler list, then check your CDN and firewall rules, which can block a crawler the file allows.
Put the content in the HTML
Most AI crawlers do not execute JavaScript. If your product pages, comparisons or FAQs render client-side, engines see an empty shell. Server-render or prerender anything you want cited, and verify by reading the raw HTML rather than the browser view.
Structured data that matches the page
Organization markup with consistent name, url, logo and sameAs links on the home and About pages; Article markup with author and dates on content; FAQPage wherever questions are answered visibly; SoftwareApplication or Product on product pages. Markup must parse, declare the schema.org context, and match what visitors see.
The small files
A sitemap that lists every page you want found, a robots.txt that declares it, and an llms.txt that describes the site and its key pages in plain markdown. None of these is a ranking lever; together they remove reasons for an engine to misunderstand you.
- robots.txt: allow search and fetch crawlers; declare the sitemap.
- sitemap.xml: every indexable page, with real lastmod dates or none.
- llms.txt: H1, one-paragraph summary, sections of links with notes.
- An AI instructions page: your key facts, dated, in one place.
Key takeaways
- Allow search and fetch crawlers; training-bot blocks are a separate decision.
- Content must be in the server HTML; most AI crawlers do not run JavaScript.
- Structured data must parse, use schema.org, and match visible content.
- Publish sitemap, robots, llms.txt and a dated facts page.
Sources and further reading
- Google Search Central: Introduction to robots.txt
- Google Search Central: Google's common crawlers: including Google-Extended
- OpenAI: Overview of OpenAI crawlers: GPTBot, OAI-SearchBot and ChatGPT-User
- Perplexity: PerplexityBot and Perplexity-User
- Anthropic: Does Anthropic crawl data from the web?: ClaudeBot, Claude-User and Claude-SearchBot
- Google Search Central: JavaScript SEO basics
- Google Search Central: Introduction to structured data
- llms.txt proposal
- RFC 9309: Robots Exclusion Protocol
Get early access to LeapScope
Join the waitlist and we will invite you in batches, with early-access pricing for everyone on the list.