Chapter 3 of 6 · updated 2026-10-02

Technical foundations: can engines read you?

Crawler access, JavaScript rendering, structured data, llms.txt and the other checks that decide whether an engine can read a page at all.

Let the right crawlers in

Each vendor runs separate crawlers for training, search indexing and user-triggered fetches. Blocking the training bot (GPTBot, ClaudeBot) does not affect citations; blocking the search bot (OAI-SearchBot, Claude-SearchBot, PerplexityBot) or the fetcher (ChatGPT-User, Claude-User) removes you from answers. Google's AI surfaces use the normal Googlebot crawl; Google-Extended only governs Gemini training and grounding.

Check the live robots.txt against the full crawler list, then check your CDN and firewall rules, which can block a crawler the file allows.

Put the content in the HTML

Most AI crawlers do not execute JavaScript. If your product pages, comparisons or FAQs render client-side, engines see an empty shell. Server-render or prerender anything you want cited, and verify by reading the raw HTML rather than the browser view.

Structured data that matches the page

Organization markup with consistent name, url, logo and sameAs links on the home and About pages; Article markup with author and dates on content; FAQPage wherever questions are answered visibly; SoftwareApplication or Product on product pages. Markup must parse, declare the schema.org context, and match what visitors see.

The small files

A sitemap that lists every page you want found, a robots.txt that declares it, and an llms.txt that describes the site and its key pages in plain markdown. None of these is a ranking lever; together they remove reasons for an engine to misunderstand you.

  • robots.txt: allow search and fetch crawlers; declare the sitemap.
  • sitemap.xml: every indexable page, with real lastmod dates or none.
  • llms.txt: H1, one-paragraph summary, sections of links with notes.
  • An AI instructions page: your key facts, dated, in one place.

Key takeaways

  • Allow search and fetch crawlers; training-bot blocks are a separate decision.
  • Content must be in the server HTML; most AI crawlers do not run JavaScript.
  • Structured data must parse, use schema.org, and match visible content.
  • Publish sitemap, robots, llms.txt and a dated facts page.

Sources and further reading

Get early access to LeapScope

Join the waitlist and we will invite you in batches, with early-access pricing for everyone on the list.