---
title: "Technical foundations: can engines read you?"
url: https://leap-scope.com/guides/aeo/technical-foundations
description: "Crawler access, JavaScript rendering, structured data, llms.txt and the other checks that decide whether an engine can read a page at all."
---

Chapter 3 of 6 · updated 2026-10-02

# Technical foundations: can engines read you?

Crawler access, JavaScript rendering, structured data, llms.txt and the other checks that decide whether an engine can read a page at all.

## Let the right crawlers in

Each vendor runs separate crawlers for training, search indexing and user-triggered fetches. Blocking the training bot (GPTBot, ClaudeBot) does not affect citations; blocking the search bot (OAI-SearchBot, Claude-SearchBot, PerplexityBot) or the fetcher (ChatGPT-User, Claude-User) removes you from answers. Google's AI surfaces use the normal Googlebot crawl; Google-Extended only governs Gemini training and grounding.

Check the live robots.txt against the full crawler list, then check your CDN and firewall rules, which can block a crawler the file allows.

## Put the content in the HTML

Most AI crawlers do not execute JavaScript. If your product pages, comparisons or FAQs render client-side, engines see an empty shell. Server-render or prerender anything you want cited, and verify by reading the raw HTML rather than the browser view.

## Structured data that matches the page

Organization markup with consistent name, url, logo and sameAs links on the home and About pages; Article markup with author and dates on content; FAQPage wherever questions are answered visibly; SoftwareApplication or Product on product pages. Markup must parse, declare the schema.org context, and match what visitors see.

## The small files

A sitemap that lists every page you want found, a robots.txt that declares it, and an llms.txt that describes the site and its key pages in plain markdown. None of these is a ranking lever; together they remove reasons for an engine to misunderstand you.

- robots.txt: allow search and fetch crawlers; declare the sitemap.
- sitemap.xml: every indexable page, with real lastmod dates or none.
- llms.txt: H1, one-paragraph summary, sections of links with notes.
- An AI instructions page: your key facts, dated, in one place.

## Key takeaways

- Allow search and fetch crawlers; training-bot blocks are a separate decision.
- Content must be in the server HTML; most AI crawlers do not run JavaScript.
- Structured data must parse, use schema.org, and match visible content.
- Publish sitemap, robots, llms.txt and a dated facts page.

## Sources and further reading

- [Google Search Central: Introduction to robots.txt](https://developers.google.com/search/docs/crawling-indexing/robots/intro)
- [Google Search Central: Google's common crawlers](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers): including Google-Extended
- [OpenAI: Overview of OpenAI crawlers](https://platform.openai.com/docs/bots): GPTBot, OAI-SearchBot and ChatGPT-User
- [Perplexity: PerplexityBot and Perplexity-User](https://docs.perplexity.ai/guides/bots)
- [Anthropic: Does Anthropic crawl data from the web?](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler): ClaudeBot, Claude-User and Claude-SearchBot
- [Google Search Central: JavaScript SEO basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics)
- [Google Search Central: Introduction to structured data](https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data)
- [llms.txt proposal](https://llmstxt.org/)
- [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html)

## Get early access to LeapScope

Join the waitlist and we will invite you in batches, with early-access pricing for everyone on the list.
