How to Audit Website for AI Readiness in 2026: 12-Step Checklist

To audit a website for AI readiness, run twelve checks in order: crawler access, indexability, raw HTML versus rendered output, llms.txt, entity clarity, headings, answer extraction, schema, freshness, citation potential, authority, and a re-test. Ten of them take under an hour with free tools. This is how to audit website for AI readiness the way an engineer would, plus the part most checklists skip: a page that passes every check can still be described wrongly, or never recommended, by ChatGPT, Gemini, Claude and Perplexity.

How to Audit Website for AI Readiness: The 12-Step Checklist
With only thirty minutes, do steps 1, 2, 3 and 8; they produce most hard failures.
- Crawler access: robots.txt allows every vendor's search fetcher.
- Indexability: indexed, snippet-eligible, not canonicalized elsewhere.
- Raw HTML versus rendering: answer text exists before JavaScript runs.
- llms.txt: present, spec-compliant, links resolve.
- Entity clarity: one name, one description, one Organization record.
- Headings: phrased as the questions buyers type.
- Answer extraction: a direct answer under each heading.
- Schema: valid JSON-LD, no required-field errors.
- Freshness: a truthful dateModified.
- Citation potential: statistics, sourced claims, quotable sentences.
- Authority: independent pages that corroborate yours.
- Re-test: same URL, same checks, after the fixes ship.
Can GPTBot, ClaudeBot and PerplexityBot Actually Reach Your Pages?
Steps 1 to 4 are the access layer. Nothing else matters if the fetchers that build AI answers cannot get a 200 response with your text in it.
Step 1: Check robots.txt for every AI user agent, not just Googlebot
Each vendor runs several bots with different jobs. Blocking a training crawler is a legitimate choice; blocking a search crawler by accident is the most expensive mistake in this audit. The table is compiled from the vendors' crawler documentation as of September 2026: OpenAI, Anthropic, Perplexity and Google. User-triggered fetchers (ChatGPT-User, Perplexity-User) may ignore robots.txt.
| User agent | Operator | Job, per the vendor's docs | robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Training foundation models | Honored |
| OAI-SearchBot | OpenAI | Surfaces sites in ChatGPT search; opted-out sites "will not be shown in ChatGPT search answers" | Honored |
| ClaudeBot | Anthropic | Collects content that may contribute to training | Honored |
| Claude-SearchBot | Anthropic | Search; blocking it "may reduce your site's visibility and accuracy in user search results" | Honored |
| PerplexityBot | Perplexity | Surfaces and links sites in Perplexity results | Honored |
| Googlebot | Search index; AI Overviews and AI Mode draw supporting links from it | Honored | |
| Google-Extended | A robots.txt token, not a crawler; controls Gemini training and grounding; no effect on Search inclusion or ranking | Token only |
Do not read robots.txt by eye. Paste the URL into the free AI crawl checker: it evaluates 15 user agents and reports the exact Allow or Disallow line behind each verdict, plus the meta robots tag and X-Robots-Tag header, since a permissive robots.txt is overruled by a noindex header.


Step 2: Confirm the page is indexed and snippet-eligible
Google's AI features documentation says a page "must be indexed and eligible to be shown in Google Search with a snippet" to appear as a supporting link in AI Overviews or AI Mode, with "no additional technical requirements." Run the URL through the index checker, which queries Google, Bing, DuckDuckGo, Yahoo and Yandex in real time (3 URLs per run anonymously, 10 signed in). Unindexed usually means a stray noindex, a canonical pointing elsewhere, a robots block or thin duplicate content; also confirm nosnippet is not limiting what an engine may quote.

Step 3: Compare the raw HTML with the rendered page
Google renders JavaScript on a queue; its docs say a page "may stay on this queue for a few seconds, but it can take longer than that." AI crawlers do not render at all. Vercel's December 2024 traffic analysis concluded that "none of the major AI crawlers currently render JavaScript"; the ChatGPT and Claude crawlers fetch JS files (11.50% and 23.84% of requests) but never execute them. The check: fetch the page with curl and search for the sentence you expect an engine to quote. If it is missing, the content is injected client-side and invisible to those crawlers; move it into the server response or prerender it.
Step 4: Publish an llms.txt file and validate it
llms.txt is a proposal by Jeremy Howard, first published September 3, 2024: a Markdown file at /llms.txt with a required H1, a blockquote summary and H2 sections of links. None of the four crawler documents in step 1 mention it, so treat it as a cheap hedge, not a lever. The llms.txt validator checks for a real 200 as text/plain, lints the H1, blockquote and section structure, samples links for dead URLs, probes for llms-full.txt and returns an A to F grade.

Does Your Page Say Who You Are and Answer the Question in the First Screen?
Steps 5 to 7 are the comprehension layer: which entity is this page about, and which passage answers the query?
Step 5: Entity clarity, one name and one description everywhere
Models resolve brands by matching name, domain and a consistent description across many sources; if the homepage, the footer and your G2 profile each describe the product differently, an engine may pull in a different company with a similar name. Google's guidance on Organization structured data says it "can help Google better understand your organization's administrative details and disambiguate your organization in search results." No properties are required; name, url, logo, description, foundingDate and a sameAs list of official profiles do most of the work.
Step 6: Headings that read like the questions people ask
Engines retrieve passages, and heading text is the strongest label a passage carries. Rewrite H2s and H3s as the question a buyer would type: "Pricing" becomes "How much does Acme cost in 2026?". One idea per heading, under about 300 words.
Step 7: Answer extraction, the 40-word test
The first paragraph under every heading should answer it in two or three sentences before any caveats; if it still makes sense without the heading, a model can quote it. Lists, tables and definitions extract better than narrative.
Which Schema, Freshness and Authority Signals Do AI Engines Actually Use?
Steps 8 to 11 are the trust layer. They separate a page an engine can read from a page it chooses to cite.
Step 8: Schema, validate what exists before adding more
Google says no special markup is required for AI Overviews, but schema labels entities and facts unambiguously, and broken schema signals low quality. Paste the URL into the schema markup checker: it lists every JSON-LD block, unwraps @graph, flags parse errors such as trailing commas, and returns field-level errors and warnings for Organization, Product, Article, FAQPage, HowTo and 16 more types. Fix errors first, then add Organization on the homepage and Article or Product where they apply; our FAQ schema guide covers why FAQPage still matters to answer engines.

Step 9: Freshness that is true, not just stamped
Recency is measured, and it decays fast. Gander's Q1 2026 study of 194,077 URLs retrieved by ChatGPT, Google AI Overviews, Gemini and Perplexity found that "Content has a 1-year half-life in AI search. Each year of age reduces a page's visibility by roughly 40-60%." Compare dateModified with the visible text and put the update in the body: a new statistic, a new year in the heading, a changelog line.
Step 10: Citation potential, give the model something to quote
Pages get cited when they hold a fact an engine wants to attribute: numbers with source and date, one-sentence definitions, named comparisons. Count the sentences a model could quote with a source attached; zero means the page reads as opinion to a retrieval system. The GEO paper (arXiv, November 2023) reported that page-level rewrites of this kind could "boost visibility by up to 40% in generative engine responses."
Step 11: Authority, what the rest of the web says
Crawlers read far more than they send back. Cloudflare's August 2025 crawl-to-refer analysis put Anthropic at 38,065.7 HTML pages crawled per referral in July 2025, OpenAI at 1,091.4 and Perplexity at 194.8, against 5.4 for Google. List the ten independent pages that mention your brand most and check whether they agree with your own description; the most cited AI sources shows which domains those tend to be.
How Do You Re-Test After Fixes Without Fooling Yourself?
Re-run the same tools on the same URLs after deployment and compare against the first run, not memory. Kairosy's AI-Ready Page Audit automates this layer: 33 checks across five dimensions (access, readability, structured data, citability, trust), stored per audit, with a "since your last audit" diff listing which checks flipped and which changed severity. Two rules: test the final URL after redirects, because a www host that 301s to the apex can pass on one and fail on the other; and re-test the answers, not only the page, because the same question to the same engine returns different text between runs.

Why Is AI Readiness Not the Same as AI Reputation Readiness?
Every step above asks whether an engine can fetch, parse and quote this page. None asks the question your CEO will: when a buyer asks ChatGPT what to use, does it name us, describe us correctly and recommend us? Crawlable is a property of a URL; recommended is a property of a brand, assembled from your pages plus every review, thread and comparison the engine has read. Most explanations of how to audit website for AI readiness stop at the URL.
The gap shows in the data. Kairosy's scan of Webflow, whose product is a website builder and whose robots.txt (checked September 2026) blocks no AI crawler on its marketing pages, reads "64/100, verdict At risk — present in 92% of answers, favourable in 56%, recommended in 50%" across 24 answers from four engines scanned 2026-08-05, with Gemini opening on "the most common complaints about Webflow center on its steep learning curve" (full report). Kairosy's scan of Netlify, whose robots.txt explicitly allows GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended and points to an llms.txt, shows the same pattern: 100% visibility, yet only 46% of answers actively recommend it and four of 24 steer buyers to Vercel, Cloudflare Pages or Render first. Both sites are open to every AI crawler. Neither is reliably recommended.
A complete audit therefore has two halves: on-site and deterministic, then off-site and probabilistic, which needs a different instrument: the answers themselves, scored per engine, with the URLs each cited so a negative sentence can be traced to its source page. Kairosy's brand monitoring runs that half every Monday for each brand and market you track, and its Fix Plan orders the actions by impact.

Technical Readiness vs Brand Evidence Readiness: What Does Each Audit Check?
Use this table to scope the work and to explain why a green technical audit did not move the answers.
| Dimension | Technical readiness | Brand evidence readiness |
|---|---|---|
| Question answered | Can an engine fetch, parse and quote this page? | Asked about the category, does the engine name, describe and recommend this brand? |
| Unit | One URL | One brand in one market |
| Inputs | robots.txt, HTTP status, raw HTML, JSON-LD, llms.txt, dateModified | Answers from ChatGPT, Gemini, Claude and Perplexity, plus the URLs each cites |
| Pass/fail signal | 33 checks in five dimensions: access, readability, structured data, citability, trust | Score 0 to 100 (average of per-engine scores); each answer positive, neutral or negative |
| Free way to check | AI crawl checker, index checker, schema checker, llms.txt validator | Ask the four engines the same buyer questions by hand, or run the 7-day trial |
| Typical failure | OAI-SearchBot disallowed; answer text rendered client-side; JSON-LD parse error | Present in 92% of answers, recommended in 50% (Webflow, scanned 2026-08-05) |
Do the technical half first; it is faster. Then measure the brand half before rewriting content, so you know which questions, engines and sources cost you the recommendation. The GEO implementation guide for technical teams goes deeper on access and rendering; the GEO audit checklist covers the content side.
How to Audit Website for AI Readiness FAQs
How to audit website for AEO readiness (AI engine optimization)?
Use the same twelve steps: AEO readiness and AI readiness describe the same property, that an answer engine can fetch the page, identify the entity and question it covers, and lift a sourced answer from it.
Does blocking GPTBot remove my site from ChatGPT answers?
Not by itself. OpenAI documents GPTBot as the training crawler and OAI-SearchBot as the one that surfaces sites in ChatGPT search, and says sites opted out of OAI-SearchBot will not be shown there. Block GPTBot if you wish; keep OAI-SearchBot allowed if you want to be cited.
Do I need llms.txt to be AI-ready?
No. It is a 2024 proposal, and the crawler documentation from OpenAI, Anthropic, Perplexity and Google does not reference it as of September 2026. Publish and validate one because it is cheap, but fix rendering and schema first.
Is an AI-ready website enough to get recommended by ChatGPT or Perplexity?
No. Readiness makes pages eligible; recommendation depends on what the engine has read about you elsewhere. Webflow and Netlify both allow every AI crawler and both score "At risk" in Kairosy's scans. Measure the answers separately; the 7-day free trial (no credit card) shows the answers from four AI platforms of your choice, for your own brand.
See what AI says about your brand
Run a free scan across ChatGPT, Gemini, Claude & Perplexity in about 30 seconds.
Run my free scan

