On this page
← All articles
Guide·Sep 24, 2026·8 min read

How to Audit Website for AI Readiness in 2026: 12-Step Checklist

Hands writing in a notebook on cream linen beside the twelve-step AI readiness checklist

To audit a website for AI readiness, run twelve checks in order: crawler access, indexability, raw HTML versus rendered output, llms.txt, entity clarity, headings, answer extraction, schema, freshness, citation potential, authority, and a re-test. Ten of them take under an hour with free tools. This is how to audit website for AI readiness the way an engineer would, plus the part most checklists skip: a page that passes every check can still be described wrongly, or never recommended, by ChatGPT, Gemini, Claude and Perplexity.

Kairosy AI-Ready Audit page section on prioritized on-page fixes ranked by their impact on AI visibility

How to Audit Website for AI Readiness: The 12-Step Checklist

With only thirty minutes, do steps 1, 2, 3 and 8; they produce most hard failures.

  1. Crawler access: robots.txt allows every vendor's search fetcher.
  2. Indexability: indexed, snippet-eligible, not canonicalized elsewhere.
  3. Raw HTML versus rendering: answer text exists before JavaScript runs.
  4. llms.txt: present, spec-compliant, links resolve.
  5. Entity clarity: one name, one description, one Organization record.
  6. Headings: phrased as the questions buyers type.
  7. Answer extraction: a direct answer under each heading.
  8. Schema: valid JSON-LD, no required-field errors.
  9. Freshness: a truthful dateModified.
  10. Citation potential: statistics, sourced claims, quotable sentences.
  11. Authority: independent pages that corroborate yours.
  12. Re-test: same URL, same checks, after the fixes ship.

Can GPTBot, ClaudeBot and PerplexityBot Actually Reach Your Pages?

Steps 1 to 4 are the access layer. Nothing else matters if the fetchers that build AI answers cannot get a 200 response with your text in it.

Step 1: Check robots.txt for every AI user agent, not just Googlebot

Each vendor runs several bots with different jobs. Blocking a training crawler is a legitimate choice; blocking a search crawler by accident is the most expensive mistake in this audit. The table is compiled from the vendors' crawler documentation as of September 2026: OpenAI, Anthropic, Perplexity and Google. User-triggered fetchers (ChatGPT-User, Perplexity-User) may ignore robots.txt.

User agentOperatorJob, per the vendor's docsrobots.txt
GPTBotOpenAITraining foundation modelsHonored
OAI-SearchBotOpenAISurfaces sites in ChatGPT search; opted-out sites "will not be shown in ChatGPT search answers"Honored
ClaudeBotAnthropicCollects content that may contribute to trainingHonored
Claude-SearchBotAnthropicSearch; blocking it "may reduce your site's visibility and accuracy in user search results"Honored
PerplexityBotPerplexitySurfaces and links sites in Perplexity resultsHonored
GooglebotGoogleSearch index; AI Overviews and AI Mode draw supporting links from itHonored
Google-ExtendedGoogleA robots.txt token, not a crawler; controls Gemini training and grounding; no effect on Search inclusion or rankingToken only

Do not read robots.txt by eye. Paste the URL into the free AI crawl checker: it evaluates 15 user agents and reports the exact Allow or Disallow line behind each verdict, plus the meta robots tag and X-Robots-Tag header, since a permissive robots.txt is overruled by a noindex header.

Kairosy AI crawl checker results for figma.com: 4 of 15 AI bots blocked, with the robots.txt rule applied to GPTBot and ClaudeBot

Claude Help Center article 'Does Anthropic crawl data from the web, and how can site owners block the crawler?'

Step 2: Confirm the page is indexed and snippet-eligible

Google's AI features documentation says a page "must be indexed and eligible to be shown in Google Search with a snippet" to appear as a supporting link in AI Overviews or AI Mode, with "no additional technical requirements." Run the URL through the index checker, which queries Google, Bing, DuckDuckGo, Yahoo and Yandex in real time (3 URLs per run anonymously, 10 signed in). Unindexed usually means a stray noindex, a canonical pointing elsewhere, a robots block or thin duplicate content; also confirm nosnippet is not limiting what an engine may quote.

Kairosy free index checker verifying whether a URL is indexed on Google, Bing, DuckDuckGo, Yahoo and Yandex

Step 3: Compare the raw HTML with the rendered page

Google renders JavaScript on a queue; its docs say a page "may stay on this queue for a few seconds, but it can take longer than that." AI crawlers do not render at all. Vercel's December 2024 traffic analysis concluded that "none of the major AI crawlers currently render JavaScript"; the ChatGPT and Claude crawlers fetch JS files (11.50% and 23.84% of requests) but never execute them. The check: fetch the page with curl and search for the sentence you expect an engine to quote. If it is missing, the content is injected client-side and invisible to those crawlers; move it into the server response or prerender it.

Step 4: Publish an llms.txt file and validate it

llms.txt is a proposal by Jeremy Howard, first published September 3, 2024: a Markdown file at /llms.txt with a required H1, a blockquote summary and H2 sections of links. None of the four crawler documents in step 1 mention it, so treat it as a cheap hedge, not a lever. The llms.txt validator checks for a real 200 as text/plain, lints the H1, blockquote and section structure, samples links for dead URLs, probes for llms-full.txt and returns an A to F grade.

Kairosy llms.txt validator page that fetches a domain's llms.txt, lints it against the spec and grades it A to F

Does Your Page Say Who You Are and Answer the Question in the First Screen?

Steps 5 to 7 are the comprehension layer: which entity is this page about, and which passage answers the query?

Step 5: Entity clarity, one name and one description everywhere

Models resolve brands by matching name, domain and a consistent description across many sources; if the homepage, the footer and your G2 profile each describe the product differently, an engine may pull in a different company with a similar name. Google's guidance on Organization structured data says it "can help Google better understand your organization's administrative details and disambiguate your organization in search results." No properties are required; name, url, logo, description, foundingDate and a sameAs list of official profiles do most of the work.

Step 6: Headings that read like the questions people ask

Engines retrieve passages, and heading text is the strongest label a passage carries. Rewrite H2s and H3s as the question a buyer would type: "Pricing" becomes "How much does Acme cost in 2026?". One idea per heading, under about 300 words.

Step 7: Answer extraction, the 40-word test

The first paragraph under every heading should answer it in two or three sentences before any caveats; if it still makes sense without the heading, a model can quote it. Lists, tables and definitions extract better than narrative.

Which Schema, Freshness and Authority Signals Do AI Engines Actually Use?

Steps 8 to 11 are the trust layer. They separate a page an engine can read from a page it chooses to cite.

Step 8: Schema, validate what exists before adding more

Google says no special markup is required for AI Overviews, but schema labels entities and facts unambiguously, and broken schema signals low quality. Paste the URL into the schema markup checker: it lists every JSON-LD block, unwraps @graph, flags parse errors such as trailing commas, and returns field-level errors and warnings for Organization, Product, Article, FAQPage, HowTo and 16 more types. Fix errors first, then add Organization on the homepage and Article or Product where they apply; our FAQ schema guide covers why FAQPage still matters to answer engines.

Kairosy schema markup checker listing JSON-LD blocks on a page with field-level errors and warnings

Step 9: Freshness that is true, not just stamped

Recency is measured, and it decays fast. Gander's Q1 2026 study of 194,077 URLs retrieved by ChatGPT, Google AI Overviews, Gemini and Perplexity found that "Content has a 1-year half-life in AI search. Each year of age reduces a page's visibility by roughly 40-60%." Compare dateModified with the visible text and put the update in the body: a new statistic, a new year in the heading, a changelog line.

Step 10: Citation potential, give the model something to quote

Pages get cited when they hold a fact an engine wants to attribute: numbers with source and date, one-sentence definitions, named comparisons. Count the sentences a model could quote with a source attached; zero means the page reads as opinion to a retrieval system. The GEO paper (arXiv, November 2023) reported that page-level rewrites of this kind could "boost visibility by up to 40% in generative engine responses."

Step 11: Authority, what the rest of the web says

Crawlers read far more than they send back. Cloudflare's August 2025 crawl-to-refer analysis put Anthropic at 38,065.7 HTML pages crawled per referral in July 2025, OpenAI at 1,091.4 and Perplexity at 194.8, against 5.4 for Google. List the ten independent pages that mention your brand most and check whether they agree with your own description; the most cited AI sources shows which domains those tend to be.

How Do You Re-Test After Fixes Without Fooling Yourself?

Re-run the same tools on the same URLs after deployment and compare against the first run, not memory. Kairosy's AI-Ready Page Audit automates this layer: 33 checks across five dimensions (access, readability, structured data, citability, trust), stored per audit, with a "since your last audit" diff listing which checks flipped and which changed severity. Two rules: test the final URL after redirects, because a www host that 301s to the apex can pass on one and fail on the other; and re-test the answers, not only the page, because the same question to the same engine returns different text between runs.

Top of a Kairosy brand scan report for Supabase: 77/100 reputation, brand sentiment, recommendation rate, AI visibility and answer counts

Why Is AI Readiness Not the Same as AI Reputation Readiness?

Every step above asks whether an engine can fetch, parse and quote this page. None asks the question your CEO will: when a buyer asks ChatGPT what to use, does it name us, describe us correctly and recommend us? Crawlable is a property of a URL; recommended is a property of a brand, assembled from your pages plus every review, thread and comparison the engine has read. Most explanations of how to audit website for AI readiness stop at the URL.

The gap shows in the data. Kairosy's scan of Webflow, whose product is a website builder and whose robots.txt (checked September 2026) blocks no AI crawler on its marketing pages, reads "64/100, verdict At risk — present in 92% of answers, favourable in 56%, recommended in 50%" across 24 answers from four engines scanned 2026-08-05, with Gemini opening on "the most common complaints about Webflow center on its steep learning curve" (full report). Kairosy's scan of Netlify, whose robots.txt explicitly allows GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended and points to an llms.txt, shows the same pattern: 100% visibility, yet only 46% of answers actively recommend it and four of 24 steer buyers to Vercel, Cloudflare Pages or Render first. Both sites are open to every AI crawler. Neither is reliably recommended.

A complete audit therefore has two halves: on-site and deterministic, then off-site and probabilistic, which needs a different instrument: the answers themselves, scored per engine, with the URLs each cited so a negative sentence can be traced to its source page. Kairosy's brand monitoring runs that half every Monday for each brand and market you track, and its Fix Plan orders the actions by impact.

Kairosy Fix Strategy suggestion 3 of 8 for Figma: no JSON-LD structured data on the homepage, with the Organization and Product markup fix

Technical Readiness vs Brand Evidence Readiness: What Does Each Audit Check?

Use this table to scope the work and to explain why a green technical audit did not move the answers.

DimensionTechnical readinessBrand evidence readiness
Question answeredCan an engine fetch, parse and quote this page?Asked about the category, does the engine name, describe and recommend this brand?
UnitOne URLOne brand in one market
Inputsrobots.txt, HTTP status, raw HTML, JSON-LD, llms.txt, dateModifiedAnswers from ChatGPT, Gemini, Claude and Perplexity, plus the URLs each cites
Pass/fail signal33 checks in five dimensions: access, readability, structured data, citability, trustScore 0 to 100 (average of per-engine scores); each answer positive, neutral or negative
Free way to checkAI crawl checker, index checker, schema checker, llms.txt validatorAsk the four engines the same buyer questions by hand, or run the 7-day trial
Typical failureOAI-SearchBot disallowed; answer text rendered client-side; JSON-LD parse errorPresent in 92% of answers, recommended in 50% (Webflow, scanned 2026-08-05)

Do the technical half first; it is faster. Then measure the brand half before rewriting content, so you know which questions, engines and sources cost you the recommendation. The GEO implementation guide for technical teams goes deeper on access and rendering; the GEO audit checklist covers the content side.

How to Audit Website for AI Readiness FAQs

How to audit website for AEO readiness (AI engine optimization)?

Use the same twelve steps: AEO readiness and AI readiness describe the same property, that an answer engine can fetch the page, identify the entity and question it covers, and lift a sourced answer from it.

Does blocking GPTBot remove my site from ChatGPT answers?

Not by itself. OpenAI documents GPTBot as the training crawler and OAI-SearchBot as the one that surfaces sites in ChatGPT search, and says sites opted out of OAI-SearchBot will not be shown there. Block GPTBot if you wish; keep OAI-SearchBot allowed if you want to be cited.

Do I need llms.txt to be AI-ready?

No. It is a 2024 proposal, and the crawler documentation from OpenAI, Anthropic, Perplexity and Google does not reference it as of September 2026. Publish and validate one because it is cheap, but fix rendering and schema first.

No. Readiness makes pages eligible; recommendation depends on what the engine has read about you elsewhere. Webflow and Netlify both allow every AI crawler and both score "At risk" in Kairosy's scans. Measure the answers separately; the 7-day free trial (no credit card) shows the answers from four AI platforms of your choice, for your own brand.

See what AI says about your brand

Run a free scan across ChatGPT, Gemini, Claude & Perplexity in about 30 seconds.

Run my free scan