# LLM Technical SEO Checklist: Seven Stages, P0 to P2 Fixes

> A technical SEO checklist for AI search: seven stages in order, every check with a pass rule and a P0 to P2 priority, backed by sources and run on a live site.

Source: https://khadijazaman.com/blog/llm-technical-seo-checklist/  ·  Last modified: 2026-10-03

[Home](/) / [Blog](/blog/) / The LLM Technical SEO Checklist: Seven Stages From Crawler Access to Citation

[Optimization](/blog/category/optimization/ "More posts on Optimization")2 Oct 2026 · 18 min read · by Khadija Zaman

# The LLM Technical SEO Checklist: Seven Stages From Crawler Access to Citation

## What is in an LLM technical SEO checklist?

Khadija Zaman's LLM technical SEO checklist runs seven stages in order: prerequisites, measurement, crawler reach, site structure, trust and entity signals, AI retrieval and citation, and a publish gate for every new URL. Each check has a pass rule and a priority from P0, which blocks indexing or AI access, to P2, which is hygiene.

The first thing this checklist caught was on my own site. Both blog category pages showed a breadcrumb trail on screen, and neither had the BreadcrumbList markup that should sit behind it. I'd built those templates myself and never noticed, because nothing looked broken.

That's the case for running a checklist in a fixed order instead of from memory. AI search adds a few new failure points to technical SEO: crawlers that don't run JavaScript, separate bots for training and for search, and firewalls that block a bot robots.txt allows. Most of them are invisible from the browser.

This is the checklist I use. It has seven stages and 75 checks. Each check has a pass rule, a way to verify it, and a priority: **P0** blocks indexing or AI access, **P1** weakens signals, **P2** is hygiene. Every stage below also carries the documentation behind its rules and an example from this site, including the rows it fails.

## How to work through it

The stages run in order because each one depends on the one before it. Fixing schema on a page an AI crawler can't reach changes nothing, which is the same reason Eligibility is the only hard gate in my [R2A Content Framework](/frameworks/r2a/).

1.  Work the stages in order, and clear every P0 in a stage before you move to the next.
2.  Mark each row Pass, Fail or N/A as you go. N/A is a real answer: a personal site has no SoftwareApplication schema to check.
3.  Run the International module only if the site serves more than one language or region.
4.  Run two loops. Stages 0 to 5 are the site audit: run them at launch, then once a quarter. Stage 6 is the publish gate: run it on every new or changed URL before it goes live.

![Sketch-note staircase of six stages: 0 Prerequisites, 1 Measure, 2 Reach, 3 Structure, 4 Trust and Entity, 5 Retrieve and Cite, each with its question. A red note says clear every P0 before moving to the next stage. A dashed loop returns from stage 5 to stage 0, labelled site audit at launch then every quarter. Below sit a dashed International SEO module box, a P0 P1 P2 legend and a shaded Stage 6 Publish gate box for every new or changed URL](/static/uploads/llm-checklist-fig1-stages.png)
*Figure 1. Stages 0 to 5 are the quarterly site audit; Stage 6 runs on every URL you publish.*

| Stage | The question it answers | What changes for AI search |
| --- | --- | --- |
| 0. Prerequisites | Do I have the tools? | Nothing yet. Without Search Console and a crawler you can't see the rest. |
| 1. Measure | Can I see what search engines see? | Bing matters as much as Google: ChatGPT's web browsing launched on Bing in 2023. |
| 2. Reach | Can crawlers fetch, render and index the right pages? | Many AI crawlers read the raw HTML only, so anything JavaScript adds may never be seen. |
| 3. Structure | Does authority reach the pages that matter, and is each page clear? | Clear headings and one canonical URL per page make passages easier to lift. |
| 4. Trust & Entity | Is a real, machine-readable author behind it? | An engine deciding whose words to repeat looks for a named, consistent identity. |
| 5. Retrieve & Cite | Will AI engines read and cite it? | The right bots allowed, and answers phrased the way buyers ask. |
| 6. Publish gate | Is this new or changed URL ready to go live? | Each page enters the index clean on day one. |
| International | Does each visitor get the right language version? | Conditional: multi-region sites only. |

## Stage 0. Prerequisites

Two tools, both P0, because every later stage reads from them.

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| Search Console access | You can open the property for the site you're auditing | Log in and confirm the property shows data | P0 |
| A crawler | A full crawl of the site completes | Screaming Frog's free tier handles up to 500 URLs; start with the issues it ranks highest | P0 |

**On this site:** 28 pages is small enough that I ran Stages 2 to 6 as one script against the built HTML instead of a crawler. Above a few hundred URLs, use the crawler.

## Stage 1. Measure: can you see what search engines see?

This stage sets the baseline that later results get compared against. On the AI side I add one thing most technical checklists leave out: an LLM referral segment in analytics, which is how I [measured the Wellows growth](/work/wellows-ai-visibility/) in the first place.

**Evidence.** Google treats a sitemap as a weak canonical signal, behind redirects and rel="canonical" ([Google Search Central](https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls)), which is why a sitemap that lists redirecting or broken URLs muddies signals instead of just wasting space.

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| Search Console verified | A Domain property is verified and the Performance report shows data | Add the Domain property, verify with a DNS TXT record, wait a day or two for data | P0 |
| Analytics connected | The tag fires in real time and is linked to Search Console; key events are defined once, site-wide | Install through GTM or the page head, then link the two under Associations | P1 |
| Bing Webmaster Tools verified | Bing shows crawl and index data | Import the site from Search Console at bing.com/webmasters | P0 |
| Sitemap submitted and clean | It lists only indexable URLs that return 200, and few submitted URLs go unindexed | Submit it under Indexing, Sitemaps, then compare submitted with indexed | P0 |
| robots.txt leaves key paths open | No Disallow on the blog, product or other revenue directories | Open yoursite.com/robots.txt and read every Disallow line | P0 |
| One host | www and non-www resolve to one version through a site-wide 301 | Request both and check where each lands | P1 |
| HTTPS enforced | http:// requests 301 to https:// at the server | Request the http:// version of a deep URL, not only the homepage | P0 |
| No mixed content | No http:// assets load on https:// pages | Look for mixed-content warnings in the browser console | P2 |
| No manual actions | Search Console reports no issues | Security & Manual Actions in Search Console | P0 |
| Sections in subfolders | The blog and docs live at site.com/blog, not blog.site.com | Keep subdomains for separate products, regions or tech stacks | P2 |

**On this site:** the sitemap lists 28 URLs and all 28 resolve, and each one's lastmod comes from the git commit that last touched its source file, so a rebuild doesn't make every page look new. The analytics row is N/A here: this site runs cookieless Cloudflare Web Analytics, not GA4, and the LLM referral segment lives on Wellows.

## Stage 2. Reach: can crawlers fetch, render and index the page?

This stage holds most of the P0s. If the copy only appears once JavaScript runs, many AI crawlers never see it, which is why the [Citeability Checker](/tools/citeability-checker/) reads what's in the HTML rather than what a browser paints.

**Evidence.** Google renders JavaScript in a separate step after crawling ([JavaScript SEO basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics)); crawlers that skip that step only get the HTML. A noindex rule only works if the crawler can fetch the page, so a URL blocked in robots.txt can still appear in results ([block indexing](https://developers.google.com/search/docs/crawling-indexing/block-indexing)). And Google only reliably follows links that are `<a>` elements with an href, which is why button-driven pagination fails ([pagination](https://developers.google.com/search/docs/specialty/ecommerce/pagination-and-incremental-page-loading)).

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| Content without JavaScript | The key copy is in the raw HTML | View Source and search for a sentence, or disable JavaScript in DevTools; compare with URL Inspection's crawled page | P0 |
| No noindex on key pages | Neither the meta robots tag nor the X-Robots-Tag header says noindex | Filter the crawl by directives, and check response headers separately | P0 |
| Canonicals correct | Key pages point to themselves, and nothing points to the homepage by mistake | Crawl report of canonical targets | P0 |
| FAQ answers in the source | Answers sit in the HTML, aren't loaded on click, and aren't aria-hidden | View Source and search for one answer | P0 |
| Sitemap only lists 200s | Zero redirecting, broken or noindexed URLs in it | Crawl the sitemap in list mode | P1 |
| No redirect chains | Every redirect lands in one hop | Redirect chain report | P1 |
| Linked dead URLs redirected | Old URLs with backlinks 301 to the closest live page, not the homepage | Broken-backlink report, sorted by referring domains | P1 |
| 404s with value fixed | 404s that still get traffic or links are redirected | Crawl 4xx report, cross-checked with Search Console's Pages report | P1 |
| No soft 404s | Empty pages return a real 404, or carry real content | Search Console Pages report, Soft 404 (usually empty categories and search results) | P1 |
| Site search kept out | /?s= pages are noindexed, disallowed and absent from the sitemap | Check the template, robots.txt and the sitemap | P2 |
| Thin URLs kept out | Tag archives, filters and parameter URLs are noindexed | List the thin URL patterns in the crawl | P2 |
| Pagination uses links | Page 2 is an <a href>, not a button that runs a script | View Source on a paginated list | P1 |
| Paginated pages self-canonical | Page 2's canonical is page 2, not page 1 | Crawl report of canonical targets | P1 |

For the FAQ row, this is the difference in the markup:

```html
<!-- Fails: the answer isn't in the HTML until someone clicks -->
<button data-faq="12">Do you work with SaaS?</button>
<div class="answer"></div>

<!-- Passes: collapsed for people, present for crawlers -->
<details>
  <summary>Do you work with SaaS?</summary>
  <p>Yes. Most of my in-house work has been B2B SaaS.</p>
</details>
```

**On this site:** zero pages carry noindex, all 28 canonicals point to themselves, and the FAQ answers on 7 pages are in the raw HTML. The six-stage method on the [work page](/work/) uses `<details>`, so its text is in the source even while it's collapsed.

## Stage 3. Structure: does authority reach the pages that matter?

Two halves: how pages link to each other, and whether each page is clear on its own. The linking half is where [topic cluster architecture](/blog/topic-cluster-architecture/) does its work, and descriptive anchors are part of what makes a [central entity](/blog/central-entity/) readable across a whole site.

**Evidence.** Google builds title links from the title element, the main heading, og:title and anchor text pointing at the page, among other sources ([title links](https://developers.google.com/search/docs/appearance/title-link)), and it sets no character limit; long titles are truncated by width on the results page. Breadcrumb markup needs at least two ListItems with a name, an item URL and a position ([breadcrumb structured data](https://developers.google.com/search/docs/appearance/structured-data/breadcrumb)). Core Web Vitals pass when 75% of page views meet LCP 2.5 s, INP 200 ms and CLS 0.1 ([web.dev](https://web.dev/articles/vitals)).

### 3A. Internal linking

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| Key pages in the nav or footer | Revenue pages are linked site-wide; low-value pages stay out of the nav | Read the nav and footer on three templates | P1 |
| No orphan pages | Every page that matters has at least one internal link in | Orphan report with analytics, Search Console or the sitemap connected to the crawler | P1 |
| Three clicks deep at most | Key pages sit at depth three or less | Crawl depth column | P1 |
| Descriptive anchors | No "click here" or "learn more"; the anchor names the target topic without stuffing | Anchor text export | P2 |
| Breadcrumbs marked up | A visible trail, with BreadcrumbList JSON-LD that matches it | Inspect each template, not just one page | P2 |
| No broken internal links | Zero internal links point to a 4xx | Inlinks to 4xx URLs in the crawl | P1 |

### 3B. On-page

For titles and descriptions, the [SERP Preview](/tools/serp-preview/) shows roughly where Google will cut them off.

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| One H1 per page | No page has two, which is usually a template problem | Crawl report of pages with multiple H1s | P2 |
| No missing H1 | Every product and landing page has a descriptive H1 | Crawl report of pages with no H1 | P1 |
| Unique descriptions | No duplicates on pages that matter, each around 155 characters | Crawl report of duplicate and long descriptions | P2 |
| Unique titles | Every indexable page has its own title, topic first, short enough to show in full | Crawl report of missing, duplicate and over-width titles | P1 |
| Dated titles current | Roundups and listicles show this year in the title, description and headings | An annual pass, only on pages that carry a year | P2 |
| Clean URLs | Lowercase, hyphenated, descriptive and stable; every change is 301'd | /technical-seo-audit passes; /page?id=8842&ref=x fails | P2 |
| One page, one URL | Parameter, filter, trailing-slash and protocol variants consolidate to one | 301s and canonicals; most work on e-commerce sites | P1 |
| Key pages self-canonical | Homepage, product and landing pages are canonical to themselves | Crawl report of canonical targets; stops ?utm_ variants competing | P1 |
| Core Web Vitals pass | LCP 2.5 s or less, INP 200 ms or less, CLS 0.1 or less at the 75th percentile | Search Console's Core Web Vitals report, on mobile and desktop | P2 |
| Speed work done | PageSpeed Insights opportunities are addressed | Serve WebP or AVIF images through a CDN, minify CSS and JS, defer render-blocking files | P2 |
| Mobile parity | The mobile page has the same copy, links and structured data as desktop, and the layout holds at phone, tablet and desktop widths | Lighthouse, or URL Inspection's crawled page | P1 |
| Alt text | Informative images are described; decorative ones use alt="" | Crawl report of images missing alt | P2 |
| Favicon eligible | Square, at least 8x8 px and ideally over 48x48, on a crawlable URL linked from the homepage head (Google) | Load the favicon URL, then check robots.txt doesn't block it | P2 |

**On this site:** every page has exactly one H1, and titles and descriptions are all unique. Two rows failed. The category pages had no BreadcrumbList markup, which I fixed with this block in the category template:

```json
{ "@type": "BreadcrumbList", "itemListElement": [
  { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://khadijazaman.com/" },
  { "@type": "ListItem", "position": 2, "name": "Blog", "item": "https://khadijazaman.com/blog/" },
  { "@type": "ListItem", "position": 3, "name": "Optimization", "item": "https://khadijazaman.com/blog/category/optimization/" } ] }
```

The other was ten pages with descriptions longer than 160 characters, the longest at 305 on a category page. The category pages were reusing their on-page intro as the description, so they now get a short one of their own; after the fix every description runs 127 to 159 characters. It was a P2, so it went last.

## Stage 4. Trust & Entity: is a real author behind it?

An engine deciding whose words to repeat needs to know who wrote them. This stage makes that easy: real people on the page, and the same identity in the markup. The [Schema Generator](/tools/schema-generator/) gives a clean starting point for the markup half.

**Evidence.** Google's structured data rules require markup to describe what's visible on the page, and treat markup that doesn't as a policy issue ([structured data guidelines](https://developers.google.com/search/docs/appearance/structured-data/sd-policies)). That's the reason for the "values match the visible page" row, and for never adding ratings you can't show.

### 4A. People and proof

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| Trust pages in the footer | Contact, Terms and Privacy are linked, with current contact details | Read the footer; missing pages are a red flag on money and health topics | P1 |
| Live social links only | No abandoned profiles are linked | A quarterly click-through of every social link | P2 |
| A real About page | It names the team, the history, the customers and why the business exists | Read it as a stranger would | P1 |
| Named authors | Every post has a real author with an author page, never "Admin" or "Team" | Spot-check ten posts | P1 |
| Checkable author boxes | A real photo, a two to three line bio of relevant expertise, and links to LinkedIn, X or a site | Spot-check ten posts | P1 |
| Claims sourced | Statistics and claims link to studies, official data or recognised industry sources | Spot-check posts; another blog's opinion doesn't count as proof | P1 |

### 4B. Schema

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| Organization on the homepage | JSON-LD with name, logo, sameAs and contactPoint | Rich Results Test on the homepage | P1 |
| SoftwareApplication (SaaS only) | name, applicationCategory and operatingSystem on product pages | Rich Results Test on a product page | P1 |
| BlogPosting on every post | headline, a linked author, datePublished, dateModified and image | Rich Results Test on a sample of posts | P1 |
| BreadcrumbList site-wide | Matches the visible trail on every template | Rich Results Test per template | P2 |
| AggregateRating only from real reviews | Ratings come from G2, Capterra or your own verified reviews | Fabricated ratings risk a manual action | P2 |
| Markup validates | Zero errors in both validators | Rich Results Test and validator.schema.org; errors first, then warnings | P1 |
| Markup matches the page | Recommended fields are filled and every value appears on the page | Read the JSON-LD next to the page copy | P1 |
| Enhancement reports clean | Search Console flags no structured data errors | Fix, then request validation | P1 |

**On this site:** a personal site uses a Person entity in place of Organization, so that row is N/A here. Every page credits one Person, `https://khadijazaman.com/#person`, whose sameAs lists the four accounts I control; the full list, with mentions kept separate, is on my [profiles page](/profiles/). Every post's BlogPosting takes dateModified from git, so it changes only when the post does.

## Stage 5. Retrieve & Cite: will AI engines read and cite it?

The shortest stage, and the one where advice most often goes wrong, nearly always on which bots to allow.

**Evidence.** OpenAI runs GPTBot for training, OAI-SearchBot for ChatGPT search and ChatGPT-User for fetches a person asks for, and each obeys its own robots.txt group ([OpenAI](https://developers.openai.com/docs/gptbot)). Anthropic splits the same way into ClaudeBot, Claude-SearchBot and Claude-User ([Anthropic](https://support.claude.com/en/articles/8896518)). Perplexity runs PerplexityBot for its index and Perplexity-User for live fetches, and its own documentation says Perplexity-User generally ignores robots.txt because a person started the request ([Perplexity](https://docs.perplexity.ai/)).

![Sketch-note grid with three columns, Training, Search and User fetch, and three rows. OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User. Anthropic: ClaudeBot, Claude-SearchBot, Claude-User. Perplexity: training not covered here, PerplexityBot, Perplexity-User. A red cross notes that blocking the search bot keeps you out of that engine's answers; a green tick notes that blocking a training bot affects training only. A footnote says Perplexity's own documentation states that Perplexity-User generally ignores robots.txt, because a person triggered the request](/static/uploads/llm-checklist-fig2-bots.png)
*Figure 2. Block training bots if you want to; the search and user-fetch bots are the ones that decide citations.*

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| AI search and user-fetch bots allowed | No Disallow applies to OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot or Perplexity-User | Read robots.txt group by group, then check the CDN and firewall bot settings | P0 |
| Buyer-phrased FAQs | Questions read the way buyers ask an assistant, such as "How do I get my SaaS cited by ChatGPT?", not "What is SEO?" | Pull questions from sales calls, support tickets and Reddit threads | P1 |
| FAQ answers in the source | Same rule as Stage 2 | View Source | P0 |
| llms.txt at the root | Markdown with the brand as the H1, a one-line summary and grouped links to key pages | Load yoursite.com/llms.txt | P2 |

The [Query Fan-Out Generator](/tools/query-fan-out/) is where my buyer-phrased questions usually start: it shows the sub-questions an engine generates from one prompt. What makes the answers liftable once they're on the page is in [how to get cited by ChatGPT, Gemini and Perplexity](/blog/get-cited-by-ai-search/).

**On this site:** robots.txt names every bot in Figure 2 explicitly, each with its own group, and disallows only `/admin/`. An excerpt, with the Content-Signal lines trimmed:

```
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Perplexity-User
Allow: /
```

Three of those groups (Claude-SearchBot, Claude-User and Perplexity-User) went in during this run. The bots were already allowed by the catch-all `User-agent: *` rule, so nothing was blocked before; naming them makes the intent readable without knowing how robots.txt groups work.

## Stage 6. Publish gate: is this URL ready to go live?

Stages 0 to 5 describe the site. This stage describes one URL on the day it ships, and most rows point back to a site-level rule, so the gate stays short.

**Evidence.** Google reads both redirects and rel="canonical" as strong signals for which URL is the real one ([canonicalization](https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls)). A URL that moves after launch has to update both, plus every internal link, which is why the slug gets settled before the page ships.

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| Final URL set | The slug is lowercase, hyphenated and in the right section folder, and the canonical points to it | Check against the site's URL map; moving it later costs a 301 and a canonical update | P1 |
| Keyword in title, H1 and URL | The primary keyword is in all three, near the start of the title | View Source | P1 |
| Open Graph tags | og:title, og:description, og:image and og:url are present and the share preview shows the right image | LinkedIn Post Inspector or Facebook's Sharing Debugger | P2 |
| HTML errors fixed | Zero errors in the W3C Nu validator; warnings are optional | validator.w3.org/nu, prioritising unclosed tags and duplicate IDs | P2 |
| Design checked | The layout matches the approved design, with nothing overlapping or shifting | Review at phone, tablet and desktop widths | P1 |
| Site rules hold for this URL | 200 status, content in raw HTML, self-canonical, valid structured data, Core Web Vitals, mobile parity, alt text and no broken links | Run each of those rows on this one URL | P0 |
| Staging protections removed | No noindex or nofollow carried over from staging, while staging itself stays protected | View Source and the response headers on the live URL | P0 |
| Page events fire | CTA clicks and form submits on this page reach analytics | Tag Manager preview plus analytics debug view | P1 |
| In the sitemap and requested | The URL is in the sitemap, and URL Inspection shows it indexed or requested | Search Console, URL Inspection, Request indexing | P0 |

**On this site:** this post went through the gate before it was published. Its URL is final at /blog/llm-technical-seo-checklist/, it has one H1, its canonical points to itself, it carries all four Open Graph tags, it's in the sitemap, and Lighthouse scored it 100 for accessibility on a local build. The page events row fails, honestly: this site doesn't track CTA clicks as events, so there's nothing to fire.

## Module: International SEO (multi-region sites only)

**Evidence.** Every language version has to list every other version, itself included, or Google may ignore the annotations ([localized versions](https://developers.google.com/search/docs/specialty/international/localized-versions)).

| Check | Pass when | How to verify | Priority |
| --- | --- | --- | --- |
| Language subfolders | /en/, /de/, /fr/ on one domain | Subdomains only when required; ccTLDs only for genuinely separate markets | P1 |
| Hreflang complete | Every version lists every version, itself included, with valid codes, in both directions | Hreflang report in the crawler | P0 |
| x-default set | It points to the global or primary version | Hreflang report in the crawler | P1 |
| html lang matches | lang="de" on pages marked hreflang="de" | Language report in the crawler; low impact on its own | P2 |

**On this site:** N/A. It serves one language, declared as `lang="en"` on every page.

## Six things most technical SEO checklists get wrong

Each ends with how sure I am. **Verified** means I checked the primary documentation, **Secondary** means trade press only, and **Inferred** means it follows from how the systems work but I haven't seen it documented.

**1\. They treat "AI bots" as one bot.** Blocking GPTBot to stay out of training is a reasonable choice, and it doesn't touch citations. Blocking OAI-SearchBot does: OpenAI says sites that opt out of it won't be shown in ChatGPT search answers. The same split exists at Anthropic and Perplexity, as Figure 2 shows. *Verified ([OpenAI](https://developers.openai.com/docs/gptbot), [Anthropic](https://support.claude.com/en/articles/8896518), [Perplexity](https://docs.perplexity.ai/)).*

**2\. They rank llms.txt too high.** llms.txt is cheap to add, so add it, and treat it as P2. It never outranks crawler access or content in the raw HTML. *My judgement, not a documented rule.*

**3\. They still promise FAQ rich results.** Google limited FAQ rich results to well-known government and health sites in August 2023, and trade press reports it stopped showing them for all sites in May 2026. FAQ content in the raw HTML still helps AI extraction, so keep the content and drop the expectation of a results-page feature. *Verified for 2023 ([Google Search Central](https://developers.google.com/search/blog/2023/08/howto-faq-changes)); Secondary for 2026.*

**4\. They stop at robots.txt.** robots.txt can allow a bot while a CDN or firewall rule blocks it, and the robots.txt check passes anyway. The fix depends on the vendor: Perplexity publishes its crawler IP ranges, so you can allowlist them, while Anthropic says it doesn't publish IP ranges and that IP blocking may not work as an opt-out. *Verified for the vendor policies ([Perplexity](https://docs.perplexity.ai/), [Anthropic](https://support.claude.com/en/articles/8896518)); Inferred for how often it happens. Confirm in your CDN dashboard and access logs.*

![Sketch-note flow of four boxes: OAI-SearchBot asks for the page, robots.txt says Allow with a green tick, a red hatched CDN or WAF box applies a block rule with a red cross, and Your page is never fetched. A caption says the robots.txt check passes and the bot still never reaches the page, so check bot management settings and confirm real AI bot hits in server logs](/static/uploads/llm-checklist-fig3-firewall.png)
*Figure 3. Two layers decide access. Most audits only read the first one.*

**5\. They end at implementation.** A checklist with no outcome metric can't tell you whether any of it worked. I add a measurement step: track citations and mentions per engine (ChatGPT, Perplexity, Gemini, AI Overviews) before and after the fixes. The [BERAP Map](/frameworks/r2a/berap/) is how I measure the recall side of that. *Inferred.*

**6\. They carry checks that outlived their tools.** FID stopped being a Core Web Vital when INP replaced it on 12 March 2024. Google retired the Mobile-Friendly Test and the Mobile Usability report on 1 December 2023, so mobile checks now run through Lighthouse and URL Inspection. And "keep titles under 60 characters" is a display guide, not a Google rule: Google sets no limit and truncates by width. *Verified ([web.dev](https://web.dev/blog/inp-cwv), [Google Search Central](https://developers.google.com/search/blog/2016/05/a-new-mobile-friendly-testing-tool), [title links](https://developers.google.com/search/docs/appearance/title-link)).*

## I ran it on this site

A checklist is easier to trust once you've seen it fail something. I ran every row I could check from this site's own build against khadijazaman.com on 2 October 2026. The rows that need live access, such as Search Console, Bing, HTTPS redirects and CDN rules, aren't in this run.

![Sketch-note scorecard titled khadijazaman.com run against the checklist, listing checks by stage with a green tick or an orange question mark. Ticks for sitemap URLs, robots.txt, noindex, canonicals, FAQ answers in raw HTML on 7 pages, one H1 on all 28 pages, unique titles and descriptions, structured data and all six AI search and user-fetch bots. Ticks marked fixed for meta descriptions, where 10 pages ran past 160 characters, and for breadcrumb markup on two category pages. A question mark for Search Console, Bing, HTTPS and CDN rules, which need live access](/static/uploads/llm-checklist-fig4-this-site.png)
*Figure 4. Two fails fixed during the run, and four rows that need live access.*

The run is the examples above in one place: a breadcrumb gap fixed in five lines, ten long descriptions rewritten, and three robots.txt groups added for readability rather than access. The build itself is written up in [built for two readers](/blog/built-for-two-readers/).

If you want this run on one of your own pages, the free [AI-search access check](/contact/#access-check) covers the Stage 2 and Stage 5 rows that decide whether AI search can reach and cite it.

## Frequently asked questions

### Should I block GPTBot to stay out of AI training?

You can. GPTBot is OpenAI's training crawler, so blocking it affects training only. ChatGPT search uses OAI-SearchBot, and OpenAI says sites that opt out of OAI-SearchBot won't be shown in ChatGPT search answers, so leave that one allowed if you want to be cited.

### Does llms.txt help a site get cited by AI?

There's no confirmed evidence that it does. I keep llms.txt at P2: it's cheap to add and harmless, but I don't treat it as a citation lever, and it never outranks crawler access or content in the raw HTML.

### Is FAQ schema still worth adding?

Not for rich results. Google limited FAQ rich results to well-known government and health sites in August 2023. Keep the FAQ content itself in the raw HTML, because that still helps AI systems extract answers.

---
Markdown twin of https://khadijazaman.com/blog/llm-technical-seo-checklist/, generated from the rendered page at build time. Cite the HTML URL. How to cite: https://khadijazaman.com/llms.txt
