Guide

How to get your website cited by ChatGPT, Perplexity and Google AI Overviews: the 43-point checklist (2026)

Short answer: AI engines cite pages they can crawl, from entities they can identify, that other sites corroborate. Start by allowing the AI crawlers in robots.txt, adding Organization schema with sameAs links, publishing an llms.txt and a public pricing page with Offer schema, then checking a fixed set of buyer prompts yourself. This guide covers 32 of the 43 checks in full. The 11 high-impact checks are counted by the free audit and revealed in paid reports.

Check your site against all 43 in 30 seconds. Free, no account.

This is the checklist behind the SearchBind free audit. We publish 32 of its 43 checks with the reason and the fix, so you can do that part yourself. The 11 checks with the largest effect on citations are named, explained and fixed for your pages in paid reports; the free scan still runs and counts them.

What "cited by AI" actually means

Three different things get lumped together. A mention is your brand name appearing in an answer. A citation is a link to your page as a source. A recommendation is the answer naming you as the choice for a need. Perplexity and Google AI Overviews cite sources on almost every answer, so citations are where to start. ChatGPT cites less often but mentions brands constantly, so entity clarity matters more there. Track all three.

How the engines pick sources

Perplexity retrieves roughly ten candidate pages per query in real time and cites three to five. Google AI Overviews draws from Google's index with a strong preference for pages already ranking plus forums, video and social. ChatGPT search combines a search index with its own crawler, OAI-SearchBot. Each engine filters the candidates before it cites: can it fetch the page, does it know who is behind it, does the page fit the prompt, and do other sites confirm it.

How AI engines pick sources A buyer prompt goes to retrieval from a search index or a live fetch, which returns about ten candidate pages. Four filters remove pages: can it be fetched, is the entity clear, does the page fit the prompt, is it corroborated elsewhere. Three to five pages are cited in the answer. BUYER PROMPT best CRM for small marketing agencies RETRIEVAL ChatGPT search: search index + OAI-SearchBot Perplexity: live fetch · AI Overviews: Google index about 10 candidate pages Can the crawler fetch it?area A Is the entity behind it unambiguous?area B Does the page fit the prompt?areas C, D Do other sites corroborate it?B2, D6 3 to 5 cited
Simplified. Each engine weighs these differently; all four filters map to checks in the list below.

A. Crawl access: can AI engines read you? (10 checks, 6 shown)

The most common failure we see is also the cheapest to fix. In 2023 and 2024 many sites added rules blocking AI crawlers. Those rules are still there.

  • A1GPTBot, OAI-SearchBot and ChatGPT-User are allowed. GPTBot feeds training, OAI-SearchBot feeds ChatGPT search, ChatGPT-User fetches pages during a conversation. Blocking the second two removes you from ChatGPT's answers.
  • A2High-impact check, revealed in paid reportsHIGH
  • A3ClaudeBot and Claude-SearchBot are allowed. Smaller traffic today, growing fastest among AI referrers in late 2026.
  • A4High-impact check, revealed in paid reportsHIGH
  • A5High-impact check, revealed in paid reportsHIGH
  • A6Applebot-Extended, CCBot and Meta-ExternalAgent are handled explicitly. Decide, do not default.
  • A7HTTPS, a 200 response, no redirect chains longer than two hops.
  • A8High-impact check, revealed in paid reportsHIGH
  • A9sitemap.xml exists, is referenced in robots.txt, and has accurate lastmod dates.
  • A10llms.txt exists at the root. See below.

A minimal robots.txt for the public checks above:

User-agent: *
Allow: /

User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Claude-SearchBot
Allow: /

Sitemap: https://example.com/sitemap.xml

And a minimal llms.txt. It is Markdown: an H1 with your name, a blockquote with one sentence, then lists of links with one-line descriptions.

# Example Co

> Example Co is a [category] for [who]. It [does what] so that [outcome].

## Key pages
- [Pricing](https://example.com/pricing/): plans and prices
- [About](https://example.com/about/): the people behind Example Co

## Docs
- [Getting started](https://example.com/docs/start/): first steps

B. Entity clarity: does AI know who you are? (7 checks, 6 shown)

Language models resolve brands as entities. If your name is spelled three ways and nothing outside your site confirms you exist, the model hedges and names a competitor it is sure about.

  • B1Organization or Person JSON-LD on the home page with name, url, logo and sameAs.
  • B2sameAs links to LinkedIn, X, Crunchbase, G2, GitHub, Wikipedia or Wikidata where applicable. These are the corroboration signals.
  • B3High-impact check, revealed in paid reportsHIGH
  • B4An About page with named people, photos and roles.
  • B5Consistent brand spelling across title tags, schema, footer and social bios.
  • B6A contact page with ContactPoint schema.
  • B7A logo file that is accessible and referenced in schema.

Minimal Organization schema:

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Example Co",
  "url": "https://example.com/",
  "logo": "https://example.com/logo.png",
  "sameAs": [
    "https://www.linkedin.com/company/example",
    "https://x.com/example",
    "https://www.crunchbase.com/organization/example"
  ]
}
</script>

C. Answer-first structure: is your content extractable? (10 checks, 6 shown)

This area covers how your pages are built: headings, tables, freshness, authorship and length. Four of its checks are high-impact and revealed in paid reports.

  • C1Exactly one H1 per key page that states the topic plainly.
  • C2High-impact check, revealed in paid reportsHIGH
  • C3High-impact check, revealed in paid reportsHIGH
  • C4High-impact check, revealed in paid reportsHIGH
  • C5Tables or lists for comparisons, pricing and specs.
  • C6High-impact check, revealed in paid reportsHIGH
  • C7Visible last-updated dates and dateModified in schema.
  • C8Over 300 words of unique text on key pages, and no thin duplicates indexed.
  • C9Author bylines with Person schema on articles.
  • C10No interstitials or cookie walls hiding content on first fetch.

D. Prompt coverage: do you answer what buyers actually ask? (6 checks, 4 shown)

If you have no page that matches what a buyer asks, you cannot be cited for it. Two of the checks here are high-impact and revealed in paid reports.

  • D1High-impact check, revealed in paid reportsHIGH
  • D2High-impact check, revealed in paid reportsHIGH
  • D3A public pricing page with Offer schema. Hidden pricing is invisible pricing.
  • D4Use-case or industry pages for your top three segments.
  • D5Glossary or "what is [term]" content for your category's terms.
  • D6Reviews or case studies with concrete numbers. First-party proof outranks generic claims.

E. Technical hygiene that still matters (10 checks, all shown)

  • E1Title tag 30 to 60 characters, unique, containing the entity name.
  • E2Meta description present, 70 to 160 characters.
  • E3Self-referencing canonical tag.
  • E4Open Graph and Twitter card tags.
  • E5Mobile viewport, no horizontal scroll.
  • E6HTML under 300 KB, no render-blocking third-party scripts before content.
  • E7Images with alt text.
  • E8Internal links from the home page to every key page with descriptive anchors.
  • E9hreflang if you publish in several languages.
  • E10An IndexNow key, so search engines that support it pick up changes within minutes instead of weeks.

What to expect, and when

Perplexity and Google AI Overviews fetch live pages, so a well-structured new page can be cited within two to six weeks. ChatGPT depends more on its search index and periodic refreshes, so plan for three to six months of consistent work before mentions become regular. Measure with a fixed prompt set, 20 to 30 prompts phrased the way buyers type them, checked by hand on each engine at a regular interval. Anything less consistent hides the signal in noise.

The order to do it in

  1. Crawl access (A). One hour. Allow the AI crawlers, fix the sitemap, publish llms.txt. Nothing else matters if the bots cannot read you.
  2. Entity clarity (B). Two hours. Organization schema with sameAs, an About page with real people, one spelling of your name everywhere.
  3. Public pricing with Offer schema (D3). One evening.
  4. The remaining public checks in C, D and E. They fit in the gaps.
  5. The 11 high-impact checks. Run the free audit to see how many of them your site fails; the paid report names them and gives the fix.

Want this run on your site automatically? The free SearchBind audit scores all 43 checks in 30 seconds. The Website plan, $69 one-time, gives you the full fix list including the 11 high-impact checks, the snippets generated for your pages and 4 to 6 content drafts, in one report within 24 hours. No developer? We implement the list for $50 an hour, capped at $150.

Frequently asked questions

How long does it take to get cited by ChatGPT?

Perplexity and Google AI Overviews fetch live pages, so well-structured new content can be cited within 2 to 6 weeks. ChatGPT relies more on its search index and periodic model updates, so plan for 3 to 6 months of consistent publishing before mentions become regular.

Do I need to block or allow AI crawlers?

If you want to be cited, allow them. Blocking GPTBot, OAI-SearchBot or ClaudeBot in robots.txt removes you from those engines' answers. Many sites blocked AI crawlers in 2023 and 2024 and forgot.

What is llms.txt and does it matter?

llms.txt is a plain-text file at your site root that lists your most important pages with one-line descriptions, in Markdown. It is an emerging convention supported by several AI agents and crawlers, cheap to add, and a useful forcing function to decide which pages matter. It does not replace schema or good content.

Why are 11 checks not described here?

They are the high-impact checks, the ones with the largest effect on whether AI engines cite a page. The free SearchBind scan runs and counts them, and paid reports name each one, explain it and give the fix for your pages.

Sources and further reading

  • OpenAI, "Overview of OpenAI crawlers" (GPTBot, OAI-SearchBot, ChatGPT-User).
  • Perplexity, crawler documentation.
  • Google, AI Overviews documentation in Search Central.
  • Meltwater, 2026 study of 9.5 million AI citations on LinkedIn content.
  • Axios, 20 August 2026, on Reddit's drop in ChatGPT citations after blocking AI crawlers.
  • llmstxt.org, the llms.txt proposal.