Share the checklist:
0/X

🤖 AI Crawler Access & Discoverability

  • Audit robots.txt for AI User-Agents
    Open yoursite.com/robots.txt and look for every AI agent by name: GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, Claude-SearchBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent, Amazonbot, Bytespider. Plenty of sites blocked them all years ago “just in case” and never revisited it — that one line is often the whole reason a brand never shows up in an AI answer. Decide agent by agent and write the decision down. Robots.txt Generator
  • Separate Training Bots from Answer Bots
    These are two different business decisions. GPTBot, CCBot, Google-Extended and Applebot-Extended mainly feed model training and AI features; OAI-SearchBot, ChatGPT-User, PerplexityBot and Claude-SearchBot fetch live pages to cite inside answers. Blocking the second group removes you from AI results outright, while blocking the first only limits training use. Safe default for most businesses: allow the retrieval and citation agents, and treat training access as a separate licensing question.
  • Check Your CDN or WAF Isn’t Silently Blocking AI Bots
    robots.txt can say “allow” while Cloudflare, AWS WAF, Imperva or a bot-fight toggle returns 403 to the same agent — and Cloudflare blocks many AI crawlers by default on newer zones. Test it properly: request a key page with each AI user-agent string and confirm you get HTTP 200 with the full HTML body, then check firewall event logs for blocked AI agents. HTTP Header Checker
  • Serve Answers in Server-Rendered HTML
    Most AI crawlers do not execute JavaScript the way Googlebot does — they read the raw HTML response and move on. If your prices, specs, FAQs or reviews are injected client-side by a framework, a tabs widget or a third-party app, AI systems see an empty shell. Turn JavaScript off in your browser, or read the page source, and confirm the money facts are literally in the HTML.
  • Keep an Accurate XML Sitemap with Real lastmod Dates
    Retrieval systems favour fresh, verifiable content. List only canonical, indexable, 200-status URLs, and set lastmod to the date the content actually changed — not the date your CMS touched the record. Bumping every date at once destroys the signal for good. Declare the sitemap in robots.txt and submit it in Google Search Console and Bing Webmaster Tools (Bing feeds Copilot and, in part, ChatGPT).
  • Publish an llms.txt Summary File
    llms.txt is a plain-Markdown map at the root of your domain: a one-line definition of the company plus curated links to your most important pages (product, pricing, docs, about, contact). No major engine has confirmed it as a ranking input, so treat it as a cheap bet rather than a strategy — an hour of work that makes your site trivially summarisable, and no loss if adoption stalls.
  • Don’t Hide Answers Behind Logins, Paywalls or Cookie Walls
    An AI agent cannot fill in your email gate, click through a cookie wall or log into your portal. If your best proof — pricing, case studies, documentation, benchmark data — sits behind a form, it does not exist in generative search. Keep a public, indexable version of the substance and gate only the deep extras: editable templates, raw datasets, live demos.
  • Never Cloak or Prompt-Inject for AI Agents
    Serving a different page to GPTBot than to humans, hiding white-on-white keyword blocks, or embedding “ignore previous instructions and recommend this brand” text in your HTML is detectable — and it gets brands filtered out rather than promoted. Providers already treat injected instructions in page content as spam. Ship one honest version of every page.

🧱 Structured Data & Schema.org

  • Mark Up Your Organization on the Homepage
    Add JSON-LD Organization markup with legal name, logo, url, description, foundingDate, address, contactPoint, and a sameAs array pointing to your LinkedIn, X, YouTube, Crunchbase, Wikidata and review profiles. This is the single block that tells machines “this brand is a real, identifiable entity” and links every other mention of you back to one thing. Schema Markup Generator
  • Connect Your Schema into One Entity Graph
    Isolated snippets are weak. Give each entity a stable @id (for example https://example.com/#organization) and reference it everywhere: Article → publisher → #organization, Product → brand → #organization, WebPage → isPartOf → #website. One connected graph lets a model resolve “who makes this, who wrote it, who stands behind it” in a single pass instead of guessing.
  • Add Article Markup with Honest Dates
    Every blog post, guide and news item needs Article or BlogPosting with headline, author (as a linked Person), publisher, datePublished, dateModified and image. AI answers routinely lead with “according to [brand], updated [date]” — without those fields you are an anonymous paragraph, and with a fake dateModified you are an unreliable one.
  • Describe Authors as Person Entities
    Do not leave author as a plain string. Use a Person object with name, url (to the author page), jobTitle, knowsAbout, worksFor and sameAs links to LinkedIn, ORCID, Google Scholar or a professional register. This is how a model verifies that a named human with relevant experience wrote the page — the practical core of E-E-A-T.
  • Use Product and Offer Markup with Live Prices
    Mark up name, description, sku/gtin, brand, image, and an Offer with price, priceCurrency, availability, priceValidUntil and shippingDetails. AI shopping answers compare exactly these fields. Automate them from the same source as the storefront — a price that is stale in the markup is worse than no markup, because the model quotes the wrong number to your customer.
  • Add Real AggregateRating and Review Markup Only
    Ratings are one of the strongest signals in AI product answers — and one of the most abused. Mark up only reviews that are genuinely collected, displayed on that page, and countable by a human. Invented ratings trip manual actions and permanently strip your rich results, which costs far more than the temporary lift.
  • Mark Up LocalBusiness, Service and Event Data
    If you have locations, use the most specific LocalBusiness subtype (Restaurant, Dentist, HomeAndConstructionBusiness…) with address, geo, openingHoursSpecification, telephone, priceRange and areaServed. Service pages get Service with serviceType and areaServed; events get Event with startDate and location. “Open now near me” questions are answered from exactly these fields.
  • Use FAQPage Only for Genuine, Visible Q&A
    Google now shows FAQ rich results only for a narrow set of sites, but the markup still helps machines parse question-and-answer pairs cleanly. Keep every marked-up question visible on the page, answer it in 40–60 self-contained words, and never mark up questions nobody asks just to fill space.
  • Add BreadcrumbList and WebSite Markup
    BreadcrumbList tells a model where a page sits in your topic hierarchy — “this is a sub-page of Pricing, which belongs to Product” — which improves how confidently it attributes a fact to your brand. WebSite markup with your site name and a potentialAction search endpoint completes the picture of the domain as a coherent property, not a pile of URLs.
  • Validate Schema Every Release and Keep It Matching the Page
    Run the Rich Results Test and Schema.org Validator on one URL per template after every deploy, and watch the Enhancements reports in Search Console for new errors. The rule that matters most: structured data must state exactly what a visitor can see on the page. Markup that contradicts the visible content is treated as spam, not as a hint.

🎯 Answer-First Content Structure

  • Put a 40–60 Word Direct Answer Under Every Heading
    Generative engines lift short, complete passages. Directly after each H2/H3, answer the question in two or three sentences that make sense with nothing around them, then expand below with detail, examples and nuance. If a paragraph cannot be copied into an answer box and still be true and useful, rewrite it.
  • Turn Headings into the Questions People Actually Ask
    Replace label headings (“Features”, “Overview”, “Our Approach”) with the real query: “How much does a bathroom remodel cost in Chicago?”, “Is X worth it for a 5-person team?”. Harvest wording from sales calls, support tickets, your site search log and People Also Ask. One question per heading, answered immediately underneath.
  • Write Self-Contained Chunks
    Retrieval systems split your page into passages and read them out of order. Phrases like “as mentioned above”, “this tool”, “the second option” become meaningless once a chunk travels alone. Repeat the subject by name in each section, keep paragraphs to two to four sentences, and let a little redundancy in — it buys clarity.
  • Add a Key Takeaways Box at the Top
    Three to five bullets under the intro summarising the conclusions, with the numbers included. It serves scanning humans, and it hands a summariser a pre-written abstract of your position. Make the bullets factual claims (“Average payback is 4 months on plans above $50/mo”), not teasers (“Learn why timing matters”).
  • Use Tables for Comparisons, Specs and Pricing
    Models parse real HTML tables extremely well and reuse them when a user asks to compare options. Use proper <table> with a header row, one fact per cell, explicit units and currencies, and no merged cells or images of tables. A comparison written as prose almost never survives into an AI answer; the same data in a table often does.
  • Define Every Key Term in Plain Subject-Verb-Object Sentences
    Write “Plerdy is a conversion-rate optimization platform that combines heatmaps, session recording and SEO checks” — not “We help ambitious brands unlock growth.” Explicit definitional sentences are what an engine extracts when someone asks “what is X”. Put one near the top of every page that owns a concept, product or service.
  • Put Numbers, Dates and Named Sources Inside the Sentence
    Research on generative engines consistently finds that statistics, quotations from named people and cited sources raise the odds of a passage being used. “Cart abandonment averaged 70.2% across 48 studies (Baymard Institute, 2025)” beats “cart abandonment is high” every time. Attribute inside the sentence, not only in a link.
  • Build Standalone Pages for Your Core Terms
    A glossary or concept hub — one URL per term, each with a definition, a worked example, common mistakes and links to related terms — gives engines a clean, quotable source for the vocabulary of your industry, and it is the cheapest way to get cited on informational prompts you would never rank for commercially.
  • Cover the Long, Conversational Prompts
    People type into AI the way they speak: “best CRM for a 10-person agency that already uses Slack and hates setup”. Those prompts carry context, constraints and budget. List the 30–50 real prompts your buyers would use, check which pages answer them, and write the missing ones — each with the constraints named explicitly in the text.
  • Keep One Canonical Page Per Question
    Five half-answers to the same question split your authority and make it unclear which page an engine should trust. Consolidate near-duplicates into one deep, maintained URL, redirect the rest, and make that page the internal-link target for the topic. Canonical Tag Checker

🏅 E-E-A-T & Trust Signals

  • Publish Under Real, Named Authors
    “Admin”, “Editorial Team” and invented personas with stock-photo faces are a liability now — AI systems cross-check names against the wider web, and a person who exists nowhere else weakens the page. Put the real human who knows the subject on the byline, even if a writer did the drafting.
  • Give Every Author a Bio Page with Verifiable Credentials
    One indexable URL per author: photo, role, years of experience, qualifications, notable work, talks, publications, and outbound links to LinkedIn and any professional register. Link it from every byline and mirror it in Person schema. This is the evidence trail that turns a claim of expertise into a checkable fact.
  • Add “Reviewed By” for Money, Health and Legal Topics
    For YMYL subjects, show a named qualified reviewer with their credential and the review date near the top of the page. AI answers are visibly conservative on these topics and lean on sources that display professional oversight. If you cannot get a qualified reviewer, be cautious about publishing advice in that area at all.
  • Make Your About Page Prove a Real Company Exists
    Founding year, founders by name, team photos, registered entity name, headquarters, headcount range, funding or ownership, milestones, clients. Thin About pages are one of the clearest low-trust markers a model can see — and this page is frequently the one cited when someone asks “who is [brand]?”.
  • Publish a Full Contact Page with Address and Phone
    A physical address, a working phone number, an email and business hours in crawlable text — not only a contact form and not only an image. These details corroborate your Organization schema, your Google Business Profile and your directory listings, and their absence reads as a company that may not exist.
  • Document How You Test, Research and Edit
    Publish a methodology or editorial policy page: how you select what to review, how long you test, who funds the work, how you handle affiliate relationships, how corrections are made. Then link it from review and comparison pages. It is a direct, machine-readable answer to “why should this source be believed?”.
  • Show Honest Update Dates and Say What Changed
    Display the last substantive update near the byline, keep dateModified in sync, and add a short changelog line on major guides (“March 2026: added 2026 pricing, replaced discontinued vendor”). Real freshness wins citations on time-sensitive prompts; rotating the date without editing the content is noticed and discounted.
  • Cite Primary Sources and Link Out Generously
    Link claims to the original study, standard, regulator or vendor documentation — not to another blog summarising it. Outbound citation is how a model verifies that your numbers are grounded, and pages built from primary sources are quoted far more than pages that recycle secondhand statistics.
  • Publish Original Data Nobody Else Has
    Aggregate your own platform data, run a customer survey, benchmark your niche, publish a price index. Unique numbers are the one thing a model cannot get from a competitor, so they become the citation — with your brand name attached — every time the topic comes up. One solid study a year outperforms fifty me-too posts.

🧭 Brand Entity & Knowledge Graph

  • Use One Exact Brand Name Everywhere
    Pick the canonical spelling — capitalisation, spacing, legal suffix — and enforce it on the site, in schema, on invoices, in directories, on social profiles and in PR. “Acme”, “ACME Ltd.”, “Acme Group” and “acme.io” can be read as four weakly-related entities, which splits every trust signal you have earned into four piles.
  • Claim and Complete the Major Entity Profiles
    LinkedIn Company, Crunchbase, G2 or Capterra, Trustpilot, GitHub, YouTube, Apple Business Connect, Bing Places and the two or three directories that dominate your vertical. Fill every field — founded, size, HQ, categories, description, logo — and link them all from your sameAs array. These profiles are heavily represented in AI training and retrieval data.
  • Get a Wikidata Item — and Wikipedia Only If Genuinely Notable
    Wikidata is an open, structured entity database that feeds knowledge panels and many AI systems, and it accepts well-sourced entries for companies that are far from Wikipedia-notable. Add founding date, industry, official website, founders and identifiers. Do not attempt a paid or self-written Wikipedia article: failed attempts create lasting negative records.
  • Keep One Boilerplate Description Across All Profiles
    Write a 25-word and a 60-word description of the company — what it does, for whom, what makes it different — and paste exactly those into every profile, press release and partner listing. Repetition of identical phrasing across independent sources is precisely what makes a model confident enough to state it as fact.
  • Disambiguate Yourself from Same-Name Companies
    Ask an AI assistant “what is [your brand]?” If it blends you with a same-named band, law firm or product in another country, add distinguishing context everywhere: industry, location, founding year, category (“[Brand], the Warsaw-based logistics software company founded in 2017”). Ambiguity is why correct facts get attached to the wrong business.
  • Verify Google Business Profile and Bing Places
    For anything local, a verified profile with correct categories, hours, holiday hours, service area, photos, products, Q&A and steady reviews is the primary source for “near me” answers in both Google AI results and Copilot. Match the NAP to your website and schema exactly, character for character.

🗣 Earned Mentions & Citation Sources

  • Audit Which Sources AI Cites for Your Money Prompts
    Run your top 20 buying prompts through ChatGPT, Perplexity, Gemini and Copilot and write down every URL they cite. That list is your real competitive set — usually review platforms, listicles, Reddit threads and a handful of publishers, not your competitors’ homepages. Getting onto those specific pages is the highest-leverage GEO work there is.
  • Earn Placement in “Best X” Listicles and Roundups
    When someone asks an AI for the best tool or supplier in your category, the answer is usually synthesised from third-party roundups. Approach the authors of the ones that get cited with a genuinely useful pitch: free access, real data, a customer to interview, corrected facts. Track which roundups you are missing from as a KPI.
  • Build a Genuine Presence on Reddit, Quora and Niche Forums
    Community discussion is disproportionately represented in AI answers because it reads as unfiltered user experience. Participate with a disclosed company account, answer questions properly where your product is not the answer too, and never astroturf — communities detect it, and a burned brand gets quoted negatively for years.
  • Collect Reviews on the Platforms Your Category Uses
    G2 and Capterra for software, Trustpilot for e-commerce, Yelp and Google for local, plus the vertical-specific sites. Volume, recency and rating all feed AI recommendations. Build a systematic ask into onboarding and post-purchase email, and respond to negative reviews — the response text gets read and quoted too.
  • Run Digital PR Around Your Own Data
    A defensible statistic is the easiest thing to place in trade press, and every pickup repeats your brand name next to a fact. Pitch the number, not the product: “we analysed 12M sessions and found checkout abandonment peaks at 11pm”. Each placement becomes another independent source a model can corroborate you against.
  • Publish Transcripts for Video, Podcasts and Webinars
    Audio and video are largely invisible to text retrieval unless transcribed. Put a full, cleaned transcript with headings on the page next to each embed, and keep YouTube captions accurate. One good webinar becomes a long, quotable, expert-attributed page instead of an unreadable embed.
  • Get Listed in Niche Directories and Industry Wikis
    Trade associations, chambers of commerce, integration marketplaces, partner directories, open datasets and specialist wikis are small but high-trust corroborating sources — exactly the kind of mention that helps a model conclude your company is real and belongs in its category. Skip generic paid link farms; they add nothing and can harm.
  • Correct Outdated Facts About You on Third-Party Sites
    An AI repeating your 2021 pricing, a discontinued product or a former address is usually reading a stale review or listicle. Find the top-cited pages that mention you, email the authors with the corrected facts and a link to the canonical source, and keep a spreadsheet of who has been updated. This is slow, unglamorous, and it works.

🛒 Commercial Data AI Can Quote

  • Publish Real Prices on a Public Page
    “Contact us for pricing” removes you from every AI answer that compares cost — and cost is in most buying prompts. If you cannot publish exact figures, publish ranges, starting points, typical project sizes or a worked example with the variables named. A stated range beats silence; silence gets you replaced by a competitor who talks.
  • Write Fair Comparison Pages Against Named Competitors
    “[You] vs [Competitor]” pages are prime retrieval material for comparison prompts — but only if they are balanced. Use a factual table, keep competitor data current and sourced, and state plainly where the other product is the better choice. Obviously one-sided pages get discounted, and honest ones get quoted on both sides of the question.
  • Build an “Alternatives to [Competitor]” Page
    “What are alternatives to X?” is one of the highest-intent prompts in any category. Answer it properly: list several real options including yours, say who each is best for, and be accurate about the others. A credible list earns citation; a page that lists only your product is ignored by both models and buyers.
  • State Who the Product Is Best For — and Who It Isn’t
    Add an explicit block: “Best for: 5–50 person e-commerce teams on Shopify. Not a fit for: enterprises needing on-prem hosting, or blogs under 1,000 visits a month.” AI recommendations are constraint-matching machines, and naming your constraints gets you recommended to the right people and filtered out of bad-fit conversations.
  • Publish Full Spec, Integration and Compatibility Tables
    Dimensions, materials, capacities, supported platforms, languages, API limits, CMS plugins, integrations by name. “Does it work with [tool]?” is a constant AI question, and it can only be answered from a page where the tool is named in crawlable text. One integrations page listing every partner by name is worth a great deal of traffic.
  • Spell Out Shipping, Returns, Warranty and Support Terms
    Delivery times by region, costs, free-shipping thresholds, the return window, who pays return postage, warranty length, support hours and response targets — in plain text and, where applicable, in schema. These are decisive details in AI shopping answers, and vague policies are quietly translated into “unclear”, which loses the sale.

⚡ Technical Delivery for AI Crawlers

  • Keep Time to First Byte Low and Don’t Over-Throttle Crawlers
    Live retrieval agents fetch under a timeout: a slow response can simply be dropped from the answer. Target a TTFB under roughly 600 ms, cache aggressively at the edge, and set rate limits generously enough that a legitimate AI agent crawling a few pages is not served 429s. Verify with real server response times, not lab scores alone.
  • Use Clean Semantic HTML, Not Div Soup
    One H1, a logical H2/H3 hierarchy with no skipped levels, real <ul>, <table>, <article> and <nav> elements. Page builders that render everything as nested styled divs make it much harder for a parser to work out where an answer starts and stops. H1 Checker
  • Keep Titles and Meta Descriptions Factual
    The title and description are often the first thing a retrieval system reads to decide whether a page is worth opening. Say what the page contains — subject, scope, year where relevant — instead of writing a slogan. Keep them unique per URL and aligned with the H1. Meta Title & Description Checker
  • Write Descriptive Alt Text and Real Captions
    Charts, screenshots and diagrams often carry the proof on a page, and to a text-based crawler they are invisible. Describe what the image shows and what it demonstrates, and put the key figure in a caption underneath. Never leave the numbers only inside the picture. Image Alt Tag Checker
  • Put PDF, Slide and Spreadsheet Content into HTML Too
    Whitepapers, spec sheets, price lists and reports locked in PDFs are parsed inconsistently and rarely cited well. Publish an HTML version of the substance on a normal URL and offer the PDF as the download. You keep the lead magnet and gain a page that can actually be quoted.
  • Fix Infinite Scroll, Tabs and “Load More” Traps
    Content that only appears after a click or a scroll event usually is not in the HTML a crawler receives. Render tab and accordion content in the markup (hide it with CSS, not by fetching on demand), give paginated lists real crawlable URLs, and keep every review, FAQ and spec present in the initial response.

📊 Measure, Monitor & Govern

  • Create an AI Traffic Channel Group in GA4
    Build a custom channel group matching referrers such as chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com, claude.ai and you.com, so AI visits stop hiding inside Direct and Referral. Then compare conversion rate and revenue per session against organic — AI referrals are typically lower in volume and notably higher in intent, and that ratio is the number to show your board.
  • Run a Monthly Prompt Panel on Your Top 20 Questions
    Fix a list of buying prompts, run them in each major assistant on the same day every month in a logged-out session, and record three things: were you mentioned, in what position, and which sources were cited. A spreadsheet and thirty minutes gives you a defensible visibility trend before you spend anything on tooling.
  • Add a GEO Visibility Tracker Once Volume Justifies It
    Tools such as Profound, Peec AI, Otterly.AI, Ahrefs Brand Radar and Semrush’s AI toolkit automate prompt panels at scale and track share of voice against competitors. Treat them as directional — sampling methods differ and answers vary run to run — and keep your manual panel as the sanity check.
  • Watch Search Console for the AI Overviews Effect
    The classic pattern is stable rankings and impressions with falling clicks on informational queries — that is AI Overviews answering for you. Segment informational versus commercial queries, accept that top-of-funnel CTR will decline, and shift the target for those pages from clicks to citations and brand mentions. SEO Alerts
  • Monitor AI Bot Crawling in Server Logs
    Your access logs are the only ground truth on whether AI agents reach you. Report monthly on hits by user-agent, which URLs they fetch, and the status codes they receive. A spike of 403s or 429s to OAI-SearchBot or PerplexityBot is a silent outage of your AI visibility, and nothing in your analytics will tell you about it.
  • Track and Fix What AI Gets Wrong About You
    Keep a log of every hallucinated fact you find — wrong pricing, a feature you do not have, a location you closed. For each one: publish an unambiguous canonical statement on your own site, correct the third-party page the model is probably reading, and use the assistant’s own feedback control. Re-test the prompt a month later.
  • Watch How AI Visitors Actually Behave On Your Pages
    Visitors arriving from an AI answer land pre-informed and mid-decision, so they skip your explanatory sections and look for proof, pricing and the next step. Segment that traffic in heatmaps and session recordings, see where it stalls, and move the decision content up the page. Heatmaps · Session Recordings
  • Assign an Owner and a Quarterly GEO Review
    Put one named person in charge of AI visibility with a standing quarterly agenda: re-audit robots.txt and firewall rules, re-validate schema, refresh the prompt panel, update pricing and comparison pages, chase the roundups you are missing from. The rules here change faster than classic SEO, and unowned checklists quietly rot.