Migration playbook for keeping AI citations during URL changes — hard 404 vs soft 404, 410 Gone, redirect chains, sitemap cleanup, and refetch monitoring.
guide•7 min read
Specification for serving Brotli, gzip, and zstd to AI crawlers via Accept-Encoding negotiation: which bots support which codecs, fallback rules, and Vary handling.
specification•6 min read
Specification for handling Accept-Language with AI crawlers: avoid auto-redirects, expose hreflang, prefer separate locale URLs, and preserve citation eligibility.
specification•6 min read
Specification for Accept-Language negotiation and html lang attribution that lets AI crawlers detect locale correctly without cross-locale citation leaks.
specification•6 min read
AggregateRating schema specification for AI citations: required fields, decimal handling, parent-type pairings (Product, Course, SoftwareApplication, LocalBusiness), Google policy violations.
specification•9 min read
Specification for AI card thumbnails: aspect ratio, minimum dimensions, file format, alt text, and ImageObject schema patterns that AI search engines extract for rich answer cards.
specification•7 min read
AI citation tracking with server log analysis: identify GPTBot, PerplexityBot, ClaudeBot hits, link them to citations, and measure crawl-to-cite latency.
guide•11 min read
AI crawl budget guide: prioritize high-value pages, reduce noise, and steer GPTBot, ClaudeBot, PerplexityBot, and Google-Extended toward citation-worthy content.
guide•7 min read
Technical reference for the signals AI systems use to discover, access, and prioritize web content — including sitemaps, llms.txt, robots.txt, structured data, and HTTP headers.
reference•4 min read
Compare allowlist vs blocklist strategies for AI crawlers across robots.txt, llms.txt, and CDN edge: trade-offs, decision matrix, and migration path.
comparison•7 min read
HTTP content negotiation (Accept, Accept-Language, Vary) for AI crawlers — serve LLM-friendly variants without cloaking penalties or cache poisoning.
specification•7 min read
Attribute infrastructure cost to GPTBot, ClaudeBot, and other LLM crawlers, then allocate allow, throttle, charge, or block budgets by citation ROI.
framework•8 min read
Reference list of official AI crawler IP range endpoints, user agents, and reverse-DNS verification methods for GPTBot, ClaudeBot, PerplexityBot, Googlebot, and more.
reference•7 min read
How to use Resource Hints, Link headers, and 103 Early Hints to accelerate AI crawler discovery while keeping origin load and crawl budget under control.
specification•7 min read
Reference table of safe rate limits for GPTBot, ClaudeBot, PerplexityBot, and other AI crawlers, with citation-impact tradeoffs and edge recipes.
reference•7 min read
Author entity markup for AI citation: Person schema, sameAs identifiers, credentials, and Wikidata linkage that lift author authority signals.
specification•8 min read
How AI search engines (ChatGPT, Perplexity, Gemini, Google AI Overviews) resolve rel=canonical, hreflang, and parameterized URLs when selecting and citing sources.
reference•8 min read
Reference of freshness signals AI crawlers track — lastmod, dateModified, version banners, changelogs, and substantive republishes — and how they influence citation recency.
reference•8 min read
AI search glossary page spec: term, definition, anchor link, and DefinedTerm schema patterns that maximize citations from ChatGPT and Perplexity.
specification•11 min read
Reference for how 301, 302, 307, 308, meta refresh, and JavaScript redirects affect AI search citation persistence on ChatGPT, Perplexity, Gemini, and AI Overviews.
reference•10 min read
Make HTML tables AI-citable: header semantics, captions, key-value pairs, and DataTable schema patterns for high-extraction tabular content.
guide•9 min read
A copy-ready ai.txt starter template for declaring AI crawler access policies, attribution requirements, and content licensing terms.
reference•5 min read
ai.txt is an emerging root-level file that declares site-wide permissions and attribution rules for AI training, citation, and inference.
reference•6 min read
Design patterns for API reference documentation that AI agents can parse, cite, and call: canonical examples, error tables, parameter schemas, and dual human/LLM delivery.
guide•11 min read
How to design API responses, OpenAPI specs, and reference docs so LLMs and AI agents can parse, cite, and generate working code from them.
guide•7 min read
Practical guide to AI answer attribution: how engines pick a source, why credit goes to the wrong page, and the canonical, structural, and entity signals that fix it.
guide•10 min read
Specification for disclosing authenticated content to AI crawlers via schema.org isAccessibleForFree, summary endpoints, and llms.txt without leaking gated material.
specification•13 min read
BreadcrumbList schema specification: required fields, position ordering, and how AI engines use breadcrumb structured data to disambiguate citations.
specification•7 min read
Brotli vs Gzip for AI crawler responses: ratio, CPU cost, Accept-Encoding negotiation, AI bot support, and CDN configuration patterns.
comparison•9 min read
How C2PA Content Credentials cryptographically prove media provenance so AI search engines can trust authorship, edits, and AI generation history.
specification•12 min read
Reference for Cache-Control directives (max-age, s-maxage, immutable, stale-while-revalidate) that influence AI crawler refresh frequency and citation freshness.
reference•6 min read
Specification for rel=canonical implementation across HTML and HTTP-header methods, with guidance on how AI engines resolve canonicals for parameterized URLs and AMP variants.
specification•6 min read
Canonicalize duplicate and conflicting sources so AI answers cite the right URL. Practical playbook with rel=canonical, redirects, sitemap, and update policies.
guide•9 min read
CDN configuration checklist for AI crawler discoverability: bot rules, cache headers, user-agent allowlists, and edge settings that keep ChatGPT readable.
checklist•8 min read
Technical specification for making sites discoverable to ChatGPT Atlas browser: user-agent behavior, fetch semantics, robots.txt directives Atlas honors, and citation eligibility rules.
specification•10 min read
schema.org Claim type spec for AI fact-checking and grounding—appearance, firstAppearance properties and ClaimReview pairing patterns.
specification•7 min read
Specification for ClaimReview schema applied to AI trust: structure, required fields, valid values, and patterns for non-fact-check publishers.
specification•4 min read
Conditional GET and ETag handling for AI crawlers: ETag generation, If-None-Match, If-Modified-Since, 304 Not Modified, and bandwidth-saving patterns.
reference•5 min read
Use RSS, Atom, and JSON Feed to help AI crawlers and agents discover, ingest, and refresh your content faster and more reliably than crawl-only approaches.
tutorial•8 min read
Practical guide to content fingerprinting for AI citation detection — SimHash, MinHash, embedding hashes, C2PA, and DMCA workflows for publishers.
guide•8 min read
Content freshness signals for AI search: dates, last-modified headers, and updated_at metadata that move ChatGPT and Perplexity citation decisions.
guide•8 min read
What independent studies say about Core Web Vitals (LCP, INP, CLS, FCP) and AI citation rates across ChatGPT, Perplexity, and Google AI Overviews.
reference•6 min read
Configure CORS headers (Access-Control-Allow-Origin and friends) so AI search engines and embedded snippet widgets can fetch your content cleanly.
specification•8 min read
Specification for Course schema markup: Course, CourseInstance, hasPart for modules, provider, offers, and AI citation patterns for 'learn X' and 'best course for Y' queries.
specification•5 min read
Specification for CSP directives that keep AI crawlers (GPTBot, ClaudeBot, PerplexityBot) able to render and cite content without weakening XSS defense.
specification•6 min read
Schema.org Dataset JSON-LD spec for AI citations: required name/description/license, distribution with DataDownload, variableMeasured, temporalCoverage, and FAIR alignment.
specification•12 min read
Reference for using dns-prefetch and preconnect resource hints with AI crawlers and browser agents: semantics, ordering, and impact on render-stage crawls.
reference•8 min read
Edge caching spec for AI crawlers: per-bot TTL matrix, vary-on-user-agent rules, surrogate-key purges, and Cloudflare/Fastly/Akamai code samples.
specification•9 min read
Edge rendering strategy for AI citation: Cloudflare Workers vs Vercel Edge vs Netlify Edge, latency targets, cache-key strategy, and content parity rules.
guide•7 min read
Schema.org Event JSON-LD spec for AI search: required name/startDate/location, virtual and hybrid events, eventStatus, performer linkage, and AI citation patterns.
specification•10 min read
Checklist of the most common FAQ schema implementation mistakes that hurt AEO/AI-citation visibility — with the fix for each, and what changed after Google's 2023 rich-results restriction.
checklist•8 min read
Specification for FAQPage schema markup optimized for AI citations: properties, validation rules, character limits, and post-rich-result-deprecation patterns.
specification•6 min read
Specification for evaluating grounded answer quality: a rubric across factuality, attribution, and coverage, plus how to design a stable test set and score it over time.
specification•6 min read
Framework for ranking RAG grounding sources by trust, freshness, and specificity to maximize evidence quality while keeping retrieval cost in check.
framework•8 min read
Specification for negotiating gzip, deflate, and brotli compression with AI crawlers via Accept-Encoding and Content-Encoding to maximize crawl throughput.
specification•6 min read
Step-by-step guide to designing an answer grounding pipeline: source selection, evidence extraction, attribution, and guardrails to reduce hallucination measurably.
guide•13 min read
Step-by-step tutorial for creating, deploying, and validating an llms.txt file so AI systems and LLMs can discover your site's most important content.
guide•12 min read
HowTo schema specification for AI search: required and recommended fields, step markup patterns, image rules, post-deprecation usage, and validator quirks.
specification•8 min read
Hreflang for AI search ensures generative engines like ChatGPT, Perplexity, and Gemini cite the right language and regional version of your content.
guide•8 min read
Specification for hreflang annotations across HTML, sitemap, and HTTP-header methods, with guidance on AI citation behavior across query languages.
specification•5 min read
Reference for semantic HTML that AI systems read well: heading order, lists, tables, definition patterns, and the anti-patterns that cause AI to extract the wrong answer.
reference•9 min read
Use HTML5 semantic elements like article, section, nav, and proper heading hierarchy to improve AI crawler extraction and citation probability.
guide•11 min read
HTTP 103 Early Hints (RFC 8297) lets origin servers send Link preload/preconnect hints to AI crawlers before the final 200, reducing TTFB on hydrated SPA pages.
specification•7 min read
Reference for HTTP cache headers (ETag, Cache-Control, Last-Modified, Vary) and how AI crawlers use them for citation freshness.
reference•11 min read
HTTP status code reference for AI crawlers: how 2xx, 3xx, 4xx, 5xx codes affect GPTBot, ClaudeBot, PerplexityBot, and Googlebot indexing.
reference•6 min read
HTTP/3 AI crawlers support is uneven: GPTBot and most AI bots still default to HTTP/2 over TCP. Compare protocols, fallback behavior, and CDN config.
comparison•9 min read
Image sitemap specification for multimodal AI citations: image:image markup, captions, license, geo-location, and signals AI engines extract for visual search.
specification•7 min read
JavaScript SPA hydration patterns for AI crawlers: rendering modes, mismatch fixes, and framework-specific strategies for GPTBot, ClaudeBot, PerplexityBot.
guide•6 min read
JobPosting JSON-LD spec for AI search and Google for Jobs: required title, hiringOrganization, jobLocation, datePosted, validThrough, baseSalary, remote work patterns.
specification•10 min read
How to implement JSON-LD structured data so AI search engines and traditional search both understand your content's type, authorship, and entities.
guide•13 min read
A production-grade specification for validating JSON-LD structured data in CI/CD: Schema Markup Validator, Rich Results Test, error vs warning triage, and regression alerting.
specification•10 min read
JSON-LD vs Microdata vs RDFa for AI search: which structured data syntax LLM crawlers parse most reliably and how each impacts citation surfacing.
comparison•8 min read
Knowledge graph markup for AI search: a schema.org pattern specification linking entities, relationships, and citations to win generative engine trust.
specification•9 min read
Per-crawler reference for how GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, GoogleOther, and Bingbot handle native and JS-driven lazy-loaded content.
reference•5 min read
Lazy loading patterns that keep AI crawlers (GPTBot, ClaudeBot, PerplexityBot) able to extract citable content while preserving Core Web Vitals performance.
guide•7 min read
How preconnect and dns-prefetch link hints reduce AI crawler latency for asset fetches, third-party endpoints, and citation extraction.
specification•10 min read
LLM-friendly FAQ schema guide: when to use FAQPage vs QAPage vs HowTo, how to write LLM-citable answers, JSON-LD examples, and how to measure impact on AI answers.
guide•6 min read
Advanced llms.txt patterns for large knowledge bases: hierarchical sections, optional metadata, llms-full.txt pairing, sitemap integration, and validation.
guide•10 min read
Ready-to-use llms.txt starter templates for SaaS, e-commerce, blog, and docs sites — annotated, spec-aligned, and copy-paste deployable in minutes.
template•8 min read
Implementation comparison of llms.txt and ai.txt: file format, hosting, crawler behavior, validation, and a decision matrix to avoid confusion with robots.txt.
comparison•5 min read
llms.txt is a proposed root-level Markdown file that gives LLMs a curated, machine-readable index of a site. Reference for spec, format, and adoption.
reference•15 min read
LocalBusiness JSON-LD spec for AI citations: required NAP fields, openingHoursSpecification, geo, sub-types, sameAs, and AI local-intent citation patterns.
specification•10 min read
Learn how to write markdown that AI parsers can reliably read, chunk, and cite — heading rules, answer-first patterns, tables, and a quality checklist.
guide•8 min read
Meta descriptions still drive AI snippet generation in ChatGPT Search, Copilot, and Bing. Length, intent alignment, and entity-density tactics.
guide•7 min read
Side-by-side comparison of microdata vs JSON-LD for AI search: parser support, ergonomics, validation, performance, and migration recommendations.
comparison•5 min read
Per-crawler reference for desktop vs mobile fetch behavior across GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, and Googlebot Smartphone, plus parity rules.
reference•6 min read
How MobileApplication schema markup helps app publishers earn AI search citations for download, rating, OS, and install-size queries.
specification•10 min read
MonetaryAmount schema specification for AI citations: required fields, ISO 4217 currency, value vs minValue/maxValue, validFrom/validThrough, and embedding in JobPosting, Offer, MonetaryGrant.
specification•8 min read
Specification for multimodal schema markup in AI search: VideoObject, ImageObject, AudioObject, transcripts, alt text, and chapter markers tuned for retrieval.
specification•6 min read
A publisher-grade specification of the Open Graph tags AI answer engines like Perplexity, ChatGPT, and Bing Chat read to render answer cards, link previews, and citations.
specification•12 min read
Organization schema specification for AI brand citations: required fields, sameAs entity linking, logo, ContactPoint, and how LLMs verify brand identity.
specification•7 min read
Permissions-Policy header reference for AI crawlers—directive list, default policy enforcement, observed crawler render-blocking behavior, and Feature-Policy migration.
reference•7 min read
Specification for Person schema markup: sameAs to Wikidata and ORCID, jobTitle, affiliation, knowsAbout, and AI entity resolution behavior for E-E-A-T.
specification•5 min read
How to use rel=prefetch, rel=prerender, and the Speculation Rules API with AI search crawlers like OAI-SearchBot and GoogleOther — what works, what they ignore.
guide•8 min read
Product schema specification for AI shopping citations: required and recommended fields, Offer/AggregateRating, GTIN/MPN, multi-variant patterns, and validator rules.
specification•9 min read
Specification for QAPage schema markup on community Q&A sites: required properties, suggestedAnswer with author, vote counts, and AI citation behavior.
specification•5 min read
schema.org Quotation type spec for marking quoted content with author and source for AI engines—properties, nesting, and citation extraction patterns.
specification•8 min read
Fixed-size, semantic, and hybrid chunking for RAG compared: how they work, when to use each, and how to evaluate retrieval quality.
comparison•5 min read
Diagnostic checklist for the seven core RAG failure modes—retrieval miss, stale sources, hallucination, citation mismatch—with a mitigation map for each layer.
checklist•7 min read
Specification for RDFa and RDFa Lite annotations for AI search: embedding patterns, AI crawler parser support, niche use cases, and migration paths to JSON-LD.
specification•6 min read
Specification for pairing Recipe and FAQPage JSON-LD schema on a single page so AI Overviews and Perplexity cite both the recipe and its Q&A pairs.
specification•8 min read
Recipe schema specification for AI search engines: required Google fields, full JSON-LD with HowToStep and NutritionInformation, and how multimodal AI extracts recipe citations.
specification•8 min read
Specification for Retry-After and RateLimit- headers that throttle GPTBot, ClaudeBot, and PerplexityBot politely while preserving AI search citation eligibility.
specification•6 min read
Review schema spec for AI shopping citations: required Review and AggregateRating fields, anti-spam policy, and validation for ChatGPT, Perplexity, and Google.
specification•6 min read
Complete robots.txt spec for AI crawlers: GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot directives, syntax rules, and validation pipeline.
specification•11 min read
How to configure robots.txt to control AI crawlers — GPTBot, PerplexityBot, Google-Extended, ClaudeBot, Applebot-Extended, and the rest — across training and retrieval use cases.
guide•15 min read
Article schema markup checklist for AI search: 30 fields LLM crawlers consume to surface citations on ChatGPT, Perplexity, and AI Overviews.
checklist•7 min read
A three-tier framework — foundation, citation, and supporting — for prioritizing JSON-LD schema markup that earns AI search citations.
framework•8 min read
A reference of the Schema.org types and properties that matter most for AI search visibility, citations in AI Overviews, and entity recognition by LLMs.
reference•18 min read
Reference for HTTP security headers (HSTS, CSP, X-Frame-Options, Referrer-Policy) that don't block GPTBot or PerplexityBot AI crawlers.
reference•7 min read
Schema.org Service spec for AI search citations: required JSON-LD properties, areaServed, provider, serviceType, hoursAvailable, and entity-linking patterns.
specification•10 min read
Specification for service-worker behavior under AI crawler traffic: cache-first vs network-first routing, crawler-aware bypass paths, and offline-fallback contracts that preserve indexable HTML.
specification•7 min read
Sitemap-index spec for AI search: 50K URL limits, partitioning by content type, lastmod accuracy, news/image/video sub-sitemaps, and AI-crawler ping behavior.
specification•9 min read
Optimize XML sitemaps for AI crawlers: URL selection rules, exclusions, lastmod freshness signals, and how to map your sitemap to llms.txt for higher cite-rate.
guide•9 min read
SoftwareApplication schema specification for AI search: required fields, AggregateRating + Offer nesting, screenshots, and ChatGPT/Perplexity extraction patterns.
specification•6 min read
Specification for SoftwareApplication schema markup: applicationCategory, operatingSystem, offers, AggregateRating, and AI citation patterns for shopping queries like 'best CRM for'.
specification•5 min read
Speakable schema markup spec marking content for AI voice assistant extraction—property fields, CSS selector targeting, and engine support matrix.
specification•8 min read
Compare static site generators and headless CMS for AI citation readiness: build-time rendering, schema injection, content freshness, and crawler access.
comparison•8 min read
SSR patterns for AI search: full SSR vs streaming SSR vs ISR vs static prerender, framework decision matrix, and AI-crawler eligibility rules.
guide•6 min read
How to implement structured data (JSON-LD / Schema.org) to improve AI search visibility. Covers TechArticle, FAQPage, HowTo, and entity definitions.
guide•16 min read
Triage structured data warnings vs errors. Which messages from Schema.org Validator and Google Rich Results Test block AI citations and which are safe to ignore.
reference•6 min read
How Subresource Integrity (SRI) hashes, C2PA manifests, and HTTP signatures combine into a verifiable trust signal for AI search engines that cite your content.
specification•7 min read
TDMRep (W3C TDM Reservation Protocol) for AI crawlers — tdmrep.json, TDM-Reservation header, HTML meta, and CDSM Article 4 alignment.
specification•8 min read
How to write title tags that win AI citation cards in ChatGPT, Perplexity, and Google AI Overviews — length, brand placement, entity prominence, examples.
guide•7 min read
URL structure affects how AI engines parse, cite, and follow links. Best practices for slug length, depth, parameters, fragments, and stability.
guide•7 min read
Specification for the Vary HTTP header when serving AI crawlers across User-Agent, Accept, Accept-Encoding, and Accept-Language to avoid cache poisoning.
specification•6 min read
Vector embedding optimization for AI citations: how chunking, density, and semantic clarity influence retrieval in RAG-powered LLM search engines.
guide•9 min read
Writer-facing specification for vector embedding optimization: chunkable structures, anchor sentences, and metadata enrichment that survive RAG retrieval and dense vector search.
specification•11 min read
How vector embeddings power AI search retrieval, why semantic similarity drives citation odds, and what publishers can control to make content retrievable.
guide•8 min read
Video sitemap specification for AI search citations: required tags, content_loc and player_loc, thumbnails, duration, and transcript pairing patterns.
specification•7 min read
VideoObject schema specification for AI search: required and recommended fields, Clip and SeekToAction patterns, transcript embedding, and how AI engines surface video citations.
specification•10 min read
How the viewport meta tag affects mobile-first AI rendering, why misconfigurations cause silent citation losses, and the safe defaults to ship.
reference•6 min read
Web Vitals and core performance metrics for AI citation eligibility: LCP, INP, CLS thresholds plus TTFB and HTML-size budgets AI crawlers respect.
guide•6 min read
WebP vs AVIF for AI image citations: format support, compression benchmarks, and fallback patterns to ensure thumbnails render in answer cards.
comparison•9 min read
Site-level WebSite schema and SearchAction specification: required fields, JSON-LD patterns, and the signals AI search engines use for site identity.
specification•7 min read
/.well-known/ai-plugin.json manifest spec—field-by-field reference, auth options, OpenAPI integration, and ChatGPT plugin sunset migration to Custom GPTs and MCP.
specification•7 min read
A specification for a /.well-known/ai-summary endpoint that exposes a site's canonical AI summary, content inventory, and crawl preferences in a single deterministic location.
specification•8 min read
Chunking for RAG explained: how splitting documents into retrievable units shapes citation accuracy across fixed-size, recursive, semantic, and sentence-window strategies.
guide•10 min read
Context window engineering is the discipline of curating, ordering, and budgeting tokens in an LLM's context to maximize accuracy and minimize hallucinations.
guide•11 min read
Fine-tuning for search adapts foundation models—rerankers, embeddings, generators—for retrieval tasks; the canonical reference for AI search engineers.
reference•13 min read
Knowledge graph grounding ties LLM answers to structured entities and relations from sources like Google Knowledge Graph and Wikidata so facts can be verified, disambiguated, and cited.
guide•14 min read
Query fan-out in RAG: when to use multi-query retrieval, how to control cost/latency, deduplicate results, and measure impact on grounded answer quality.
reference•7 min read
Query fan-out is how AI search engines decompose a single question into many parallel sub-queries to retrieve diverse sources and synthesize a grounded answer.
guide•12 min read
RAG (retrieval-augmented generation) pairs a retriever and an LLM so answers are grounded in fresh, citable sources rather than the model's parametric memory alone.
guide•13 min read
Reranking refines retrieval results before grounding by scoring query-document pairs with a cross-encoder, sharply improving citation accuracy in RAG.
guide•12 min read
Semantic search uses meaning, not keywords, to retrieve results. Learn how vector embeddings, dense retrieval, and AI models power modern search.
reference•13 min read
A vector embedding is a fixed-length list of numbers that captures the meaning of text so similar concepts sit close together, powering semantic search and RAG.
guide•16 min read
Specification for XML sitemap fields (priority, changefreq, lastmod) in the AI search era — what is honored, what is ignored, and how to use IndexNow.
specification•5 min read