Haven Research
The Ground Truth of Modern SEO: What Search Engines Actually Document

Key findings
A systematic reading of more than 330 primary documents from Google, Microsoft, web standards bodies, and AI providers, current to September 28, 2026, finds that 17 of 32 common SEO rules are contradicted by the organizations they are attributed to, 12 more have no primary-source support, and no engine publishes a ranking weight, link value, or AI citation criterion. The documentation that does exist is consistent and specific: it names Google’s ranking systems, sets eligibility floors for indexing, rich results, and AI features, and disclaims every guarantee. This report distills that record, chapter by chapter, from The Ground-Truth Guide to Modern SEO.
- Of 32 widely repeated SEO rules tested against first-party documentation, 17 are explicitly contradicted by the engine they are attributed to, 12 have no primary-source support, 2 are narrower than claimed, 1 is context dependent, and none is confirmed as popularly phrased (research cutoff September 28, 2026).
- Google names 17 current ranking systems and 4 retired ones, calls the list “some of our more notable ranking systems,” and publishes no factor list or weights. Bing, by contrast, orders six ranking parameters “in general order of importance,” led by relevance and including user engagement.
- Every guarantee is disclaimed: following Search Essentials does not guarantee crawling, indexing, or serving; a sitemap is “merely a hint”; rel=“canonical” is “a hint, not a rule”; valid structured data “does not guarantee” a rich result; and a “URL is on Google” status does not guarantee appearance in results.
- For AI Overviews and AI Mode, Google’s only requirement is that a page be indexed and snippet-eligible; it states there are “no additional requirements” and that Google Search ignores llms.txt. No provider reviewed (Google, Microsoft, OpenAI, Anthropic, Perplexity) documents how a retrieved page is chosen for citation.
- Google’s Business Profile policies list “Rental or for-sale properties such as vacation homes” as ineligible, local ranking is documented as relevance, distance, and prominence with reviews and inbound links named, and LocalBusiness markup is not documented as a local ranking input.
Executive summary
Most SEO advice does not cite anything. It circulates as rules of thumb, correlation studies, vendor dashboards, and recommendations that have been repeated until they sound like requirements. The Ground-Truth Guide to Modern SEO, a 499-page primary-source reference written by Haven cofounder Dustin Hofer and completed in second draft with a research cutoff of September 28, 2026, asks a narrower question than the industry usually does: what can actually be established from the documentation that Google, Microsoft, Schema.org, the web standards bodies, and the AI providers publish about their own systems?
This report distills that book. It is a synthesis of a synthesis, and we say so plainly: every quotation below is a quotation the book verified against a live primary document on September 28, 2026, and every classification is the book’s. Where the book marks something Unknown, so do we.
The findings are more useful than they are flattering to the industry:
- The engines guarantee nothing, in writing. Google states that it “doesn’t guarantee that it will crawl, index, or serve your page, even if your page follows the Google Search Essentials.” The strongest outcome language it offers is that compliant sites “are more likely to show up.” Every intermediate promise (a sitemap, a canonical tag, valid structured data, a passing URL Inspection test) is disclaimed in the same way.
- Google documents ranking systems, not ranking factors. Its guide names 17 current systems and 4 retired ones and calls the list “some of our more notable ranking systems.” No weights, no formula, no count. Bing publishes an ordered list of six parameters, including user engagement, that Google has no equivalent for.
- Seventeen of 32 common rules are contradicted by their own supposed source. Title-length limits, meta-description limits, word-count minimums, duplicate-content penalties, E-E-A-T as a ranking factor, llms.txt as an AI requirement, and “blocking GPTBot removes you from ChatGPT search” are all directly contradicted by a current first-party document. Twelve more, including keyword density, domain age, bounce rate, and “schema boosts rankings,” have no primary support at all.
- AI search requires nothing new, and nobody documents how citations are chosen. Google’s eligibility rule for AI Overviews and AI Mode is one sentence: the page must be indexed and eligible to show with a snippet. Google Search ignores llms.txt. OpenAI, Anthropic, and Perplexity document which crawler to allow and that placement is not guaranteed. None of the five providers reviewed documents how a retrieved page becomes a cited one.
- For lodging operators, one local rule matters more than the rest. Google’s Business Profile policies list “Rental or for-sale properties such as vacation homes” as ineligible. Local ranking is documented as relevance, distance, and prominence; LocalBusiness structured data is not documented as a local ranking input.
The book’s closing argument is methodological rather than tactical: read the source, keep its exact strength, notice which engine said it, and treat everything else as a hypothesis. The rest of this report lays out what that method produces.
Scope and what this report is not
This is not a ranking study, an experiment, or a factor list. The book it summarizes excludes, by design, every industry study, correlation analysis, tool-vendor metric, social-media remark by a search engine employee, leaked document, and patent. Only documentation published by the organization whose system or standard is being described was admissible. Where that documentation is silent, the book writes “Unknown” rather than filling the gap, and its 48th chapter is devoted to mapping where the record stops.
Two consequences follow for readers of this report. First, “not established” is not “false”: a claim no primary source supports is unproven, not disproven, and the book calls a claim contradicted only when a primary source directly contradicts it. Second, Google’s documentation is far larger than Microsoft’s, so the evidence is Google-heavy. A Google statement settles nothing about Bing, ChatGPT search, Claude, or Perplexity, and the book keeps the engines separate on every page.
How the evidence was assembled
The book’s method is the reason its findings can be checked. Research came first and writing second.
| Element of the method | What the book reports (as of September 28, 2026) |
|---|---|
| Admissible sources | Google Search Central and Google help centers; Microsoft and Bing first-party pages and blogs; Schema.org; WHATWG, W3C, IETF, and IANA; Chrome and web.dev for performance metrics; each AI provider’s own crawler documentation |
| Excluded | Industry tools, agencies, trade publications, social media, correlation studies, leaked-document interpretations, and patents as evidence of ranking behavior |
| Research streams | 12, each reading the primary documents for its chapters and logging every cited claim alongside a verbatim excerpt |
| Claim ledger | More than 5,700 entries across more than 330 distinct primary documents |
| Adversarial review | Independent reviewers re-read every chapter for unsupported claims, secondary-source contamination, outdated guidance, recommendations dressed as requirements, and Google behavior generalized to Bing; re-fetched live sources for strong classifications, numbers, dates, and Ground Truth rows; corrected more than 300 issues |
| Second-draft pass | Chapters cut by roughly a third; every remaining quotation checked by script against the verified text |
| Source hierarchy when documents disagree | Current product documentation, then current standards, then current webmaster documentation, then current announcements, then older documentation, then statements by company representatives; where two current Google pages differ, both wordings are reported |
Every recommendation and most factual statements carry one of eleven evidence labels, applied so that a label never claims more than its source: Required; Documented Recommendation; Documented Signal (reserved for explicit statements that a ranking system uses the concept); Eligibility Requirement; Supported; Optional; Search-Engine-Specific; Standard; Not a Documented Ranking Factor; Explicitly Dismissed; and Unknown / Not Publicly Documented. The distinction the labels enforce most often is the one the industry most often collapses: a recommendation is not a requirement, and eligibility is not a guarantee.
Finding 1: Every guarantee is disclaimed
The book’s first chapter establishes the boundary that the rest of it operates inside. Google’s definition of SEO, from its own Starter Guide, is “about helping search engines understand your content, and helping users find your site and make a decision about whether they should visit your site through a search engine.” The same guide adds: “There are no secrets here that’ll automatically rank your site first in Google.”
What follows is a chain of disclaimers, each from a current Google page, each verified on September 28, 2026:
| Stage | What is commonly assumed | What Google documents |
|---|---|---|
| Compliance | Following the guidelines gets a page indexed | “Google doesn’t guarantee that it will crawl, index, or serve your page, even if your page follows the Google Search Essentials.” Meeting the three technical requirements (Googlebot not blocked, HTTP 200, indexable content) “doesn’t mean that a page will be indexed; indexing isn’t guaranteed.” |
| Discovery | A sitemap gets URLs crawled and indexed | “Submitting a sitemap is merely a hint”; “There is no guarantee that a page URL discovered in a sitemap has been or will be crawled or indexed by Google.” Google “ignores <priority> and <changefreq> values.” |
| Submission | Requesting indexing gets a page indexed faster | Repeated requests “won’t get it crawled any faster”; “Requesting a crawl does not guarantee that inclusion in search results will happen instantly or even at all”; there is a daily limit. |
| Canonical selection | rel=“canonical” tells Google which URL to index | “Indicating a canonical preference is a hint, not a rule.” Redirects and rel=“canonical” are “strong” signals; sitemaps are “weak.” |
| Indexing | All crawled pages should be indexed | “You should not expect all URLs on your site to be indexed, only the canonical pages.” |
| Verification | “URL is on Google” means the page appears in results | That status “doesn’t actually guarantee that your page will appear in Search results”; the live test cannot “predict indexing success with 100% confidence.” |
| Rich results | Valid structured data produces a rich result | “Google does not guarantee that your structured data will show up in search results, even if your page is marked up correctly.” “Using structured data enables a feature to be present, it does not guarantee that it will be present.” |
| Paid influence | Advertising helps organic rankings | “Participation in an advertising program doesn’t positively or negatively affect inclusion or ranking in the Google search results.” Bing: “does not allow websites to improve their organic ranking positions (the ‘blue links’) through payment.” |
Timing is the second trap the book flags. Google warns that “Some changes might take effect in a few hours, others could take several months,” suggests waiting “a few weeks” before judging an SEO change, and says recovery from a core-update decline “could take several months,” possibly until the next core update. A change judged after two days may simply not have been processed.
Finding 2: Systems, not factors
SEO discussion treats ranking as a scoreboard of discrete, weighted factors. Google’s documentation describes something structurally different. The unit it names is the ranking system, and its guide covers “some of our more notable ranking systems,” described as working on the page level while “site-wide signals and classifiers are also used.” The guide is explicitly partial; absence from it proves nothing, and presence on a third-party factor list proves nothing either.
As of the research cutoff, the guide names these 17 current systems: BERT; crisis information systems; deduplication systems; the exact match domain system; freshness systems; link analysis systems and PageRank; local news systems; MUM (“not currently used for general ranking in Search but rather for some specific applications”); neural matching; original content systems; removal-based demotion systems; the passage ranking system; RankBrain; reliable information systems; the reviews system; the site diversity system (“we generally won’t show more than two web page listings from the same site in our top results”); and spam detection systems.
Four systems are listed as retired, having “either been incorporated into successor systems or made part of our core ranking systems”:
| Retired system | Google’s dating | Implication |
|---|---|---|
| Helpful content system | Announced 2022; “In March 2024, it evolved and became part of our core ranking systems” | Advice treating it as a separate filter with its own recovery schedule reflects the pre-2024 description |
| Hummingbird | “a major improvement to our overall ranking systems made in August 2013” | Historical |
| Panda | Announced 2011; part of core ranking systems since 2015 | Not a separate active system |
| Penguin | “designed to combat link spam”; announced 2012; integrated 2016 | Not a separate active system |
Google uses the word “signal” sparingly, and the book reserves its Documented Signal label for those uses. The Starter Guide calls PageRank one of Google’s “many ranking signals” and the country-code TLD “usually a low impact signal.” The consumer-facing How Search Works site states that “We also use aggregated and anonymized interaction data to assess whether search results are relevant to queries,” which the book classifies as a documented signal while noting that it names no metric, no weight, and nothing measured on the publisher’s own site. Google’s page experience page states that “Core Web Vitals are used by our ranking systems.” Beyond statements of that kind, the “ranking factor” vocabulary has no footing in Google’s documentation.
Google also dismisses several popular factors by name. Its Starter Guide states that “Google Search doesn’t use the keywords meta tag,” that “The length of the content alone doesn’t matter for ranking purposes,” that keywords in a domain name “have hardly any effect beyond appearing in breadcrumbs,” that heading order does not matter to Google Search, and, asked whether E-E-A-T is a ranking factor: “No, it’s not.”
Microsoft documents Bing differently, and the difference is the single most useful search-engine divergence in the book. Instead of naming systems, Microsoft publishes “a high-level overview of the main parameters Bing uses to rank pages,” listed “in general order of importance”: relevance; quality and credibility; user engagement; freshness; location and language; and page load time. User engagement is described concretely: “Did users click through to search results for a given query, and if so, which results? Did users spend time on these search results they clicked through or quickly return to Bing?” No reviewed Google source makes an equivalent statement. Microsoft qualifies its own order, noting that relative importance “may vary from search to search and evolve over time.”
Finding 3: The crawl–index boundary is where technical SEO breaks
The book separates search into seven steps (discovery, crawl, render, index, retrieve, rank, display) because each fails in its own way, and success at one buys nothing at the next. Most technical errors it catalogs come from collapsing two adjacent steps. The following facts are all Google-documented and current to September 28, 2026.
| Mechanism | What the documentation establishes |
|---|---|
| Crawlable links | “Google can only crawl your link if it’s an <a> HTML element with an href attribute” that resolves to a URL. Click handlers without href, <span href>, and javascript: URLs are listed as not reliably crawlable. Google’s crawlers “don’t ‘click’ buttons.” |
| Sitemaps | Optional. A site of “about 500 pages or fewer” that is “comprehensively linked internally” may not need one. Limits per file: 50,000 URLs or 50MB uncompressed. Google uses <lastmod> only “if it’s consistently and verifiably” accurate; Bing calls it “a key signal” for recrawl prioritization. Both engines ignore priority and changefreq. The sitemaps ping endpoint was deprecated in June 2023. |
| IndexNow | Supported by Bing, Yandex, Seznam, Naver, Yep, Internet Archive, and AmazonBot per indexnow.org. Google is not listed and no Google page documents support. A 200 response “only indicates that the search engine has received your URL.” |
| robots.txt | Controls crawling, not indexing: “it is not a mechanism for keeping a web page out of Google.” A disallowed URL can still be indexed from links, without its content. Applies only to its own protocol, host, and port. Google caches it generally up to 24 hours. If robots.txt returns 5xx: “For the first 12 hours, Google stops crawling the site,” then uses the last good copy for 30 days. Crawl-delay and noindex in robots.txt are unsupported by Google (noindex support retired September 1, 2019). |
| noindex | Works only if Googlebot can fetch the page. A noindex on a robots.txt-blocked page is never seen. Google “may skip rendering” when it encounters noindex, so removing it with JavaScript “may not work as expected.” |
| Fetch limits | “When crawling for Google Search, Googlebot crawls the first 2MB of a supported file type, and the first 64MB of a PDF file,” uncompressed, per resource. The 15MB figure is the general crawling-infrastructure default, not the Search limit (clarified February 3, 2026). |
| Rendering | Pages with a 200 status are queued for rendering by an evergreen headless Chromium; a page “may stay on this queue for a few seconds, but it can take longer than that.” “Google Search does not interact with your page.” Dynamic rendering is “a workaround and not a recommended solution.” |
| Redirects | Googlebot follows “up to 10 redirect hops” by default. 308 is “Equivalent to 301.” Redirecting all old URLs to the homepage “might be treated as a soft 404 error.” Keep migration redirects “generally at least 1 year.” HTTP-to-HTTPS moves should not use the Change of Address tool. |
| Error codes | “The 4xx status codes, except 429, have no effect on crawl rate.” 500, 503, and 429 slow crawling but are not recommended beyond 1 to 2 days; indexed URLs behind persistent errors are eventually dropped. Unreachable indexed URLs “will be removed from Google’s index within days.” No speed difference between 404 and 410 is documented. |
| Crawl budget | The guide targets sites with 1M+ unique pages changing weekly, 10k+ changing daily, or many URLs “Discovered - currently not indexed”; others “don’t need to read this guide.” Google still requests noindexed pages, “wasting crawling time,” so noindex does not save budget. The URL Parameters tool was deprecated in 2022 and the crawl rate limiter in 2024; “You cannot request an increase in crawl rate.” |
| Faceted navigation | robots.txt and URL fragments are Google’s primary methods for keeping facets from being crawled; canonical and nofollow are “generally less effective in the long term.” “Google no longer uses” rel=next/prev. Do not canonicalize paginated pages to page 1. |
| Duplicate content | “Some duplicate content on a site is normal and it’s not a violation of Google’s spam policies.” Bing: does not trigger penalties on its own. Google “might hold pages in a duplicate cluster for up to two weeks” after a fix. |
The pattern across the table is consistent: the engines document controls and their limits, not outcomes.
Finding 4: Thirty-two claims, one verdict each
The book’s 43rd chapter takes 32 widely repeated SEO rules and gives each exactly one verdict against first-party documentation. The tally, as of September 28, 2026:
| Verdict | Count | Meaning |
|---|---|---|
| Explicitly contradicted by primary source | 17 | A current first-party document states the opposite or plainly disclaims the promised effect |
| Not supported by available primary documentation | 12 | No reviewed primary source establishes the claim; not proof it is false |
| Partially true, but commonly overstated | 2 | A primary source documents a narrower version |
| Context dependent | 1 | Primary sources describe when the claim holds and when it does not |
| Currently documented | 0 | A current primary source states the claim substantially as popularly phrased |
None of the 32 survives in its popular form. The individual verdicts, grouped as the book groups them:
| Claim | Verdict | The decisive primary-source wording |
|---|---|---|
| Google has exactly 200 ranking factors | Not supported | Relevancy “is determined by hundreds of factors”; no count is published, and “the weight applied to each factor varies depending on the nature of your query” |
| Google never uses clicks or user behavior | Contradicted | “We also use aggregated and anonymized interaction data to assess whether search results are relevant to queries” |
| Bounce rate is a direct Google ranking factor | Not supported | No Google document names it; the interaction statement concerns search results, not site analytics |
| Google Analytics data is a ranking signal | Not supported | Google’s guide to using the two tools together says nothing about Analytics feeding ranking |
| Buying Google Ads improves organic rankings | Contradicted | “Investment in paid search has no impact on your organic search ranking” |
| E-E-A-T is a ranking factor | Contradicted | “No, it’s not.” “E-E-A-T itself isn’t a specific ranking factor” |
| Quality raters’ scores change a site’s rankings | Contradicted | “These ratings do not directly impact ranking” |
| The meta keywords tag improves rankings | Contradicted | “not used by Google Search, and it has no effect on indexing and ranking at all” |
| Every page needs exactly one H1 | Not supported | “no magical, ideal amount of headings”; the HTML Standard says at least one heading “should” be level 1, not at most one |
| Title tags must be exactly 60 characters | Contradicted | “there’s no limit on how long a <title> element can be”; truncation is “typically to fit the device width” |
| Meta descriptions must be exactly 160 characters | Contradicted | “There’s no limit on how long a meta description can be” |
| Keyword density should be X percent | Not supported | No density is published; repetition “is tiring for users, and keyword stuffing is against Google’s spam policies” |
| Schema markup automatically boosts rankings | Not supported | Structured data “enables a feature to be present, it does not guarantee that it will be present”; a structured data manual action “doesn’t affect how the page ranks in Google web search” |
| Longer articles inherently rank better | Contradicted | “The length of the content alone doesn’t matter for ranking purposes”; “There’s no ideal page length” |
| Every page needs at least X words | Contradicted | Preferred word count? “(No, we don’t.)” |
| Changing the publication date improves rankings | Not supported | Freshness systems apply “for queries where it would be expected”; date changes without substantive change are listed as a search-engine-first warning sign |
| AI-generated content is automatically penalized | Context dependent | Scaled content abuse applies “no matter how it’s created”; “Appropriate use of AI or automation is not against our guidelines,” and “Using AI doesn’t give content any special gains” |
| Having more pages automatically improves SEO | Not supported | Google flags “producing lots of content on many different topics in hopes that some of it might perform well” |
| Duplicate content automatically causes a penalty | Contradicted | “It’s inefficient, but it’s not something that will cause a manual action” |
| Outbound links to authoritative sites improve rankings | Not supported | External links “can help establish trustworthiness (for example, citing your sources)”; no ranking gain for the linking page is stated |
| nofollow means Google can never follow the link | Partially true | Such links “will generally not be followed,” and targets “may still be crawled”; nofollow became a hint for crawling and indexing on March 1, 2020 |
| Sitemaps improve rankings directly | Not supported | Documented role is discovery; no ranking effect is described |
| Submitting a URL guarantees indexing | Contradicted | “Submitting a request does not guarantee that the page will appear in the Google Index” |
| Indexing guarantees ranking | Contradicted | “Google doesn’t guarantee that it will crawl, index, or serve your page” |
| A perfect PageSpeed score guarantees rankings | Contradicted | Good report results do not “guarantee that your pages will rank at the top”; “trying to get a perfect score just for SEO reasons may not be the best use of your time” |
| Core Web Vitals outweigh content relevance | Contradicted | “Google Search always seeks to show the most relevant content, even if the page experience is sub-par” |
| HTTPS gives a direct ranking boost | Not supported | 2014: “a very lightweight signal” for “fewer than 1% of global queries.” Current page: aspects beyond Core Web Vitals “don’t directly help your website rank higher” |
| Domain age guarantees authority | Not supported | No current Google or Bing document mentions domain age as a ranking input |
| A .com domain inherently outranks other TLDs | Contradicted | “Does my top-level domain (TLD) impact my site’s performance in Google Search? No.” |
| Exact-match domains automatically rank better | Partially true | Domain words are “one of many factors,” capped by a system that avoids giving “too much credit”; alone they have “hardly any effect” |
| llms.txt is required for AI search visibility | Contradicted | llms.txt “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them” |
| Blocking GPTBot removes a site from ChatGPT search | Contradicted | “Each setting is independent of the others”; OAI-SearchBot governs search, GPTBot governs training |
The book’s verdict logic bears restating. Where sources conflict, current product documentation outranks older announcements, so the 2014 HTTPS announcement does not override the current page-experience page, and the 2019 nofollow blog post does not override the current link documentation. Verdicts describe what is established today, not what was once true.
Finding 5: Content guidance is real; the measurable rules are not
Google’s content documentation is detailed about what to do and silent about how it is measured. The book treats both halves as findings.
What is documented: Google’s Search Essentials ask for “helpful, reliable, people-first content,” and its helpful-content page offers self-assessment questions it says are meant to help creators “gauge” their own work. Google “strongly encourage[s] adding accurate authorship information” where readers would expect it. “Original content systems” aim to show original reporting prominently. Trust is described as the most important component of E-E-A-T, and for topics that could affect “health, financial stability, or safety,” Google’s systems “give even more weight” to content aligning with strong E-E-A-T. Links or references from prominent sites are “one of the factors used to determine quality.”
What is not: no weights, no thresholds, no automated evaluation criteria, and no statement that author bios, bylines, word counts, or publication dates move rankings. E-E-A-T itself “isn’t a specific ranking factor,” and the “mix of factors” Google says can identify it is not listed. Quality raters’ ratings “do not directly impact ranking”; Google says they benchmark the quality of results and help improve its systems, without documenting how. On freshness, Google’s systems are query-dependent (“query deserves freshness”), and Bing defines freshness by substance: “A page that consistently provides up-to-date information is considered fresh.” Neither engine documents a benefit from altering a date without altering the content.
The AI-content position follows the same shape. The method of production is not the test; purpose and value are. Scaled content abuse covers “many pages generated for the primary purpose of manipulating search rankings and not helping users,” and Google’s generative-AI guidance says that using AI “to generate many pages without adding value for users may violate” that policy. The book records one Required item here, and it is not in Search: Google Merchant Center requires AI-generated product images to carry IPTC DigitalSourceType TrainedAlgorithmicMedia metadata. For web pages, Google says to “consider” adding information on how content was created; disclosure is described as useful, not required.
Finding 6: AI search needs nothing new, and citation selection is undocumented everywhere
The book’s AI chapters are its most consequential for anyone being sold “answer engine optimization,” because Google addressed the vocabulary directly in a guide published May 15, 2026 and last updated July 10, 2026: “From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.” Its closing priorities include “Prioritize effective SEO strategies over ‘AEO/GEO hacks’.”
Google’s eligibility rule and mythbusting list
Google’s eligibility rule for AI Overviews and AI Mode is one sentence: “To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements.” The next sentence reads “There are no additional technical requirements,” and the page’s introduction states that there are “no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” Both features are “rooted in our core Search ranking and quality systems” and “may use a ‘query fan-out’ technique,” defined as “A set of concurrent, related queries generated by the model.”
Google’s guide includes a section titled “Mythbusting generative AI search: what you don’t need to do.” Its five entries, as of July 10, 2026:
| Practice sold as AI optimization | Google’s documented position |
|---|---|
| llms.txt and other “special” files or markup | “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn’t use them.” Maintaining llms.txt for other services “will neither harm nor help” Google visibility |
| “Chunking” content into small pieces | “There’s no requirement to break your content into tiny pieces for AI to better understand it.” “There’s no ideal page length” |
| Rewriting content for AI systems | “You don’t need to write in a specific way just for generative AI search”; AI systems “can understand synonyms and general meanings” |
| Seeking inauthentic “mentions” across the web | “isn’t as helpful as it might seem,” because core ranking systems focus on quality “while other systems block spam; our generative AI features depend on both” |
| Overfocusing on structured data | “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add,” though it remains “a good idea” for rich-result eligibility |
Google’s spam policy definition was clarified on May 15, 2026 to cover “attempting to manipulate generative AI responses in Google Search.” Seeing Googlebot fetch an llms.txt file in server logs is not evidence Google uses it; Google says it “may discover, crawl, and index many kinds of files in addition to HTML on a website: this doesn’t mean that the file is treated in a special way.”
Controls
Google treats AI features as part of Search for control purposes: “robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search.” To limit what appears, the documented page-level controls are nosnippet (prevents use “as a direct input for AI Overviews and AI Mode”), max-snippet, data-nosnippet (on span, div, and section only), and noindex. Every page-level control is a trade-off across all of Search; no reviewed documentation shows a page-level control that removes a single page from AI features while leaving its ordinary snippet untouched. A property-level “Search generative AI control” in Search Console, rolled out worldwide by August 31, 2026, excludes a site from AI Overviews, AI Mode, and generative features in Discover, and “isn’t used as a ranking or inclusion signal affecting other parts of Search.”
Google-Extended is the most frequently misread control in this area. It is a robots.txt token with no separate user agent, governing whether crawled content may be used to train Gemini models and for grounding in Gemini Apps and Vertex AI. Google states that it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” The book notes that Google’s pages differ on exactly which models the token covers; all three agree it is not a Search inclusion or ranking control.
Other providers: crawler access is not citation
The other AI providers document their crawlers by purpose, and the book sorts them into four jobs: search-index crawler, user-triggered fetcher, training crawler, and control token. The distinction decides what a robots.txt rule actually does.
| Provider | Search crawler | Training crawler or token | User-triggered fetcher and its robots.txt stance | Documented limit |
|---|---|---|---|---|
| OpenAI | OAI-SearchBot (allow it and OpenAI’s published IPs to be eligible for ChatGPT search; changes take about 24 hours) | GPTBot (“Each setting is independent of the others”) | ChatGPT-User: “robots.txt rules may not apply” | “Placement is not guaranteed.” Opted-out sites “can still appear as navigational links” |
| Anthropic | Claude-SearchBot (blocking “may reduce your site’s visibility and accuracy in user search results”) | ClaudeBot | Claude-User: Anthropic states its bots honor robots.txt, with no exception for the user agent | Search provider and ranking behind the developer tool are not named |
| Perplexity | PerplexityBot (“not used to crawl content for AI foundation models”) | None documented | Perplexity-User “generally ignores robots.txt rules” | How sources are selected and ranked is not documented |
| Googlebot (AI Overviews and AI Mode draw on the same crawl) | Google-Extended (token; no effect on Search inclusion or ranking) | User-triggered fetchers are a separate category | Supporting-link selection is not documented | |
| Apple | Applebot (Spotlight, Siri, Safari; follows Googlebot rules if not named; no crawl-delay) | Applebot-Extended (token; disallowed pages “can still be included in search results”; “not considered in ranking”) | Not documented | Not documented |
| Meta | Meta-WebIndexer | Meta-ExternalAgent | Meta-ExternalFetcher “may bypass robots.txt” | No verification method on the reviewed page |
| Microsoft | Bingbot (“Many LLMs rely on data grounded in the Bing index or other search indexes”) | NOARCHIVE meta directive keeps content out of Copilot answers and model training; NOCACHE limits answers to URL, title, snippet | Not documented | Bing’s AI Performance report (public preview since February 10, 2026) counts citations and “does not indicate ranking, authority, or the role of any page” |
The common thread is the book’s most repeated AI finding: for every provider reviewed, eligibility, crawler controls, and citation formats are documented, and the criterion by which a retrieved page becomes a cited one is not. OpenAI’s own note that “the number of sources is often greater than the number of citations” confirms a selection step exists without describing it. On llms.txt specifically, no reviewed OpenAI, Anthropic, Bing, or Perplexity document states that its answer engine reads other sites’ llms.txt files; the fact that those companies publish llms.txt indexes of their own documentation says nothing about reading anyone else’s. We covered the practical side of these crawler decisions for hosts in our guide to getting a vacation rental recommended by ChatGPT.
Finding 7: Local search, and the vacation-rental exception
Two objects are easy to conflate in local search: the listing a search engine holds (a Google Business Profile) and the business’s own website, which is crawled, indexed, and ranked like any other. The listing’s rules come from Business Profile guidelines and Maps content policies, not from Search Central.
The core eligibility test, in Google’s words: “To qualify for a Business Profile on Google, a business must make in-person contact with customers during its stated hours.” Ineligible businesses include “Rental or for-sale properties such as vacation homes, model homes, or vacant apartments.” The book notes the implication for hospitality operators directly: an individual rental property is listed as ineligible, even though a management company with an office that receives customers may itself qualify. Our Insights post on who qualifies for a Business Profile walks through the rule for hosts.
| Local claim | Classification | What Google documents |
|---|---|---|
| You can pay for better local ranking | Explicitly dismissed | “There’s no way to request or pay for a better local ranking on Google” |
| Local ranking inputs | Documented signal | Local results are “mainly based on relevance, distance, and popularity”; “More reviews and positive ratings can help your business’s local ranking”; prominence is “based on info like how many websites link to your business and how many reviews you have.” No weights |
| Categories affect ranking | Documented signal | “The categories you select affect your local ranking on Google” |
| Keywords in the business name help | Required (prohibited) | Service, product, and location information are not permitted in the name; violations “could result in the suspension” of the profile |
| Replying to reviews or adding photos improves ranking | Not a documented ranking factor | Replies “can help your business stand out”; no ranking effect is stated |
| Incentivized or selectively solicited reviews are fine if disclosed | Required (prohibited) | Paid reviews, incentives, and selective solicitation of positive reviews are prohibited |
| LocalBusiness structured data improves local-pack ranking | Not a documented ranking factor | Documented as enabling a possible knowledge panel display; no ranking or Business Profile effect is stated. Self-controlled reviews on LocalBusiness or Organization pages are ineligible for star rich results |
Finding 8: Page experience is a signal, not a lever
The Core Web Vitals are Largest Contentful Paint, Interaction to Next Paint, and Cumulative Layout Shift; INP replaced First Input Delay on March 12, 2024. The thresholds web.dev and PageSpeed Insights publish, measured at the 75th percentile of page loads:
| Metric | Good | Poor |
|---|---|---|
| Largest Contentful Paint (LCP) | 2.5 seconds or less | More than 4 seconds |
| Interaction to Next Paint (INP) | 200 milliseconds or less | More than 500 milliseconds |
| Cumulative Layout Shift (CLS) | 0.1 or less | More than 0.25 |
Google’s position is stated in two sentences on one page. “Core Web Vitals are used by our ranking systems,” and “Google Search always seeks to show the most relevant content, even if the page experience is sub-par.” Good report results do not guarantee top rankings; “There is no single signal” for page experience; and aspects beyond Core Web Vitals (mobile display, interstitials, HTTPS) “don’t directly help your website rank higher in search results.” The Lighthouse performance score is lab data and is not what the ranking statement refers to. Field data requires a page or origin to be publicly discoverable and to have enough visitors to appear in the Chrome UX Report, so a new site has no field Core Web Vitals at launch.
On HTTPS the book is careful. Google announced it as “a very lightweight signal” in August 2014, said in April 2023 that “not all” former page-experience signals “may be directly used,” and today says aspects beyond Core Web Vitals do not directly help rank. Neither statement names HTTPS as dropped, so the book does not say it was; it classifies the direct-boost claim as Unknown and notes that Google does prefer HTTPS URLs as canonicals.
Finding 9: Where Google and Bing document different behavior
Because most published SEO evidence is Google’s, Bing behavior is routinely assumed from it. The book marks each divergence it found.
| Topic | Bing (Microsoft) | |
|---|---|---|
| Ranking description | 17 named systems, “some of our more notable”; no ordering | Six parameters “in general order of importance” |
| User engagement | Aggregated, anonymized interaction data “to assess whether search results are relevant”; no metric named | Explicitly asks whether users clicked, which results, and whether they stayed or “quickly return to Bing” |
| IndexNow | Not a participating engine; no documented support | Recommended “as the primary method for real-time URL submission” (November 2025 note); the older URL Submission API is “a legacy option” |
| Sitemap lastmod | Used “if it’s consistently and verifiably” accurate | “a key signal, helping Bing prioritize URLs for recrawling”; sitemaps processed “at least once every 24 hours” |
| robots.txt group handling | Specific and global (*) groups are not combined | A bingbot group causes other groups to be ignored |
| data-nosnippet | Only on span, div, and section | Any element (added October 2025) |
| noarchive / nocache | noarchive moved to a historical reference section on October 2, 2024 | NOARCHIVE excludes content from Copilot answers and model training; NOCACHE limits answers to URL, title, and snippet (September 2023) |
| Syndication canonicals | Cross-domain canonical not recommended; block indexing of the copy instead | Ask partners for a canonical “when agreements allow” |
| Link disavow | Tool exists; “most sites will not need to use this tool” | Disavow feature removed in October 2023; Bing says it can discount unnatural links itself |
| AI visibility reporting | AI Overviews and AI Mode counted in the Performance report; separate Generative AI performance report (impressions) | AI Performance report counts citations; “not a ranking system or a competitive scoreboard” |
A documented limit applies to every Bing row: Bing’s help-center pages, including the Bing Webmaster Guidelines, render client-side and could not be retrieved as text during the research, so Bing statements rest on Bing Webmaster Blog posts, Microsoft Support’s “How Bing delivers search results,” IndexNow documentation, and Merchant Center documentation. Several are old, and the book dates each one.
Finding 10: What remains unknown
The book’s closing chapter maps the space where SEO folklore lives: the questions the documentation cannot answer. Google gives its reason for withholding detail: it has “to be careful not to reveal too much detail that would allow people to game our search results.” It also reports 4,781 launches to Search in 2023, so a published formula would be stale almost immediately. As of September 28, 2026, the following are Unknown / Not Publicly Documented for Google and, where noted, Bing:
- The numeric or relative weight of any documented Google signal; the magnitude of Bing’s parameters beyond their general order; whether signals combine additively or multiplicatively.
- How Google’s named systems are combined into a ranking; whether the systems guide is exhaustive; the full set of factors.
- The value of any link; how value divides among outbound links; how much passes through redirects or canonicals; the current PageRank computation.
- Which signals identify E-E-A-T; what Google’s “site-wide assessments” measure; whether either engine computes a single site authority value. No Google or Bing document refers to the “authority” metrics sold by SEO tool vendors, which are those vendors’ constructs.
- Which interaction metrics Google uses, at what aggregation, over what windows, and how manipulation is guarded against.
- The threshold at which any quality, spam, or helpfulness classifier changes a page’s treatment; whether any ranking threshold exists for Core Web Vitals.
- What any specific core update changed, which systems it touched, or why a particular site moved. Dates and general character are published; the May 2026 core update ran from May 21 to June 2, 2026.
- For every AI provider reviewed, the criteria by which a retrieved page becomes a cited source.
Two kinds of material are sometimes offered to fill these gaps and the book declines both. Material presented in 2024 as leaked Google documentation is known only through secondary coverage, with no primary statement from Google found on its official domains, so its contents are not treated as established. Patents establish that a technique was patented, not that it is used in production, in what form, or with what weight.
What changed between 2024 and the research cutoff
Search documentation moves without announcement. The book lists the dated changes that alter practical guidance; several of them retire advice still widely sold.
| Date | Change | Source |
|---|---|---|
| 2024-03-12 | Interaction to Next Paint replaced First Input Delay as a Core Web Vital | web.dev |
| March 2024 | The helpful content system “became part of our core ranking systems” and is now listed among retired systems | |
| 2024-11-29 | Sitelinks search box documentation removed; the feature “is no longer available in Google Search results” | |
| 2025-09-09 | Google removed documentation for course info, estimated salary, learning video, special announcement, and vehicle listing structured data | |
| November to December 2025 | Google moved robots.txt, crawler, crawl budget, faceted navigation, and HTTP status documentation to developers.google.com/crawling | |
| 2026-02-03 | Google clarified that Googlebot “crawls the first 2MB of a supported file type, and the first 64MB of a PDF file” for Search | |
| 2026-02-10 | Bing launched AI Performance in Bing Webmaster Tools in public preview | Bing |
| 2026-04-13 | Google added “back button hijacking” to the malicious practices spam policy, with a corresponding manual action | |
| 2026-05-07 | “FAQ rich results are no longer appearing in Google Search”; documentation removed June 2026 | |
| 2026-05-15 | Google published “Optimizing your website for generative AI features on Google Search” and clarified that spam policies apply to generative AI responses | |
| 2026-06-15 | Google added a note that llms.txt files “aren’t needed for Google Search (and won’t negatively or positively impact your visibility or rankings)” | |
| By 2026-08-31 | Search Console’s Generative AI performance report and “Search generative AI control” rolled out to all websites worldwide |
Implications for direct-booking websites
The book was written because its author had to answer these questions in software, for every Haven customer’s site at once, and it keeps its practitioner notes separate from its evidence under a “From the Build” label. We follow the same discipline here. The following are implications we draw, not findings the engines document.
A direct-booking site lives at the crawl–index boundary more than most. The same property page can be reached with different dates in the query string, calendars generate URLs without limit, and a site with forty properties and a few hundred pages is, by Google’s own threshold, not a crawl-budget site at all. The documentation’s answer is unglamorous: crawlable <a href> links, one canonical URL per property with the strong signals (redirects and rel=“canonical”) pointing at it, robots.txt rather than noindex or canonical for parameter sprawl Google should never fetch, and an accurate sitemap lastmod. None of it is a ranking lever; all of it determines whether a page is eligible to be ranked.
For AI search, the documented work is defensive rather than additive. Allow the search crawlers (Googlebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot), make the training decision separately (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended), and do not buy an llms.txt file, a chunking pass, or AI-specific schema expecting a Google effect, because Google says there is none. Whether AI discovery then converts to supplier-direct bookings is a different question, which our report on AI, distribution, and the future of direct versus intermediated lodging booking examines with the channel-share evidence available.
For local presence, the Business Profile rule is the one to internalize: an individual vacation home is listed as ineligible. The documented local inputs (relevance, distance, prominence via reviews and inbound links) attach to an eligible business, and no structured-data type substitutes for that eligibility. Our Insights post on vacation rental SEO minus the folklore translates the book’s findings into a thirty-day sequence for hosts.
Limitations
This report inherits the book’s limitations and adds one of its own.
- It summarizes a second draft. The Ground-Truth Guide to Modern SEO is a working manuscript with a research cutoff of September 28, 2026; the book itself instructs readers to open each cited page and confirm its current wording before acting. Documentation changes without announcement, and several of the dated changes above happened within the four months before the cutoff.
- Bing evidence is thinner and older. Bing’s help-center pages could not be retrieved as text, so Bing claims rest on blog posts and Microsoft Support pages, some dating to 2019 or earlier.
- RFC 9110 wording is not quoted. Status codes are identified by RFC 9110 section numbers and IANA registry names only.
- Displayed dates may not reflect substantive change. Many Google pages share the same “last updated” date, which may reflect site-wide maintenance.
- This report is a synthesis of a synthesis. We did not independently re-fetch each primary document; we relied on the book’s claim ledger, which it reports was verified against live pages on the cutoff date and script-checked for quotation fidelity. Readers wanting the primary wording should follow the source links below.
- No ranking effects are measured here. Neither the book nor this report reports experiments, reverse-engineers ranking, or orders factors by importance, because no primary source does so for Google.
How we know
This report is a structured distillation of The Ground-Truth Guide to Modern SEO (second draft, Dustin Hofer, research cutoff and access date September 28, 2026), a 499-page reference that admits only first-party documentation as evidence: Google Search Central and Google help centers, Microsoft and Bing first-party pages and blogs, Schema.org, WHATWG, W3C, IETF and IANA, Chrome and web.dev, and each AI provider’s own crawler documentation. The book reports a claim ledger of more than 5,700 entries across more than 330 distinct primary documents, assembled by twelve research streams, checked by independent adversarial reviewers who corrected more than 300 issues, and script-verified for quotation fidelity in the second draft.
We read the full manuscript and extracted the 48 chapter-closing Ground Truth tables, which the book describes as a condensed version of itself, plus the full text of its methodology, its 32-claim myth audit (Chapter 43), its AI search chapters (35 to 37), its local search chapter (27), and its closing chapter on undocumented territory (48). Verdict counts in Finding 4 were tallied directly from the 32 individual verdicts in Chapter 43. Every quotation in this report is a quotation the book attributes to a named primary document, and we have carried over the book’s evidence classification wherever we report one. Where the book marks a matter Unknown, we have not filled it.
Excluded: any claim the book draws from a secondary source (it draws none as evidence); any statistic without a stated as-of date; experiments, correlation studies, and vendor metrics; and the book’s “From the Build” practitioner notes, except where this report explicitly labels a passage as our own implication. All dates and figures are as of September 28, 2026 unless a different date is given in the text.
Sources
- Google Search Central — Search Engine Optimization (SEO) Starter Guide (updated 2025-12-10)
- Google Search Central — In-depth guide to how Google Search works (updated 2025-12-18)
- Google Search Central — Google Search technical requirements (updated 2025-12-18)
- Google Search Central — A guide to Google Search ranking systems (updated 2025-12-10)
- Google — How Search Works: Ranking results
- Google — How Search Works: Rigorous testing
- Google Search Central — AI features and your website (updated 2025-12-10)
- Google Search Central — Optimizing for generative AI features on Google Search (updated 2026-07-10)
- Google Search Central — Google Search’s guidance on using generative AI content (updated 2025-12-10)
- Google Search Central — Spam policies for Google web search (updated 2026-08-28)
- Google Search Central — Creating helpful, reliable, people-first content (updated 2025-12-10)
- Google Search Central — Understanding page experience in Google Search results (updated 2026-09-22)
- Google Search Central — Understanding Core Web Vitals and Google search results (updated 2025-12-10)
- Google Search Central — Influencing your title links in search results (updated 2025-12-10)
- Google Search Central — Control your snippets in search results (updated 2026-04-20)
- Google Search Central — Meta tags and attributes that Google supports (updated 2025-12-10)
- Google Search Central — What is URL canonicalization (updated 2026-08-20)
- Google Search Central — How to specify a canonical URL with rel="canonical" and other methods (updated 2026-07-10)
- Google Search Central — Introduction to robots.txt (updated 2025-12-10)
- Google Crawling Infrastructure — How Google interprets the robots.txt specification (updated 2026-08-31)
- Google Search Central — Block Search indexing with noindex (updated 2025-12-10)
- Google Search Central — Googlebot (updated 2026-02-03)
- Google Crawling Infrastructure — Google’s common crawlers (updated 2026-07-14)
- Google Crawling Infrastructure — Large site owner’s guide to managing your crawl budget (updated 2026-07-22)
- Google Search Central — Understand JavaScript SEO basics (updated 2026-03-04)
- Google Search Central — SEO link best practices for Google (updated 2025-12-10)
- Google Search Central — Qualify your outbound links to Google (updated 2025-12-10)
- Google Search Central — What is a sitemap (updated 2025-12-10)
- Google Search Central — Build and submit a sitemap (updated 2026-07-08)
- Google Search Central — General structured data guidelines (updated 2026-07-10)
- Google Search Central — Google Search’s core updates and your website (updated 2025-12-10)
- Google Search Central — FAQ: Site position in Google Search (updated 2026-05-06)
- Google Search Central — Latest Google Search documentation updates
- Google Search Central Blog — HTTPS as a ranking signal (August 2014)
- Google Search Central Blog — Evolving “nofollow”: new ways to identify the nature of links (September 2019)
- Google Search Central Blog — Google Search’s guidance about AI-generated content (February 2023)
- Google Search Console Help — Page indexing report
- Google Search Console Help — URL Inspection tool
- Google Search Console Help — Search generative AI control
- Google Search Console Help — Generative AI performance report (Search)
- Google Business Profile Help — Overview of Google Business Profile policies
- Google Business Profile Help — Guidelines for representing your business on Google
- Google Business Profile Help — Tips to improve your local ranking on Google
- Microsoft Support — How Bing delivers search results (updated March 2025)
- Microsoft Bing Webmaster Blog — Keeping Content Discoverable with Sitemaps in AI Powered Search (July 2025)
- Microsoft Bing Webmaster Blog — Introducing AI Performance in Bing Webmaster Tools (Public Preview) (February 10, 2026)
- IndexNow — Documentation
- IndexNow — Participating search engines (searchengines.json)
- OpenAI — Overview of OpenAI crawlers
- Anthropic (Claude Help Center) — Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated 2026-04-07)
- Perplexity — Perplexity crawlers
- Apple Support — About Applebot (updated 2026-09-04)
- web.dev — Web Vitals
- web.dev — Interaction to Next Paint (INP) (updated 2025-09-02)
- IETF / RFC Editor — RFC 9309: Robots Exclusion Protocol (September 2022)
- WHATWG — HTML Living Standard: Sections (Headings and outlines)
- sitemaps.org — Sitemaps XML format (protocol)
- llmstxt.org — The /llms.txt file (proposal, September 3, 2024)
How to cite
Kate Swanson. (2026). The Ground Truth of Modern SEO: What Search Engines Actually Document. Haven Research. https://www.bookwithhaven.com/research/the-ground-truth-of-modern-seo-what-search-engines-actually-document
Quote with attribution. Prefer the canonical URL when citing this report.




