Grounded SEO: What an Audit Should Be Able to Prove

Run the same rental website through two SEO audit tools and you will usually get two scores, two color schemes, and two lists of urgent problems that barely overlap. One docks you for a title that runs past sixty characters. The other wants a second H1 removed. A third flags "thin content" on a page that describes a two-bedroom cottage in exactly as many words as a two-bedroom cottage needs.
None of the three tells you where its rules came from.
That silence is the problem. Search advice circulates the way folk remedies do, repeated with confidence and rarely traced to a source, and a host with a direct booking site has no easy way to sort the documented requirement from the inherited habit. So when we set out to tune the sites Haven builds for hosts, we adopted a rule for ourselves before we adopted any rule for the pages: a recommendation counts only if an engine, a standards body, or an AI crawler operator has put it in writing, and the audit has to show the passage.
That discipline is what we mean by grounded SEO. Over the past year it reshaped how our templates are built, and it eventually became a public tool, the Ground-Truth SEO Scanner, which runs the same checks on any website for free. This article is about the principles behind both.
Three tests for a recommendation
A grounded audit asks three questions of every rule before it is allowed to produce a finding.
Is it sourced? Not "everyone knows," not "a study found," but a specific document from Google, Microsoft, the sitemaps protocol, Schema.org, or the operator of a crawler. If the record is silent, the rule does not exist, however widely it is repeated. We've written before about how many popular rules fail this test when held against the engines' own words; the earlier tour of SEO folklore and our research report on what search engines actually document cover that ground in detail.
Is its strength preserved? A documented must, an engine recommendation, and a web standard are different kinds of claims. Google saying that a page returning a 404 will not be indexed is a mechanism; it either works or it does not. Google recommending a meta description on your important pages is advice. The HTML specification requiring a lang attribute is a standard with no documented search consequence. Flattening the three into one "errors and warnings" list is how a broken robots.txt ends up looking as serious as a missing favicon.
Is it repeatable? If two people run the audit on the same site on the same afternoon and get different results, at least one of them is measuring the weather.
Score the distance, not the decoration
The most common flaw in SEO scoring is subtler than a bad rule. Many audits grade only the parts of a site that exist. A page with no structured data cannot have malformed structured data; a site with no sitemap cannot have a sitemap with a stale lastmod. Measured that way, a bare template looks immaculate, and the host who has actually invested in their setup gets penalized for every slip.
A grounded score inverts the denominator. It measures how far a site stands from the best setup the engines describe, which means a documented practice counts whether or not the site has started on it. A missing meta description is a gap. A missing sitemap is a gap. The report then does the honest thing and separates the two kinds of gap: what to fix, because it is broken, and what to add, because it is absent.
The scanner expresses this with a fixed weight ratio of 8 to 4 to 2 to 0 across its evidence classes: mechanism, recommendation, standard, informational. Those numbers are not a guess at Google's ranking weights, which no engine publishes. They encode how strongly the documentation speaks, and every report prints the sentence that says so. A high score means the site has most of the documented setup in place. It is not a ranking prediction, and a tool that claims otherwise is selling you a hypothesis.
Eligibility comes before everything
Some rules are not about quality at all. They decide whether a page can be in the index in the first place: it answers 200, it is publicly reachable, the search crawlers are allowed to fetch it, and nothing on it says noindex. A grounded audit handles these as gates rather than line items, so the headline number cannot look healthy while half the property pages are blocked.
For rental sites this matters more than it sounds. A one-click "block AI" setting on a hosting platform or CDN can quietly turn away the crawlers that power AI search along with the ones that gather training data. A booking calendar can spawn hundreds of near-identical addresses, one for every date and guest count, until the engine cannot tell which is the house. A sitemap can declare pages that return redirects. Each of those is an eligibility failure, and each should be fixed before anyone touches a title tag. Our direct booking website builder checklist is a reasonable way to ask a platform whether it has handled this layer for you.
Same input, same answer
Determinism sounds like an engineering nicety until you try to use an audit in a deploy pipeline. If a score can drift because a third-party API was slow or the clock rolled over, you cannot tell a regression from noise, and you will stop trusting the number within a week.
The scanner solves this by touching the network once. It fetches robots.txt, the sitemaps, and the pages they declare, respects the crawl rules it finds, and freezes everything into a content-addressed snapshot. Every rule then runs over that snapshot with no network, no clock, and no randomness, using exact rational arithmetic so floating point never rounds a score. Re-evaluate the same snapshot with the same ruleset and the report comes back byte for byte. When two scans differ, one of three things changed: the site's responses, the rules, or the settings.
That is the property that makes a before-and-after diff meaningful. For a platform shipping many sites from shared templates, as Haven does, it is the difference between "the score went down" and "this template change removed the canonical tag on 140 property pages."
What an audit cannot see, and should say so
A grounded audit knows the edge of its own evidence. From outside a site, nobody can read Search Console, inspect server logs, or know the intent behind a redirect. Google's own documentation describes several requirements that depend on exactly that context.
Rather than guess, the scanner moves nineteen such questions to a manual review checklist with their sources attached and keeps them out of the score. The same principle governs the AI crawler report: which search bots and training bots robots.txt admits, whether llms.txt is published, what each operator documents for AI answers. All of it is reported. None of it is scored, because no provider documents how a retrieved page becomes a cited one, and Google says plainly that Search ignores llms.txt.
A tool that refuses to score what it cannot verify will produce a less dramatic report than one that invents a number for everything. We think that trade is worth making.
How it changed our templates
Working this way reshaped what a Haven site ships with by default. One canonical address per house, with date and guest variations pointing back to it. Titles that lead with the property's own name, because the most winnable search for most hosts is the one a past guest types after remembering the place. Image descriptions written by the host rather than left blank. Sitemaps that list what the host wants found and nothing else. Search crawlers allowed in; training crawlers left to the host's decision.
None of that is clever. All of it is documented, and that is the point. Clever is where the folklore lives.
Reading a report
If you run your own site through the scanner, the order of operations follows the pipeline from the top down. Clear the indexing gates first; nothing downstream matters for a page the engine cannot keep. Then fix what is broken, highest weight first. Then add what is missing, again by weight. Each finding names its rule, quotes the passage behind it, and states the fix, so whoever maintains your site can act on it without translating.
And then wait. Google's own guidance is that some changes take effect in hours and others take months. Judge a fix in weeks, not days, and resist the urge to chase the number itself. The score was never the goal. A property page that the engines can find, keep, and understand, written by the person who knows the house, is the goal, and that is a thing you can actually check.


