SEO Audit

An audit is not a list of problems. It is an argument for a sequence of work, and it either survives contact with a CFO, an engineering backlog, and six months of hindsight, or it does not.

Severity is not priority. An eyeballed register almost never produces the right order.

Why Most Audits Fail

Most audits fail all three tests above. The four failures below are not stylistic complaints. Each one is the specific reason a particular audit stopped being used.

Severity Is Scored By Offence, Not Cost

Findings get ranked by how technically offensive the defect is rather than by what it costs. A rendering quirk on a template nobody visits outranks a measurement error that has been misdirecting budget for eight months.

Everything Is Recommended, So Nothing Is Prioritised

An audit that recommends everything has prioritised nothing. Without an order the client can defend, the roadmap collapses into whichever item the loudest stakeholder remembers from the readout.

Findings Carry No Derivation

A number with no stated basis collapses at the first challenge from a skeptical stakeholder. So does an observation written as an interpretation, because once the interpretation is disputed there is nothing underneath it left to defend.

It Is A Document, Not Data

An audit structured as prose cannot be sorted, filtered, assigned, or tracked. Nothing in it survives delivery, because there is nothing in it that a tracker can hold.

Three Parts, Twenty-Three Sections

Context and method come first, because a finding is only as trustworthy as the data behind it, and the reader deserves to know what that data was before being asked to act on it.

3

Parts, in reading order

23

Sections across the document

11

Diagnostic pillars, fixed question sets

40

Fields recorded per finding

Front matter

Brief And Sources

Establishes who the audit is for and what it has to achieve, then declares the data it stands on before a single finding appears.

The brief
Business model and revenue mechanics, the trigger that prompted the engagement, the measurable outcome the work is judged against, what is explicitly out of scope, and the constraints that shape what an acceptable answer looks like.
The sources
Every tool, dataset, and window used, with a confidence rating per source. Every later finding cites one of them by name, so any claim can be re-verified after delivery.
The gaps
An honest account of what was not available. An audit missing log files is a weaker audit, and the document says so upfront rather than hiding it.

Part one

The Baseline

Establishes where the site actually stands, so that nothing later in the document rests on a number nobody agreed to.

Site inventory snapshot
Dates the audit and anchors every later count. A finding that says thirty-five percent needs a denominator captured on a known day.
Historical timeline
Migrations, redesigns, incidents, and algorithm updates plotted against the traffic trend. This kills most speculation about causes before anyone starts speculating.
Search landscape
Where the site competes, for what, and against whom.
Performance baseline
Brand and non-brand separated, so brand strength cannot mask organic weakness.

Part two

The Diagnostic

Eleven pillars, each answering a fixed question set, plus the conditional modules the business model requires.

Fixed question sets
Every pillar answers the same questions on every engagement, which makes coverage provable rather than asserted.
Conditional modules
Six modules attach where the business model requires them. A single-market SaaS product does not carry an international section.
Removed, not left empty
Modules that do not apply are deleted from the document, and the sources section records which were excluded and why.

Part three

The Action Plan

Where the diagnostic becomes a sequence of work with owners, effort, and dates attached.

Findings register
Every finding as a structured record. Nothing reaches the roadmap without a finding ID, and no finding is recorded without a disposition.
Sequenced roadmap
Ordered by priority band and by dependency, in waves, with a validation gate between each.
Opportunity sizing
What the work is worth, with the derivation attached or an explicit statement that no reliable model exists.
Glossary, rubric, appendix
The scoring rubric so the order can be audited, and the appendix so the raw evidence outlives the readout.

The Eleven Diagnostic Pillars

Each pillar exists to answer one question. The question is the same on every engagement, which is what makes the answer comparable across them.

The eleven pillars and the question each one answers
PillarThe question it answers
Crawlability and indexationCan the pages that matter be reached, and are the pages that are indexed the ones that should be?
Architecture and internal linkingDoes authority reach the pages that convert, and is anything orphaned from the sections that should feed it?
Technical health and performanceDo the templates hold up in the field, on the devices the traffic actually arrives on?
Structured data and entityIs anything on the page machine-readable as an entity, and does the estate resolve to one organisation?
On-page and templatesDoes the template make each page distinguishable, or does the distinctive term sit past the truncation point?
Content and keyword strategyDoes each topic have exactly one owner, or does the site compete against itself?
AI search readinessAre the pages retrievable but not selected, and if so, what is the selected source offering that these are not?
Off-page authorityWhere does authority come from, what is genuinely at risk, and what should be left alone?
Analytics and measurementDo the numbers the client reports mean what the client thinks they mean?
What is workingWhat is driving traffic today, and what must not be touched?
Competitive benchmarkWhere does the estate actually stand against the sites it loses to?

Why the tenth pillar earns its place: a section on what is working, and specifically on what not to touch, is cheap to write and stops a new engineering team from refactoring away the one thing driving traffic.

Six Conditional Modules

Modules attach where the business model requires them, and are removed rather than included empty when it does not.

Local and multi-location International Ecommerce Migration risk UX and conversion Governance

Every Finding Is A Record

The register is the spine of the document. Nothing appears in the roadmap that does not carry a finding ID, and no finding is recorded without a disposition. Each one is a structured record across forty fields, grouped into eight sets.

Identity

An immutable ID, a headline that states the problem rather than the fix, and a pillar and category so recurring failure types can be counted across engagements.

4 fields: ID, headline, pillar, category

Scope

The level the problem lives at — portfolio, cohort, site, template, or page set — plus what it applies to, how many URLs are affected, what share of the site or of organic sessions that represents, and clickable examples. A finding without examples is an assertion.

5 fields: scope, applies to, affected count, affected share, example URLs

Evidence

Separates what was observed from what it means. The observation is stated neutrally, with no interpretation, because that is the part defended under challenge. Confidence is Confirmed when it was reproduced, Probable when the data is consistent with it, and Suspected when it is a pattern and not a proof.

4 fields: observed, source, captured, confidence

Impact

The causal chain from defect to harm, the type of harm, a quantified estimate where one is honestly possible, and the basis for it. Severity is business impact if nothing is done; cost of inaction states what worsens, and over what horizon.

6 fields: mechanism, impact type, estimate, basis, severity, cost of inaction

Remedy

Splits three things most audits merge: the recommendation, which is the decision in one sentence; the implementation, which is what a developer works from without a follow-up call; and the acceptance criteria, which verify the work was done correctly — a separate question from whether it worked.

6 fields: recommendation, implementation, acceptance criteria, effort in days, effort type, dependencies

Risk

Attaches to the recommendation rather than the finding, and captures what breaks if the fix goes wrong. This is what lets a client sequence a low-risk quick win ahead of a high-impact frightening one, which is usually the correct order.

4 fields: blast radius, failure likelihood, reversibility, mitigation

Priority

How many properties one execution covers, the leverage and risk factors derived from that, the score, the band, and the wave it lands in. Every field in this group except the wave is a formula, and none of them is ever set by hand.

6 fields: sites fixed, leverage, risk factor, priority score, priority band, wave

Lifecycle

Owner as a role rather than a name, status, and how the fix will be verified and when. Shipped and Verified are different states. When a recommendation is declined, the decision note records why, which is the field that protects everyone involved six months later.

5 fields: owner, status, decision note, verify method, verify date

The mandatory field: impact basis is required even when the estimate is null. A number without a stated derivation is a number a client can dismiss, and an admission that no reliable model exists is stronger than a figure invented to fill a cell.

One Finding, End To End

A worked example: one record from a six-property healthcare portfolio audit, expanded into the eight groups above. This is what a single row of the register holds.

Sample record P08-01

The Staging Subdomain Is Fully Indexed And Duplicating The Flagship Site

Scope
Site level, one property. 3,140 URLs, thirty-five percent of the site, with three example URLs attached.
Observed
The staging host serves 200 to all crawlers and carries no robots exclusion, no noindex, and no HTTP authentication. A site query returns 3,140 indexed staging URLs. 214 of them rank in the top twenty for queries where the production equivalent does not rank at all.
Source and confidence
A full crawl, a Search Console site query, and manual header inspection, captured on a stated date. Confirmed, meaning it was reproduced rather than inferred.
Mechanism
An exact duplicate of the flagship is competing with it. Google is choosing the staging URL as canonical for 214 queries, sending users into an environment with test data and broken forms, and splitting signals across two hostnames for everything else.
Impact and severity
Revenue. 214 queries currently served by staging, and 3,140 duplicate URLs diluting the flagship. Critical, with a disclosure exposure attached: three staging pages carry unreleased service-line names.
Recommendation
Put staging behind HTTP authentication and remove the indexed URLs.
Implementation
Add basic auth at the load balancer for the staging host, all paths. Do not rely on robots.txt alone, it will not remove what is already indexed. Submit a removal request for the host prefix. Add a pre-release check to the deploy pipeline that fails if a non-production host returns 200 without auth.
Acceptance criteria
The host returns 401 to an unauthenticated request. The site query returns zero results within thirty days. The pipeline check is present and fails correctly against a deliberate test.
Effort and risk
Half a day of developer time. Contained blast radius, low failure likelihood, instant reversibility, because auth is a load-balancer rule. Mitigation: provision QA access before the rule goes live.
Priority
Score 4.168, band P1 Quick win, wave one.
Lifecycle
Owned by platform engineering. Verified by the site query for the staging host plus a scheduled uptime check asserting 401, re-run on a dated deadline.

Priority Is Computed, Not Asserted

Severity is not priority. A Suspected Critical affecting two percent of pages should rank below a Confirmed Medium affecting sixty percent, and an eyeballed register almost never produces that result.

score = (severityWeight × reach × confidenceFactor × leverage)
        / (sqrt(effortDays) × riskFactor)

The Scoring Inputs

Six terms, each one an input rather than a constant. The model can be retuned to a client's risk appetite, and the whole register re-prioritises when it is.

The six terms in the priority score
TermWhat it does, and what it is worth
Severity weightBusiness impact if nothing is done.

Critical 8 · High 5 · Medium 3 · Low 1

ReachAffected share, floored. Measured against organic sessions wherever that represents the harm better than a URL count.

Floor 0.05

Confidence factorScales the whole score by how well the finding is proven.

Confirmed 1.00 · Probable 0.70 · Suspected 0.40

LeverageRewards a single fix that covers multiple properties.

1 + 0.15 × (sites fixed − 1)

EffortDays of work, dampened rather than ignored.

sqrt(effortDays)

Risk factorPenalises work that is dangerous and hard to undo.

Ranges 0.95 to 1.45

Risk factor: 1 + blast radius + failure likelihood − reversibility credit
ComponentWhat it asks, and what it weighs
Blast radiusWhat breaks if the fix goes wrong.

Contained 0.00 · Template 0.05 · Sitewide 0.15 · Portfolio 0.25

Failure likelihoodThe odds it goes wrong, given this team and this stack.

Low 0.00 · Medium 0.10 · High 0.20

Reversibility creditHow fast it can be undone. An instant revert justifies shipping sooner.

Instant −0.05 · Same day −0.02 · Hard 0.00

Why Effort Sits Under A Square Root

Linear effort in the denominator crushes every large initiative toward zero, and a register built that way recommends nothing but quick wins. The square root dampens effort without ignoring it: a twenty-day project is penalised about four and a half times against a one-day project, not twenty.

The Priority Bands

Bands fall out of thresholds on the score. They are computed, never set by hand, and they move when an input moves.

Band thresholds on the priority score
BandScoreWhat it means
P1 Quick win1.20 and aboveHigh return against what it costs and what it risks. Ship first.
P2 Structural0.45 and aboveReal work with real return. Sequenced behind the quick wins and any dependency.
P3 Strategic0.15 and aboveLarge, slower, often needing a pilot and a held-back control before full rollout.
P4 MonitorBelow 0.15Recorded and watched. Acting would cost more than the exposure is worth.

A Worked Register

Ten findings from the same worked example, ordered as the model orders them rather than as severity alone would.

Findings register excerpt, ordered by computed priority score
IDFindingScoreBand
P16‑01Conversions are double counted, so every performance number is inflated

Critical, confirmed · 1.0 days

8.421P1 Quick win
P08‑01The staging subdomain is fully indexed and duplicating the flagship site

Critical, confirmed · 0.5 days

4.168P1 Quick win
P11‑04No physician or organisation markup anywhere in the estate

High, confirmed · 9.0 days

1.600P1 Quick win
P12‑02Titles lead with the brand and truncate before the condition name

Medium, confirmed · 1.5 days

1.279P1 Quick win
M1‑02Forty-one location listings carry wrong hours or duplicate entries

High, confirmed · 6.0 days

0.520P2 Structural
P07‑02Provider profiles are orphaned from the specialty sites that should feed them

High, confirmed · 4.0 days

0.508P2 Structural
P10‑01Provider photos push the template past the LCP threshold on mobile

Medium, confirmed · 3.0 days

0.352P3 Strategic
P13‑03The flagship and the specialty sites compete against each other on the same queries

High, probable · 22.0 days

0.255P3 Strategic
P14‑01Clinical content is absent from AI answers on the queries that convert

High, probable · 18.0 days

0.185P3 Strategic
P15‑02A retired directory campaign left 240 low-quality links pointing at the flagship

Low, suspected · 5.0 days

0.011P4 Monitor

Read the fourth row against the eighth. A Confirmed Medium on thirty-six percent of a cohort, costing a day and a half and reverting instantly, scores 1.279 and ships in week two. A High that is only Probable, on thirty-one percent, costing twenty-two days and carrying a portfolio-wide blast radius that is hard to undo, scores 0.255 and waits for a controlled batch. Severity alone would have inverted that order, and the client would have spent a quarter on the wrong thing.

The last row is the one an audit is usually least willing to write. 240 low-quality links, seven years old, no manual action, no ranking discontinuity traceable to them: the recommendation is to document and monitor, and explicitly not to disavow. Disavowal is the risky action there, not inaction.

Where the effort actually goes
BandFindingsEffort (days)Share of effort
P1 Quick win412.017%
P2 Structural210.014%
P3 Strategic343.061%
P4 Monitor15.07%
Total1070.0100%

Multi-Site Engagements

A portfolio audit is not the same audit run six times. Above the eleven pillars sits a portfolio layer that no single-site audit needs.

  • Estate inventory. Every property, and whether it should exist at all.
  • Overlap analysis. Which query clusters the client competes with itself for.
  • Consolidation strategy. Keep, merge, retire, or rebrand, called per property.
  • Shared infrastructure. What is common across the estate, and what only looks like it is.
  • Authority distribution. Where authority actually sits across the properties.
  • Entity coherence. Whether the estate resolves to one organisation or to six strangers.

On most portfolio audits the largest single finding is that three properties should be one.

How The Diagnostic Changes

The diagnostic runs once at estate level with per-property columns, and each property receives a short annex carrying its own snapshot, findings, and roadmap slice. Site-level stakeholders read only their annex. Nobody reads eleven full pillar stacks.

Leverage

Findings pick up a scope dimension and, critically, a leverage figure expressed as one fix to N sites. A medium severity finding that ships once and repairs eleven properties outranks a high severity finding on one, and without that term the roadmap sequences wrong.

What Gets Delivered

Two artefacts, and one of them is alive.

The Audit Document

Three parts, the sections that apply, and the modules the business model requires. Sections that do not apply are removed rather than left empty, and the sources section records what was excluded and why.

The Live Findings Register

A spreadsheet with the scoring model exposed as editable inputs. Change a severity, an effort estimate, or any weight, and the bands recalculate. It is built to live in the tracker the client actually uses, not to be admired once and filed.

The Roadmap Sequences In Waves

Ordered by band and by dependency. For portfolio work, waves run pilot, then cohort, then full estate, with a validation gate between each.

Key takeaway: that wave structure is also what turns a Probable finding into a Confirmed one. A controlled batch against a held-back control is both the safe way to ship and the only way to prove the causal claim.

Commission An Audit

Single site or portfolio, with the register handed over as a live document your team keeps working after the readout. If that is the shape of what you need, the conversation starts here.