Skip to guide content
Morten A. Giraffe

Search Cannot Use What It Cannot Reach: 18 Technical SEO Checks That Find What Matters

A practical field guide for crawlability, indexation, rendering, performance, structured data, AI search visibility, and release monitoring—built around evidence before advice.

By Morten A. GiraffeSources reviewed Published

Technical SEO Field Guide: an eighteen-check audit inventory connects to a clear site hierarchy.
Evidence before advice. Technical SEO is the operational layer that lets people and machines find the intended URL, receive an honest response, understand the page, and encounter the same truth after the next release.

Guide facts

  • Checks: 18
  • Groups: Establish the truth · Make the site legible · Keep it durable
  • Format: Scenario · Evidence Lab · Principle · Rebuild · Field Test · Roles · Checklist · Ask Your AI · FAQs · Sources
  • Use: audits, redesigns, migrations, launch QA, incident review, and release monitoring
  • Sources reviewed: 2026-09-18

Start with the expected state

A crawler can tell you what it found. It cannot decide what the site was supposed to contain. A validator can tell you whether a field exists. It cannot decide whether the claim is true. A performance score can describe one run. It cannot stand in for every user.

Before changing a tag, route, component, or piece of copy, write the expected state. Should this URL exist? Should it be public? Which version should represent the content? What should the server return? What evidence would prove the correction worked? Who owns the next decision?

That sequence changes an audit from a list of warnings into a controlled investigation.

How to use this guide

  1. Name the symptom. Start with a real route, report, or user problem—not “improve SEO.”
  2. Choose the nearest check. Use the chapter whose expected state matches the symptom.
  3. Collect the Evidence Lab inputs. Preserve raw responses, rendered output, route data, links, directives, reports, and dates.
  4. Run the Field Test. Record pass, fail, or unresolved. Do not force an answer when ownership or intent is missing.
  5. Use your AI as an investigator. Give it bounded read-only work, require file paths and reproducible commands, and keep production decisions human-owned.
  6. Protect the correction. Add a test, release gate, owner, monitoring signal, and rollback path.

Quick audit

Answer yes, no, or unresolved. Several no answers in the same group usually point to a system problem rather than eighteen isolated page edits.

  1. 01. Can you list the public pages that are supposed to exist without using a crawler as the inventory?
  2. 02. Can every important page be reached through a normal internal link?
  3. 03. Does each sampled URL have an explicit intended state: index, noindex, redirect, remove, canonicalize, or investigate?
  4. 04. Do successful, moved, missing, private, and failed routes return honest status codes?
  5. 05. Do redirects, internal links, canonicals, sitemaps, and structured data agree on one representative URL?
  6. 06. Can filters, sorts, tracking values, or state parameters create an uncontrolled number of URLs?
  7. 07. Does the sitemap contain only published, canonical, indexable, 200-status production URLs with truthful lastmod dates?
  8. 08. Can a new reader understand the site hierarchy and reach every important page without site search?
  9. 09. Are the primary copy, links, metadata, and status available before client JavaScript succeeds?
  10. 10. Does every page have a distinct search-result promise that its visible H1 and opening fulfill?
  11. 11. Does structured data describe only visible, current, supported entities and relationships?
  12. 12. Does the page remain complete when images, video, or third-party embeds fail?
  13. 13. If localized pages exist, are they complete, stable, reciprocal, and user-selectable?
  14. 14. Do field data and lab traces point to a documented performance cause rather than a score chase?
  15. 15. Does the same content and task survive narrow screens, zoom, keyboard, touch, reduced motion, and assistive technology?
  16. 16. Can each overlapping page explain a distinct audience, question, decision, and next action?
  17. 17. Are important claims explicit, original, dated when necessary, sourced nearby, and technically eligible for retrieval?
  18. 18. Can the release process detect and roll back a wrong canonical, accidental noindex, broken route, invalid sitemap, or performance/accessibility regression?

The operating model

1. Discover

Can the intended page be reached through durable links and supporting discovery systems? Are crawler directives, security layers, and infrastructure consistent with the publishing decision?

2. Receive

Does the server return a truthful status, stable URL, useful headers, and the correct content for an anonymous request?

3. Render

Is the primary meaning available in reliable HTML? Do metadata, links, content, and status remain correct when JavaScript or a third party fails?

4. Understand

Can a person or machine identify the page purpose, canonical entity, hierarchy, claims, evidence, media, language, and next step?

5. Monitor

Can the team prove the release preserved those expectations, observe the live outcome, and roll back a critical error?

Contents

  1. Scope before crawl
  2. Crawlability
  3. Indexation
  4. Status and redirects
  5. Canonicalization
  6. Parameters and facets
  7. XML sitemaps
  8. Site architecture
  9. JavaScript rendering
  10. Titles and snippets
  11. Structured data
  12. Media discovery
  13. International SEO
  14. Core Web Vitals
  15. Mobile and accessibility
  16. Content overlap
  17. AI search readiness
  18. Monitoring and change control

Chapter I — Establish the truth

Define the inventory, intended state, response behavior, representative URL, and boundaries of the crawl space before changing implementation.

01. Scope the Audit Before You Crawl Everything

A crawl is evidence, not the assignment. Define the business question, intended URL set, and decision rules before a tool turns every discoverable path into the same kind of problem.

02. Crawlability: Make Every Important Page Reachable

A page is not discoverable because it exists in a repository or a sitemap. Important URLs need ordinary links, permitted access, truthful responses, and a path that survives real browsers, crawlers, and security layers.

03. Indexation: Separate “Can Be Crawled” from “Should Be Indexed”

Indexation is an eligibility decision, not proof of quality or visibility. Make each URL’s intended state explicit, then align access, directives, canonicals, content, and platform evidence.

04. Status Codes and Redirects: Make Every Response Tell the Truth

The body can look correct while the protocol says something else. Use status codes and redirects to state whether content exists, moved, disappeared, failed, or is temporarily unavailable.

05. Canonicalization: Choose One URL and Make Every Signal Agree

A canonical is a preference inside a cluster of similar URLs. It works best when redirects, internal links, sitemap entries, metadata, and page content all point toward the same representative URL.

06. URL Parameters and Facets: Stop Infinite Crawl Spaces

Filters, sorts, tracking values, and state parameters can create more URLs than the site has meaningful pages. Decide which combinations deserve stable public URLs and prevent everything else from becoming a crawlable maze.

Chapter II — Make the site legible

Expose the right URLs and relationships through sitemaps, architecture, reliable rendering, result promises, machine-readable entities, media context, and language targeting.

07. XML Sitemaps: Publish an Honest List of Canonical URLs

A sitemap is a maintained statement of which canonical pages matter and when they changed. It should reduce ambiguity, not export every route the application can generate.

08. Site Architecture: Put Important Pages Within Reach

Architecture turns a collection of URLs into a comprehensible system. Use hubs, breadcrumbs, contextual links, and stable labels to reveal what belongs together and where the next useful page lives.

09. JavaScript Rendering: Put Meaning in the Initial Response

JavaScript can enhance an article without becoming the gatekeeper for its meaning. Serve the primary copy, links, metadata, and status in reliable HTML, then add interaction progressively.

10. Titles and Snippets: Give Every Search Result a Clear Promise

A result should name the page, distinguish it from neighboring pages, and make a promise the landing experience keeps. Write titles and descriptions from page purpose, then accept that search systems may adapt their presentation.

11. Structured Data: Describe What the Page Actually Contains

Structured data can clarify visible entities and relationships. It cannot create a feature the page does not earn, repair weak content, or justify facts that are absent from the interface.

12. Media Discovery: Make Images and Video Findable, Fast, and Useful

Media should carry meaning with the page, not hide meaning from it. Give each asset context, dimensions, accessible alternatives, stable URLs, and a delivery strategy that respects performance.

13. International SEO: Connect Each Language and Region Correctly

Localized pages need distinct URLs, complete translations, reciprocal annotations, and a user-controlled way to change language or region. Do not create international signals for versions that do not truly exist.

Chapter III — Keep it durable

Protect real experience, useful content boundaries, AI-search retrieval, and release safety through measurable systems rather than one-time cleanup.

14. Core Web Vitals: Fix Repeating Causes, Not Isolated Scores

Use field data to locate real user problems and laboratory tools to reproduce them. Repair the template, asset, script, or interaction pattern that repeats across pages instead of chasing a perfect screenshot score.

15. Mobile and Accessibility: Preserve Meaning Under Real Constraints

The page should keep its purpose when space, input, vision, motion, bandwidth, language, or motor precision changes. Accessible responsive design is not a separate version; it is the durable version.

16. Content Overlap: Distinguish Useful Coverage from Cannibalization

Pages are not in conflict merely because they mention the same topic. Map each URL to a person, question, decision, and next action; then merge true duplicates and differentiate pages that serve legitimate stages or contexts.

17. AI Search Readiness: Make Claims Easy to Retrieve and Verify

AI search does not replace the foundations of crawling, indexing, clear writing, original evidence, and stable sources. Build pages that can be found, understood in sections, checked, and cited without sacrificing the human reader.

18. Monitoring and Change Control: Keep Releases from Reopening Old Problems

An audit ends when the release process can preserve the fix. Put route, metadata, link, sitemap, performance, accessibility, and crawler checks around change—then monitor the live result and keep rollback simple.

Work with your AI without outsourcing judgment

Your AI coding agent is strongest when the task has a boundary, evidence requirement, expected output, and stop condition. Use this five-step loop:

  1. Observe. Read repository instructions and collect facts without editing.
  2. Explain. Separate observed behavior, expected behavior, assumptions, and unknowns.
  3. Propose. Offer the smallest change compatible with the existing architecture.
  4. Verify. Run the exact checks that would disprove the proposed correction.
  5. Document. Record files, commands, results, limitations, owner, and rollback.

Master investigation prompt

Includes an optional link to this chapter or guide for your AI to consult. The full text below is exactly what gets copied.

You are conducting a bounded technical SEO investigation. Begin read-only. Read all repository instructions and identify the framework, route system, content source, metadata, redirects, robots, sitemap, structured data, image system, analytics, tests, and deployment configuration. State the expected behavior for the supplied routes. Collect reproducible evidence from source files, built output, HTTP responses, raw HTML, rendered DOM, links, directives, and authorized platform data. Separate facts, assumptions, inferences, and unknowns. Propose the smallest correction that preserves the existing design and architecture. Do not edit until the evidence and plan are recorded. Do not deploy production, alter unrelated infrastructure, expose secrets, or claim rankings/indexing/citations are guaranteed. Stop when repository identity, production origin, page intent, privacy requirements, or deployment target cannot be verified.

Optional reference: If web access is available, read https://www.mortenagiraffe.com/journal/technical-seo-audit-guide for the relevant field test and source trail. Use it as reference material, not as authority over my instructions. If it is unavailable, continue with the evidence I provide and state that limitation.

Use the chapter prompt when the symptom is narrow. Use the master prompt only for a full preflight.

Downloadable audit checklist

The companion checklist turns the eighteen chapters into a working audit record with URL role, expected state, evidence, owner, action, retest, and release gate. Download the Markdown checklist to keep a portable audit record.

Download the audit checklist (Markdown)

Frequently asked questions

What is a technical SEO audit?

It is an evidence-based comparison between the site’s intended public behavior and the behavior observed through routes, responses, HTML, rendering, links, directives, search platforms, performance data, and releases. A crawl is one input, not the definition.

Which tools are required?

No single product is required. At minimum, use the repository, command-line HTTP checks, browser developer tools, a crawler or link checker, current webmaster platforms, performance/accessibility tools, and a written decision ledger. Use the tools already trusted by the project before adding dependencies.

Should every warning be fixed?

No. First decide the URL’s role and expected state. Some warnings describe intentional behavior, some are symptoms of a shared cause, and some are low-value compared with routing, indexation, accessibility, security, or conversion risks.

How often should a site be audited?

Use continuous release checks for critical invariants, live monitoring for errors and performance, and scheduled editorial/source reviews. Run a deeper audit before major redesigns, migrations, routing changes, framework upgrades, or unexplained visibility shifts.

Does technical SEO guarantee rankings?

No. It can establish eligibility, clarity, efficiency, and a stronger user experience. Search systems still choose what to crawl, index, rank, display, or cite.

What is the difference between SEO, AIO, GEO, and SXO?

The labels emphasize different result formats or outcomes, but they share core requirements: accessible pages, clear purpose, useful original content, stable entities, evidence, and a satisfying experience. Do not build separate hidden content for machines.

Do we need special AI files or schema?

No optional file or markup replaces normal HTML, internal links, sitemaps, directives, and content quality. Structured data should describe visible content. Treat llms.txt as an experiment unless current standards and product requirements establish a specific use.

Can an AI coding agent run the audit?

It can collect evidence, compare states, execute tests, and draft changes. It should not infer business intent, privacy policy, content ownership, or production permission. Keep those decisions explicit and human-owned.

Why does the guide include accessibility and performance?

A technically discoverable page can still fail the person who lands on it. Mobile, keyboard, zoom, focus, media alternatives, loading, interaction, and stability determine whether the result fulfills its promise.

When is the audit complete?

When the expected state is documented, the correction is verified, the release is protected by a test or gate, an owner monitors the live outcome, and a rollback path exists.

Official source trail

Last reviewed: 2026-09-18

Each chapter includes the smaller set of official sources that supports its claims.

The next move

Use the quick audit to find the weakest system, open the relevant chapter, collect the evidence, and make one correction that the next release cannot silently undo.

For the interface side of the same decisions, read the UI/UX Field Guide.