A website can publish excellent content and still struggle in search because search engines can’t reliably discover, render, interpret or index it. That is the job of technical SEO: making sure the platform gives search engines a clear, efficient and consistent path from discovery to indexing, and gives users a fast and stable experience when they arrive.
This blueprint explains how to audit a site’s architecture, from server responses to internal linking, JavaScript rendering, canonicalization, structured data and performance. It is written for business websites, online stores, publishers and SaaS platforms that need more than a surface-level scan. For a pass/fail checklist you can work through, use our technical SEO checklist; this article explains the reasoning behind it.
The search pipeline
Google describes Search as crawling, indexing and serving results, and not every page makes it through every stage. A failure early in the pipeline makes later optimization irrelevant: a page that can’t be crawled can’t be rendered, and a page that is indexed under the wrong canonical may not receive the signals intended for it. Select a stage to see what typically goes wrong there.
What an audit must answer
A complete audit answers six business questions:
- Can search engines discover every page that should be found?
- Can they crawl those pages without wasting requests on low-value URLs?
- Can they render and understand the primary content?
- Are duplicate, redirected and canonical versions handled consistently?
- Does the site deliver a strong mobile and performance experience?
- Can the team detect technical regressions before traffic declines?
The audit isn’t finished when issues are listed. It is finished when each issue has evidence, a business impact, an owner, a priority and a practical fix.
The twelve audit phases
Expand a phase to see what to check. The phases run roughly in order, because later checks depend on earlier ones.
01Build an accurate URL inventory
Before diagnosing architecture, establish what actually exists. Sites often contain more URLs than anyone expects, from filters, parameters, tags, old campaigns, staging remnants and CMS archives. No single source is complete, so combine several: XML sitemaps, a full crawl, Search Console page indexing, Bing Webmaster Tools, analytics landing pages, server logs, CMS exports, backlink data and old migration maps.
Classify each URL by type, status, indexability, canonical target, organic value and business purpose. That classification becomes the audit’s source of truth.
| URL type | Recommended action |
|---|---|
| Valuable and unique | Keep indexable, linked, fast and in the relevant sitemap |
| Duplicate of a preferred version | Consolidate with canonical signals and consistent internal links |
| Permanently replaced | 301 or 308 redirect to the closest equivalent |
| Temporarily unavailable | Appropriate temporary status or informative response |
| Thin or operational, no search value | Consider noindex, authentication, consolidation or removal |
| Invalid URL | Return a genuine 404 or 410, not a soft 404 |
02Audit crawlability
Robots.txt controls crawler access, not indexing. Google notes that a URL blocked by robots.txt can still appear in results if other signals point to it. Check for accidentally blocked directories, blocked CSS or JavaScript needed for rendering, over-broad wildcards, staging rules copied to production and missing sitemap declarations.
Server responses should be stable and accurate. Review the spread of 200, 301/308, 302/307, 404, 410, 429 and 5xx responses. Collapse redirect chains so old URLs point straight to the final destination, and remove redirected URLs from sitemaps and navigation.
Crawl waste comes from calendar paths, faceted navigation combinations, internal search results, tracking parameters, session IDs and broken relative links that create endless URL spaces. Control it based on business value, not blanket blocking.
03Audit indexability
Crawlable doesn’t mean indexable. A page can return 200 and still be excluded by directives, canonicalization, duplication, quality evaluation or rendering problems.
Google supports noindex through a meta tag or an X-Robots-Tag header, but the crawler must be able to access the page to see it. Common mistakes: templates applying noindex to a whole section, development settings left in production, PDFs missing the intended header, and conflicting directives between HTML and headers.
Compare intended with actual indexation in four groups: intended and indexed (healthy), intended but not indexed (missed opportunity), not intended but indexed (index bloat) and not intended and not indexed (correct). Aim to maximize the share of valuable, unique pages in the index, not the total number indexed.
04Canonicalization and duplicate control
Canonicalization selects one representative URL from a group of duplicates, and Google may choose its own even when you declare a preference. The strongest setup makes every signal agree: the canonical tag, internal links, redirects, sitemap entries, protocol, hostname, trailing slashes, letter case, parameter handling and hreflang.
Link internally to canonical URLs, not redirected or duplicate versions. Online stores need particular care, because colour, size, sort, filter and tracking parameters multiply URLs. The fix is rarely one global rule.
05Site architecture and internal linking
Architecture decides how users and crawlers move through the site. Google uses links to discover pages and as a relevance signal, and recommends crawlable links with meaningful anchor text.
Check click depth from the homepage or relevant hub, orphan pages, navigation consistency, breadcrumbs, topic hubs, contextual links in body content, vague anchors, links to redirects or errors, and pagination. Important pages should be linked from relevant pages, not only from a sitemap or footer, and internal links should mirror real topical relationships and customer journeys.
06XML sitemap quality
A sitemap is a discovery hint, not a guarantee. It should list only canonical, indexable URLs that return 200, have search value and use the preferred protocol and hostname. Segmenting sitemaps by content type (products, categories, articles, locations) makes submitted-versus-indexed comparisons easier. Update lastmod only when a page meaningfully changes.
07JavaScript rendering
Google can render JavaScript, but rendering happens in a queue and adds steps, and blocked resources won’t render. Compare the raw HTML, the rendered DOM and what Search Console’s URL Inspection shows. Titles, canonicals, links, structured data and directives should be stable before and after rendering.
Avoid content that appears only after interaction, links built from click handlers instead of real anchors, near-empty initial HTML and metadata that changes after hydration. Google describes dynamic rendering as a workaround; server-side rendering or static generation is usually clearer.
08Mobile-first architecture
Google indexes and ranks using the mobile version of a page, so mobile pages must carry the same primary content, headings, metadata, internal links, images and structured data as desktop. Test real devices and network conditions, not just a narrowed desktop browser.
09Core Web Vitals and performance
Google’s current “good” thresholds, measured at the 75th percentile of real visits:
| Metric | Measures | Good |
|---|---|---|
| Largest Contentful Paint (LCP) | Loading | 2.5 seconds or less |
| Interaction to Next Paint (INP) | Responsiveness | 200 milliseconds or less |
| Cumulative Layout Shift (CLS) | Visual stability | 0.1 or less |
Use field data from real users wherever possible; lab tools help diagnose but don’t replace it. Fix problems at the template or component level, because correcting one shared template helps every page built from it.
10Structured data
Markup must match visible content, use the correct type, include required properties and pass validation. Supported markup can make pages eligible for richer results, but eligibility is not a guarantee, and Google stopped showing FAQ rich results in May 2026. Never mark up reviews or claims that aren’t on the page.
11Security and HTTPS
Verify HTTPS on every public page, no mixed content, valid certificates, secure redirects from HTTP, protected staging and admin areas, current software and no indexable sensitive data.
12Prioritization
Audits fail when they produce hundreds of issues with no order. Fix systemic problems before isolated symptoms: a template-level canonical error on 20,000 pages matters more than a missing tag on one low-traffic article. The calculator below shows one way to score issues.
An annotated issue log
This is what the output of an audit should look like. The example below is fictional, for an imaginary online store, but the structure is the one we use. Select a column heading to see why it’s there.
| ID | Issue | Evidence | Scope | Impact | Priority | Owner | Acceptance test |
|---|---|---|---|---|---|---|---|
| T-01 | Category template outputs noindex | Crawl shows meta robots noindex on /category/*; Search Console lists them as “Excluded by noindex” | Category template · 140 URLs | Main product-discovery pages missing from search | Critical | Developer | Template outputs “index, follow”; URL Inspection shows indexable on 5 sampled URLs |
| T-02 | Sort parameters indexed | site: search and crawl show ?sort= variants with self-canonicals | Listing templates · about 2,000 URLs | Crawl waste and duplicate signals | High | Developer + SEO | Sort URLs canonicalize to the clean category URL; removed from sitemap |
| T-03 | Redirect chains from old URLs | Crawl finds 3-hop chains from the 2024 migration | Legacy URLs · 380 | Slower crawling, diluted signals | Medium | Developer | Each old URL returns one 301 to its final destination |
| T-04 | Hero image slows product pages | Field data: LCP 3.8s at p75 on mobile product pages | Product template | Poor loading experience on key pages | High | Developer + design | Field LCP ≤ 2.5s at p75 after 28 days |
| T-05 | Orphaned guides | 12 guides in sitemap but no internal links in crawl | Blog · 12 URLs | Useful content rarely discovered | Medium | Content team | Each guide linked from at least one relevant hub and service page |
| T-06 | Missing sitemap declaration | robots.txt has no Sitemap line | Site-wide · 1 file | Minor discovery hint missing | Low | Developer | robots.txt lists the sitemap index URL |
Prioritizing fixes
One practical scoring method multiplies the likely search impact, the number of affected URLs, the business value of those pages and your confidence in the diagnosis, then divides by the effort to fix. Adjust the inputs to see how the score and band change. Treat the result as a conversation starter, not a verdict.
A 90-day plan
Build the URL inventory, crawl, export Search Console data, inspect logs, benchmark Core Web Vitals and talk to the development and content teams.
Fix critical crawl blocks, index directives, server errors, redirect loops, canonical conflicts and broken deployment settings.
Strengthen internal linking, navigation, sitemaps, faceted-navigation controls, rendering, mobile parity and key templates.
Add release checklists, automated tests, monitoring dashboards, performance budgets, structured-data validation and clear ownership.
A sequence, not a promise: timing depends on the site’s size, access and development capacity.
Monitoring after the audit
Technical SEO is not a one-time project. Releases, CMS updates, campaigns, migrations and third-party scripts all change the site. Monitor indexed versus submitted URLs, server-response trends, canonical mismatches, Core Web Vitals by template, bot activity in logs, redirect and 404 growth, structured-data errors and mobile parity. Alert on abnormal change, not only fixed thresholds: a sudden rise in excluded pages can matter more than a long-standing warning.
