Home ›
Guides › How to Read a Sitemap Audit Report
How to read a sitemap audit report
What the pages table, beacon table, status codes, and per-URL beacon rows in a sitemap audit report mean — and which findings to act on first.
What the report contains
A sitemap audit report has two tables. The pages table is the site inventory: every URL the audit enumerated, with per-page checks. The beacon table is the tracker inventory: every tracking script found, broken out by individual page URL, with identifiers extracted. Both export to CSV for deeper analysis.
Reading the pages table
- URL — the audited page. Compare this list against your sitemap: URLs in the report but missing from your sitemap are orphans the search engines may not know about.
- Status — the HTTP status. 200 is healthy; 404s are dead pages still listed in the sitemap or linked internally (remove them or redirect); 403s are pages blocking the crawler (often bot protection — verify they're reachable for real visitors and search bots).
- Title / meta description / H1 — the on-page fields that flag indexation problems at a glance. Missing or duplicated titles across many pages is the most common cheap fix.
- Canonical — where the page says its canonical URL is. A canonical pointing elsewhere means the page is telling search engines not to index it.
- Word count — thin pages (very low word counts) are candidates for consolidation.
- Source — how the URL was discovered: sitemap, crawl, or a specific subdomain. Sitemap-less discoveries deserve attention.
Reading the beacon table
- Tracker / vendor — which tracker was found (e.g. Google Tag Manager, Meta Pixel, Hotjar) and which company operates it.
- Identifier — the extracted ID (container ID, measurement ID, pixel ID). Cross-reference this against your tag manager: an ID nobody recognizes is the start of a cleanup investigation.
- Pages — every page URL where the tracker appears. A marketing pixel on checkout pages but not blog pages is fine; an analytics tag on only half the site is a measurement gap.
- Load method — direct script vs. injected via a tag manager. Rows marked "via Google Tag Manager (client-side tag)" mean the static scan saw the container, not the individual tag — the deployment path is an inference, confirmed only by a rendered scan.
What to fix first
- Dead and blocked pages — 404s in the sitemap and 403s from bot protection. These are live quality problems.
- Unrecognized trackers — unknown pixels and duplicate properties. The full triage method.
- Missing titles and meta descriptions — cheap, high-volume SEO fixes.
- Pre-consent tracker firing — re-render as a first-time visitor; trackers firing before consent are compliance findings. Consent basics.
To see what a finished report looks like before running your own, see the sample report walkthrough.
Run a free audit on your own site.
Enumerate every URL, audit each page, and scan for 40+ marketing trackers — no sign-up needed to try.
Audit my site