What the website check covers
Most checks only look at the homepage. Whether a subpage from the sitemap still responds, whether a footer link leads nowhere or whether the staging setting “noindex” survived the relaunch usually shows up weeks later in Search Console or in falling visitor numbers. The website check does in one run what a crawler does first:
- robots.txt: does “Disallow: /” for all crawlers block the whole website? Which sitemap does it name?
- Sitemap: is there a readable XML sitemap at /sitemap.xml (or at the address from robots.txt)? A sitemap index is read up to three child sitemaps deep.
- Every URL via GET: status code, redirect target and response time. A HEAD request would be faster, but many servers answer it incorrectly – hence the real fetch.
- noindex: a robots meta tag or an X-Robots-Tag header on a page that is listed in the sitemap – a contradiction that search engines resolve by not indexing.
- Canonical: is the tag missing, or does it point to a different address? A trailing slash and http instead of https do not count as a mismatch.
- Internal links of the homepage: links not already in the sitemap are checked until the time budget runs out; a dead link names its source.
robots.txt and sitemap done right
A robots.txt belongs in the root folder and names the sitemap. It rarely needs more; block individual paths only if they really should not be crawled (search result pages, cart, internal areas):
- “Disallow: /” under “User-agent: *” blocks everything. After a relaunch this is the most commonly forgotten setting – in WordPress it hides behind “Discourage search engines from indexing this site”.
- The sitemap has to be valid XML (<urlset> or <sitemapindex>) and contain final addresses: with https and the correct spelling with or without www, without redirects.
- Only pages that should be indexed: no 404 pages, no noindex pages, no duplicates whose canonical points elsewhere.
User-agent: *
Disallow: /cart/
Disallow: /search
Sitemap: https://example.com/sitemap.xmlHow to read the result
The check distinguishes between errors that cost visitors and rankings and hints you should know about:
- robots.txt blocks everything: critical. Google stops adding new pages and gradually drops existing ones from the index.
- Broken URL (4xx, 5xx or no answer) in the sitemap or as an internal link: warning. Five or more broken URLs: critical.
- noindex on a sitemap URL and a canonical pointing to a different address: warning – the page is in the sitemap but will not be indexed.
- Missing sitemap, redirect in the sitemap and missing canonical: hints. Without a sitemap Google discovers new pages through links only.
- Incomplete: the time budget ran out before the end, for instance on a slow website. Homepage and sitemap always come first.
Common mistakes
- Staging settings in production: noindex or “Disallow: /” from the test environment survive the move. The damage shows weeks later.
- Sitemap with old addresses: after a change of the URL structure the old paths are still in the sitemap and answer with 404 or redirect.
- Canonical on the wrong domain: after switching from http to https or from www to the bare domain the canonical still points to the old address because the base URL in the CMS was never changed.
- Footer links into the void: a privacy or terms page was renamed, the footer link on every page was not. One dead link, a hundred times.
- Sitemap plugin disabled: after a plugin switch /sitemap.xml serves an error page with status 200 – an invalid sitemap for Google, silently ignored.
Monitor your website continuously
Dead links and wrong noindex settings do not only appear at relaunches but every time an editor deletes or renames a page. DomainWarn checks the website of every client domain daily with up to 50 sitemap URLs and 100 internal links, reports every newly broken URL as a change in the timeline and opens an incident when robots.txt blocks everything or many pages fail at once. The table of all checked URLs lives in the domain’s dashboard.
Frequently asked questions
- Why only 20 URLs?
- The free tool checks the homepage, up to 19 sitemap URLs in their order and then internal links of the homepage as long as the time budget lasts. That is enough to spot the typical mistakes. The daily check in monitoring covers up to 50 sitemap URLs and 100 links.
- Does the check cover images, PDFs and external links?
- Internal links to PDFs and other files yes, as long as the homepage links to them. External links no: checking foreign servers would be slow and says nothing about your own website. Images and scripts are not fetched.
- My website has no sitemap – is that bad?
- A hint, not an error. Google finds small websites through links as well. From a few dozen pages on a sitemap pays off, and almost every CMS generates one by itself: WordPress 5.5+ at /wp-sitemap.xml, Yoast and Rank Math at /sitemap_index.xml, Shopware and TYPO3 at /sitemap.xml.
- What does “canonical points to a different address” mean?
- The page tells search engines: “The real address is a different one.” That is correct for intentional duplicates (print view, filter parameters) and wrong when it affects every page – then the website’s base URL in the CMS is usually outdated, for example after the switch to https.