⚠️ DEV TOOL — This website intentionally contains SEO issues for testing. Not for public use.Issues Index →

Missing Crawl Configuration

3 Intentional Issues

Two deployments of one site: the primary keeps every existing fixture, and a second host serves no robots.txt and no sitemap.xml.

#26Issue #26: xml_sitemap_created — no XML sitemap at any probed location (variant host)#314Issue #314: robots_txt_missing — /robots.txt answers 404 (variant host)#38Issue #38: sitemap_directive_declared — no Sitemap: directive in robots.txt (both hosts)

Why two hosts and not one page

All three findings are answered from new URL(seedUrl).origin, so no page on this site can change them. One origin gives one answer per code, and two of the answers this list asks for contradict fixtures the site already has — and contradict /media-library, whose #37 and #41 conditions require a sitemap that exists.

Rather than delete working fixtures, the same source is deployed a second time under a second worker name. Both are real sites to the scanner, so both detections happen on a first scan.

Detected on the primary host

  • #38 sitemap_directive_declared

    /robots.txt declares no Sitemap: directive. checkSitemapDirective() raises the finding when robotsParser.sitemapUrls is empty.

    Fires on BOTH hosts. Removing the directive costs nothing: /sitemap.xml is still found by the conventional-path probe, so the orphan diff, #40 and #37/#41 all keep the input they had.

Detected on the no-crawl-config host

npx @opennextjs/cloudflare build && npx wrangler deploy --env no-crawl-config

Deploys to seo-test-site-no-crawl-config.<subdomain>.workers.dev. Every page is identical; only /robots.txt and /sitemap.xml differ, and both answer 404. Scan that hostname as its own site.

  • #26 xml_sitemap_created

    /sitemap.xml answers 404, and none of /sitemap_index.xml, /sitemap-index.xml, /sitemap/sitemap.xml or /sitemap1.xml exist. With robots.txt also 404 there are no declared targets to try first, so checkXmlSitemapExistence() ends with sitemapFound = false.

    Cannot be combined with #37/#41, which require a sitemap to exist. That is the whole reason for the second deployment.

  • #314 robots_txt_missing

    /robots.txt answers 404. getDetectedIssues() treats 404, 5xx and an empty 200 body as the same finding.

    All three of those make fetchRobotsTxt() return an empty rawText, which would retire #44, #324, #175 and the /private/ block if it happened on the primary host.

What the variant deliberately does not change

  • Every page, every meta tag and every structured-data block is the same build, so page-level fixtures behave identically on both hosts.
  • /llms.txt is still served on the variant, so #307, #308 and #309 remain testable there too.
  • With no robots.txt on the variant, nothing is disallowed there — the orphan fixtures that depend on Disallow: /private/ are a primary-host condition and should be read from that scan.