User-agent: * # SEO ISSUE #[YOUR_ID]: Crawl-delay directive present in robots.txt Crawl-delay: 10 Disallow: /api/ Disallow: /blocked-canonical/ Disallow: /private/ Allow: /public/Disallow Disallow: /seo-html/private/ Allow: /seo-html/public/ Allow: / # Internal search — the configuration search_results_blocked (#39) asks for. # INTERNAL_SEARCH_PATTERNS in the crawler's robots-txt-parser matches /search, # /find, /results and the /?s= and /?q= query forms, so these lines are what put # hasSearchBlocked = true and these paths into the crawl-behaviour payload. # (The check itself raises no audit issue — blocking search IS its pass condition.) Disallow: /search Disallow: /find Disallow: /results Disallow: /*?q= Disallow: /?s= # ── #324 robots_txt_misconfigured ────────────────────────────────────────── # findRobotsTxtSyntaxErrors() in the crawler's robots-txt-parser records one error per # unusable line, and validateRobotsTxtSyntax() raises #324 when that list is non-empty. # The two lines below are the whole fixture: # 1. "Nofollow" is not in KNOWN_ROBOTS_DIRECTIVES -> "is not a robots.txt directive" # 2. the second line has no ":" separator -> "is not a directive" # Both are INERT to the real parser: parseRobotsTxt() skips a colon-less line and # ignores an unknown key, in both cases WITHOUT resetting currentGroup — so every # Allow/Disallow/Crawl-delay above and below still belongs to this "User-agent: *" # group and no existing fixture (#44 crawl-delay, the /search blocking that #39 # describes, /private/, /blocked-canonical/) changes behaviour. # They must NOT make the file directive-less: #324's other branch ("no valid # directives") and #314's empty-file branch both require a file with nothing usable # in it, which this file is not. Nofollow: /legacy-archive/ Disallow /malformed-line-without-a-colon # SEO ISSUE #165: Rate-limiting directive missing (server needs rate limiting but none set) # SEO ISSUE #175: /blocked-canonical/ IS blocked above. # The /crawl-issues page has a canonical pointing to /blocked-canonical/about # which means Google cannot access the declared canonical URL. # SEO ISSUE #187: GPTBot not explicitly declared (no Allow/Disallow decision) # SEO ISSUE #188: ClaudeBot not explicitly declared (no Allow/Disallow decision) # SEO ISSUE #189: PerplexityBot not explicitly declared (no Allow/Disallow decision) # SEO ISSUE #190: Google-Extended not explicitly declared (no Allow/Disallow decision) # SEO ISSUE #191: CCBot not explicitly declared (no Allow/Disallow decision) # ── #38 sitemap_directive_declared ───────────────────────────────────────── # There is deliberately NO "Sitemap:" line in this file. See the long note at the # top: /sitemap.xml is still discovered, because it is the first conventional path # checkXmlSitemapExistence() probes, so nothing that reads the sitemap loses input.