⚠️ DEV TOOL — This website intentionally contains SEO issues for testing. Not for public use.Issues Index →

Edge-Case False-Positive Fixtures

0 Intentional Issues

The same five codes as the false-positive batch, written the way real sites write them. Every code on this page must come back NOT DETECTED.

Correct, and written the awkward way

/false-positive-fixtures proves the crawler stays quiet when a site does the obvious correct thing. Almost no real site is in that shape: sitemaps sit behind an index, llms.txt files use the bare-URL dialect, referrer policies come from a CDN header, and structured data arrives as one @graph from a plugin. Each of those is equally correct and each is a different path through the detector.

Nothing here replaces anything. The five fixtures in that batch are untouched and still assert what they always did; these are second cases for the same five codes. If any of them is reported, the defect is in the scanner, not on this site — the fixture stays as it is and the finding is recorded as a false positive.

Page-level fixtures (this host)

  • Referrer Policy by HTTP Header
    /fp-edge-referrer-header
    #22 meta_name_referrer — NOT DETECTED
    Plain case
    /fp-meta-referrer declares one specification token in one <meta name="referrer">.
    Edge case
    No referrer meta tag of any kind. The policy is the Referrer-Policy response header, and its value is the comma-separated fallback list the specification defines — no-referrer, strict-origin-when-cross-origin.
    Why a detector could get it wrong
    Two ways. A check looking for the tag finds none and reports a page whose policy is set the more robust way — which is the double count #22 was retired over. A validator matching the value against the token list rejects it, because a LIST is not a token; both members of this one are valid.
  • Our Locations
    /fp-edge-localbusiness-graph
    #74 localbusiness_schema_location — NOT DETECTED
    Plain case
    /fp-localbusiness-location declares one LocalBusiness, plain @type, plain nested address, in its own script.
    Edge case
    Two branches in ONE @graph: an array @type of ["LocalBusiness","ProfessionalService"], an address given as an @id pointer into the same graph, and a second branch typed with the Store subtype. Plus every prose signal — two <address> elements, two tel: links, an opening-hours table.
    Why a detector could get it wrong
    A reader that does not flatten @graph sees one untyped wrapper and no business at all. One that takes @type as a string skips the array. One that demands streetAddress on the address object reports a pointer. One keyed on the literal "LocalBusiness" misses the Store.

Site-level fixtures (the fp-edge deployment)

These three read one file per origin, so they need a host of their own — the same reason the false-positive batch has one. Deploy with npm run deploy:fp-edge and scan seo-test-site-fp-edge.<subdomain>.workers.dev. Locally: SEO_FIXTURE_VARIANT=fp-edge npx next start -p 3980.

  • /sitemap.xml → /sitemap-images.xml (fp-edge deployment)
    #37 image_sitemap_created — NOT DETECTED
    Plain case
    The false-positive host declares <image:image> inline, in the one document the crawler discovers.
    Edge case
    /sitemap.xml is a sitemap INDEX. It contains no media markup at all; the images are in a dedicated child file — the layout Google recommends for a site with a media library. Entries also vary: <image:loc> alone (title and caption were deprecated in 2022), a <image:loc> wrapped across indented lines, and several <image:image> children under one <url>.
    Why a detector could get it wrong
    A check that parsed the file it DISCOVERED, without following the index, sees zero image entries on a site that has declared every image it owns. The crawl still finds twelve images, so the check applies and is not merely skipped.
  • /sitemap.xml → /sitemap-videos.xml (fp-edge deployment)
    #41 video_sitemap_created — NOT DETECTED
    Plain case
    The false-positive host declares one complete <video:video> inline, duration included.
    Edge case
    Same index indirection, plus a second entry that is a LIVE STREAM: <video:live>yes</video:live> and no <video:duration>, because a stream has none. The extension asks for a duration on recorded video and not on live.
    Why a detector could get it wrong
    The index indirection, as above. And a check that treated <video:duration> as required would reject a correctly declared live broadcast.
  • /llms.txt (fp-edge deployment)
    #309 llms_txt_coverage — NOT DETECTED
    Plain case
    The false-positive host lists every page as an absolute Markdown link.
    Edge case
    Every page again, in four rotating valid forms: a relative target, the bare-URL list-item dialect, a www. host and an http:// scheme. Plus two URLs that must NOT be counted — one inside a fenced code block, one in prose — both pointing at paths that 404.
    Why a detector could get it wrong
    Each form is a fold canonicalUrlKey() has to perform before the comparison is right; miss one and a listed page is reported missing. And if either uncounted URL leaked into the link set it would be probed, so llms_txt_broken_links (#308) would fire instead.

One scenario that is expected to expose a gap

public/fp-edge-fixtures/media-alt-prefix.xml declares images and videos under the prefixes img: and vid:, bound to the same two namespace URIs Google's extensions define. XML Namespaces says a prefix is a local alias and carries no meaning, so those elements are the same elements image:image and video:video name — the document is valid and its media is declared.

The crawler matches on the literal qualified name ($(urlEl).find('image\\:image, image'), and cheerio in xmlMode does no namespace resolution), so #37 and #41 are expected to fire on it. That is recorded as a potential false positive in the scanner, not fixed here: renaming the prefixes would delete the test. The document is parsed on an origin of its own so a valid entry beside it cannot mask the result.

How to verify

node scripts/verify-fixtures/verify-fp-edge-cases.mjs feeds every fixture above to the real crawler modules and prints one report block per issue code. It also re-asserts the false-positive batch, because the cheapest way to make a second set of fixtures pass is to break the first.

Other hubs

/false-positive-fixtures (the plain false-positive cases), /malformed-fixtures (present and structurally invalid), /crawl-audit-fixtures (third batch) and /audit-fixtures (second batch).