⚠️ DEV TOOL — This website intentionally contains SEO issues for testing. Not for public use.Issues Index →

Internationalised False-Positive Fixtures

The same four issue codes as the three batches before it, encoded the way an internationalised site behind a media CDN encodes them. Every one of them must come back NOT DETECTED.

Why a fourth batch

The plain batch, the edge-case batch and the CMS batch all vary the same thing: the shape of the document tree. One flat file, one index, two roots. Between them they prove the walk follows an index and unions every candidate root.

None of them varies the bytes. On all three the sitemap is LF-terminated ASCII starting with <, the llms.txt is LF-terminated ASCII starting with #, and every media URL is an unescaped same-origin path. A site generated on Windows, serving localised URLs and hosting its media on a CDN, is in none of those states — and a byte-order mark is the cheapest way there is to make a perfectly valid file unreadable to a check that anchors at the start of it.

Nothing here replaces anything. The fixtures in all three earlier batches are untouched and still assert what they always did. If any code below is reported, the defect is in the scanner, not on this site — the fixture stays as it is and the finding is recorded as a false positive.

Page-level fixture (this host)

  • Our Locations
    /fp-intl-localbusiness-locations
    #96 localbusiness_schema_location — NOT DETECTED
    Earlier batches
    /fp-localbusiness-location declares one LocalBusiness in one script; /fp-edge-localbusiness-graph declares two branches in one @graph; /fp-cms-localbusiness-branches declares two branches across two script blocks. All three put the business at or near the top of the document.
    This batch
    A store locator. The top-level node is an ItemList, and the businesses are its ListItems' items — two property hops down, behind itemListElement and item. A fourth business hangs off the Utrecht showroom as hasPOS, one hop deeper again. Two branches use the LEGACY openingHours string form (a day range on one, a comma-separated day list on another); one is a ProfessionalService, which the location check grades and the schema-missing check does not; one spells addressCountry as a country name rather than an ISO code; and the street and region names are non-ASCII.
    Why a detector could get it wrong
    A reader that graded schema['@type'] and stopped sees an ItemList and finds no business at all. One that walked only @graph and mainEntity — the properties the three earlier pages use — finds none either. A walk that stopped at the ItemList's items misses the pickup point. A validator that only understood openingHoursSpecification would call two correct hour declarations invalid, and one that required a two-letter addressCountry would report the Lisbon desk.

Site-level fixtures (the fp-intl deployment)

These three read one resource per origin, so they cannot share a hostname with the positive fixtures the primary host serves. They live on a sixth deployment of this same source — npx wrangler deploy --env fp-intl — and nothing was removed from any existing host to make room. Locally: SEO_FIXTURE_VARIANT=fp-intl npx next start -p 3982.

  • #37 image_sitemap_created — NOT DETECTED
    /sitemap.xml on the fp-intl deployment
    This batch
    A flat urlset — the shape is deliberately ordinary — served with a UTF-8 BOM and CRLF line endings, behind an <?xml-stylesheet?> processing instruction, with <xhtml:link rel="alternate" hreflang> siblings interleaved between the <image:image> children of the same <url>. Every <image:loc> is on a separate CDN hostname and carries a percent-encoded non-ASCII path and an &amp;-escaped resize query string. Every third entry declares its location and nothing else.
    Why a detector could get it wrong
    body.startsWith('<?xml') is false on a file with a BOM, and so is body[0] === '<'. A reader that tested either rejects a complete sitemap and reports #26, #37 and #41 at once. A walk that assumed the <image:image> elements of a <url> were contiguous stops at the first <xhtml:link>. A check that required the descriptive children, or that rejected an off-origin image, drops the rest.
  • #41 video_sitemap_created — NOT DETECTED
    the same document
    This batch
    One video in the embed shape: <video:player_loc> carrying allow_embed and autoplay attributes, and NO <video:content_loc>. Alongside it, <video:uploader info>, <video:restriction relationship>, <video:family_friendly>, <video:live> and an expiration date — the full entry a site whose video is on a third-party player publishes.
    Why a detector could get it wrong
    The protocol requires ONE OF content_loc and player_loc. A check reading only content_loc records the entry as a video with no address and reports a site that has declared the video it owns — which the extractor's own comment calls the majority of video sitemaps in the wild.
  • #309 llms_txt_coverage — NOT DETECTED
    /llms.txt on the fp-intl deployment
    This batch
    Complete again, with a UTF-8 BOM and CRLF endings. Four entries link the Markdown MIRROR of their page — …/head-tags.md rather than /head-tags — which is the form llmstxt.org asks for and which coveredKeys() exists to fold back onto the HTML page. Those four mirrors are served, so #308 has nothing to report either. Bullet markers vary between -, * and +; some entries are tab-indented; some carry a #fragment.
    Why a detector could get it wrong
    A BOM ahead of the H1 makes title null under an anchored heading pattern, and #307 reports missing_h1 on a correct file. A reader splitting on \n alone leaves a \r on every URL. If coveredKeys() stopped folding .md onto the HTML page, the four most spec-conformant entries in the file would each read as listing a page that does not exist while the page they mirror reads as unlisted.

← Back to the false-positive fixtures hub