Crawl Audit Fixtures
24 Intentional IssuesThird batch. Each entry isolates one issue code, or a small set that can only be expressed together on one document.
Page fixtures
- Acme Analytics vs Beacon Metrics
/comparison-no-table#115 comparison_content_tablesTitle and H1 both say 'vs' / 'comparison', 300+ words, and no <table> or role=table anywhere.
- Media Library
/media-library#37 #41 #63 image_sitemap_created / video_sitemap_created / figure_figcaption_wrapTwelve <img> tags and one <video>, none declared in /sitemap.xml; plus a <figure> with no <figcaption>.
- Warehouse-native attribution, three years on
/social-article-gaps#97 #100 #102 #103 #107 og_image_width / og_image_alt / article_modified_time / article_author_meta / twitter_creator_handleog:image declared with no width, height or alt; an article with a published time but no modified time, author or twitter:creator.
- Structured Data Validation Gaps
/schema-validation-gaps#89 #76 #86 website_schema_searchaction / person_schema_author / sitelinkssearchbox_schema_missingWebSite without url, Person without name, SiteLinksSearchBox without potentialAction. Absence is never a finding — each type must be declared and wrong.
- Acme Link Checker (WebApplication + Product)
/webapp-product-schema#154 merchant listing: shippingDetails / hasMerchantReturnPolicy; product_schema_review_fieldsOne JSON-LD block typed ["WebApplication", "Product"] with an Offer and a brand, and no shippingDetails, hasMerchantReturnPolicy, review or aggregateRating.
- Acme Link Checker (WebApplication only, control)
/webapp-schema-controlcontrol none (no Product or merchant-listing finding)The same block typed WebApplication only, with no brand. Nothing about shipping, returns, reviews or ratings may be reported.
- Issue Fixtures #154–#205
/issue-fixtures#154–#205 canonical / robots / link-integrity / H1 / HSTS / compression / Lighthouse codesHub for the #154–#205 batch: one raw HTML fixture per issue code, plus the canonical targets and redirect hops they need.
- Broken External Links
/broken-external-links#95 broken_external_linkFour outbound destinations on four independent hosts, each verified to ANSWER 404. An unreachable host is recorded as unchecked, not broken.
- Mixed Active Content
/mixed-active-content#66 no_mixed_activeHTTPS page loading a script, a stylesheet and an iframe over http://. Active subresources only — an http:// image would raise #65 instead.
- Broken Web App Manifest Link
/manifest-broken#59 link_rel_manifestrel=manifest declared with no href. A page with NO manifest link is silent — the check grades configuration, not adoption.
- Canonical Over HTTP
/canonical-http#3 canonical_points_httpsThe first rel=canonical on the page is an absolute http:// URL.
- Page With No H1
/h1-missing#326 h1_heading_missingNo <h1> anywhere in the document — PageHeader, the only component that renders one, is deliberately not used.
- Legacy Documentation
/Legacy_Docs/Old_Page#13 structure_clean_lowercaseURL path carries uppercase letters and underscores. The URL must be crawled, not merely linked — the audit skips URLs it cannot resolve to a urlId.
- HTTP Scheme Audit
/http-scheme-audit#321 page_not_served_over_httpsLinks the same-domain http:// form of /http-scheme-page. pageUrl is request.url, the URL as queued, so the finding survives an edge redirect to https.
- See-other redirect (303)
/redirect-see-other#28 all_redirects_placeA single 303 hop. Valid codes are 301/302/307/308; 303 is deliberately excluded. One hop, so it does not also raise #34.
- Two-hop redirect chain
/chain-hop-1#34 no_redirect_chains/chain-hop-1 -> /chain-hop-2 -> /status-200, chainLength 2. Both hops are 301, so it does not also raise #28.
- Deep page chain (5 levels)
/depth/level-1#35 crawl_depth_shallowEach level links only the next. Five levels so the chain crosses maxDepth 3 whether the scan is seeded here or at the site root.
- Missing crawl configuration
/crawl-config-missing#38 #26 #314 sitemap_directive_declared / xml_sitemap_created / robots_txt_missing#38 is live on this host — /robots.txt declares no Sitemap: line, and /sitemap.xml is still found by the conventional-path probe. #26 and #314 contradict #37/#41 and the robots.txt fixtures respectively, so they are raised by the second deployment of this same source (wrangler deploy --env no-crawl-config), where both files answer 404.
- Character encoding not declared
/head-fixtures/charset-missing.html#312 charset_declaration_missingStatic file: no <meta charset> and no http-equiv content-type. Next.js injects a charset into every rendered page, so this cannot be a route.
- Character encoding is not UTF-8
/head-fixtures/charset-windows-1252.html#313 charset_declaration_not_utf8Static file declaring windows-1252. The other branch of the same if/else as #312, so it needs its own document.
- No viewport meta tag
/head-fixtures/viewport-missing.html#310 viewport_meta_tag_missingStatic file with no viewport tag. The guard tests the CONTENT, so an empty content attribute would land here too.
- Incomplete viewport meta tag
/head-fixtures/viewport-incomplete.html#311 viewport_meta_tag_incompleteStatic file declaring width=1024, initial-scale=0.8 — short both required directives.
- No icon link of any kind
/head-fixtures/no-favicon.html#60 link_rel_iconStatic file with no icon, shortcut icon or apple-touch-icon. The root layout gives every rendered page a favicon link.
- No main landmark
/head-fixtures/no-main-landmark.html#62 key_regions_useStatic file whose content sits in a div with role=main. The selector matches tag names, so the role does not count. The root layout wraps every page in <main>.
Site-level resources
- /llms.txt#307 #308 #309 llms_txt_malformed / llms_txt_broken_links / llms_txt_coverage
Opens with a > summary and no # title (malformed), lists two URLs that 404 (broken links), and lists four of ~50 crawled pages (coverage). Deliberately NOT linked as an anchor — the crawler fetches it at the origin root.
- /robots.txt#324 robots_txt_misconfigured
Two unusable lines inside the User-agent: * group. parseRobotsTxt skips both without ending the group, so every existing directive still applies.
Codes with no fixture, and why
Five requested codes are isActive: false and emit nothing. Two more are active but unreachable from any site content: ssl_certificate_expired needs a lapsed certificate on the crawled host, and moved_deleted_returnneeds the previous scan's URLs, which the crawler passes as an empty array. All seven are written out on /head-tag-gaps. The two that need a second origin are on /crawl-config-missing.
False-positive fixtures
/false-positive-fixtures is the inverse of this hub: conditions that resemble a finding while being correct as written, so the crawler must stay silent on them. It covers link_rel_sitemap, meta_name_author, meta_name_referrer, localbusiness_schema_location, moved_deleted_return, and — on the third deployment of this source — image_sitemap_created, video_sitemap_created and llms_txt_coverage. Every positive fixture on this page is untouched by it.
Malformed fixtures
/malformed-fixtures is the third corner of the same square. This hub removes the condition, the false-positive hub writes it correctly, and that one writes it PRESENT and STRUCTURALLY INVALID — the only case that can tell a check reading a resource's structure from one that merely notices the right element name is somewhere in the file. It covers referrer_policy_header and organization_localbusiness_schema on this host — the live successors of the retired meta_name_referrer and localbusiness_schema_location — and, on a fourth deployment of this source, image_sitemap_created, video_sitemap_created and llms_txt_coverage. Every positive fixture on this page is untouched by it.
Other hubs
/audit-fixtures (second batch) and /orphan-audit (orphan-page fixtures). Nothing on this page duplicates either.