Issue Fixtures #154–#205
50 Intentional IssuesEach page below contains the markup, header or status that one audit finding fires on. They are deliberate faults: every one of them must be DETECTED.
Indexability
- /issue-fixtures/page-nofollow-directive#154 page_nofollow_directiveverified by: crawler page check
<meta name="robots" content="index, nofollow"> on an indexable page.
apps/crawler/toggleGroups/pageSeoBasics.js — indexabilityNoindexAbsent && (metaNofollow || headerNofollow). A noindexed page is never reported, so the tag must keep "index".
- /issue-fixtures/noindex-with-external-canonical#155 noindex_with_external_canonicalverified by: crawler page check
noindex, plus a canonical that points at a different URL (the hub).
apps/crawler/toggleGroups/pageSeoBasics.js — !indexabilityNoindexAbsent && a canonical (HTML or Link header) whose normalised target differs from the page URL.
Also raises, by design: indexability_audit_critical (#21) — the noindex itself
- /issue-fixtures/robots-directives-conflict#156 robots_directives_conflictverified by: crawler page check
Meta robots "index, follow" while the X-Robots-Tag response header says "noindex".
apps/crawler/toggleGroups/pageSeoBasics.js — both routes present and robotsStance() disagrees on the index or follow dimension.
Also raises, by design: indexability_audit_critical (#21) — the header noindex wins
- /issue-fixtures/multiple-canonical-tags#157 multiple_canonical_tagsverified by: crawler page check
Two <link rel=canonical> tags in the head naming two different URLs.
apps/crawler/toggleGroups/pageSeoBasics.js — more than one DISTINCT resolved target among head canonicals (two tags naming the same URL do not count).
- /issue-fixtures/relative-canonical-url#158 relative_canonical_urlverified by: crawler page check
<link rel="canonical" href="/issue-fixtures/relative-canonical-url"> — a path with no scheme or host.
apps/crawler/toggleGroups/pageSeoBasics.js — canonicalUrl && !/^[a-z][a-z0-9+.-]*:/i.test(canonicalUrl).
- /issue-fixtures/canonical-to-error-page#159 canonical_to_error_pageverified by: API post-crawl judge
Canonical → /issue-fixtures/targets/gone, which answers 404 and is linked from the hub so the crawl fetches it.
apps/api/src/page-seo-basics-data/canonical-targets.ts judgeCanonicalTargets — fetched target with statusCode >= 400.
- /issue-fixtures/canonical-to-redirect#160 canonical_to_redirectverified by: API post-crawl judge
Canonical → /issue-fixtures/targets/moved, which answers 301.
canonical-targets.ts judgeCanonicalTargets — fetched target with a 3xx status or a recorded redirect chain to another URL.
- /issue-fixtures/canonical-to-noindex-page#161 canonical_to_noindex_pageverified by: API post-crawl judge
Canonical → /issue-fixtures/targets/noindex, a 200 page carrying noindex.
canonical-targets.ts judgeCanonicalTargets — fetched 2xx target whose noindexAbsent === false.
- /issue-fixtures/canonical-chain#162 canonical_chainverified by: API post-crawl judge
Canonical → /issue-fixtures/targets/chain-middle, whose own canonical points on to /issue-fixtures/targets/chain-final.
canonical-targets.ts judgeCanonicalTargets — the target's own canonical resolves to a third URL.
- /issue-fixtures/paginated-canonical-to-first-page/page/2#163 paginated_canonical_to_first_pageverified by: crawler page check
Page 2 at …/page/2 with rel=prev to page 1 and a canonical to page 1.
apps/crawler/toggleGroups/pageSeoBasics.js — URL page number > 1 (PAGE_NUMBER_PATH /page/N) and the canonical is page 1 of the same series.
- /issue-fixtures/canonical-outside-head#201 canonical_outside_headverified by: crawler page check
The only canonical <link> is written inside <body>.
apps/crawler/toggleGroups/pageSeoBasics.js — any canonical tag for which isInDocumentHead() is false.
Also raises, by design: canonical_all — the head declares no canonical
- /issue-fixtures/link-header-canonical-conflict#202 link_header_canonical_conflictverified by: crawler page check
HTML canonical = this page; HTTP Link header canonical = the hub.
apps/crawler/toggleGroups/pageSeoBasics.js — head canonical and Link-header canonical both present and resolving to different URLs.
Crawlability
- /issue-fixtures/hashbang-url-link#200 hashbang_url_linkverified by: crawler page check
Internal links written as "#!/reports" and "/issue-fixtures#!/settings".
apps/crawler/utils/links.js detectLinkIssues() — any internal href containing "#!".
- /issue-fixtures/css-js-blocked-by-robots#164 css_js_blocked_by_robotsverified by: API post-crawl judge
A stylesheet and a script served from /private/fixture-assets/ — /private/ is Disallowed for User-agent: * in robots.txt.
apps/api/src/site-crawl-behaviour/site-crawl-behaviour.service.ts checkRobotsBlockedResources — a same-host css/js page_resources row matched by a Disallow of the googlebot (else *) group. Site-level: one finding per scan.
- /issue-fixtures/page-failed-to-load#193 page_failed_to_loadverified by: HTTP response
On the primary host the response is held for 60 s — past the crawler's 45 s navigation timeout, on all 3 attempts. Other deployments answer 404 at once.
apps/crawler/utils/fetchError.js buildPageFailedToLoadIssue — a transport failure (here class 'timeout') that survived every retry. 4xx/5xx never reach it.
- /issue-fixtures/internal-link-to-robots-blocked-url#194 internal_link_to_robots_blocked_urlverified by: crawler page check
A followed internal link to /private/blocked-page.html (Disallow: /private/).
apps/crawler/robots-txt-parser.js detectInternalLinksToRobotsBlocked — an internal link without rel=nofollow whose target matchRule(url, 'googlebot') disallows.
- /issue-fixtures/sitemap-lists-robots-blocked-url#195 sitemap_lists_robots_blocked_urlverified by: crawler page check
/sitemap.xml on the primary host lists <loc>…/private/blocked-page.html</loc>, which robots.txt disallows. This page documents it.
apps/crawler/toggleGroups/crawlBehaviour.js detectSitemapUrlsBlockedByRobots — a sitemap <loc> for which robots.matchRule(loc, 'googlebot').allowed === false. Site-level.
Also raises, by design: the sitemap-vs-crawl orphan diff — the listed URL is never crawled, because robots.txt blocks it
Link Integrity
- /issue-fixtures/internal-server-error#165 internal_server_error_pageverified by: crawler page check
The page answers HTTP 500 with an HTML error body.
apps/crawler/utils/brokenPageIssue.js — integer status 500–599 (reported on the 5xx page itself).
- /issue-fixtures/redirect-loop#166 redirect_loopverified by: HTTP response
301 → /issue-fixtures/redirect-loop-b, which 301s straight back here.
apps/crawler/utils/redirects.js analyzeRedirectChain (hasLoop) / fetchError.js too_many_redirects → apps/api redirect-chain.service.ts.
- /issue-fixtures/dead-end-page#167 dead_end_pageverified by: API post-crawl judge
A 200 text/html page with no <a> element at all.
apps/api/src/page-internal-link/page-internal-link.service.ts — COMPLETED 2xx HTML page with no internal_links row to another URL.
- /issue-fixtures/link-to-local-or-staging-host#168 link_to_local_or_staging_hostverified by: API post-crawl judge
Links to http://localhost:3000, http://192.168.1.20, http://127.0.0.1:8080 and https://staging.acme-analytics.com.
apps/api/src/page-external-link/link-host-classification.ts classifyLinkHost — loopback / private-ip / staging-host.
- /issue-fixtures/malformed-link-href#196 malformed_link_hrefverified by: crawler page check
hrefs "http://acme analytics.com/" (space in the host), "https://docs.acme-analytics.com:99999/" (port out of range) and "http://[2001:db8::1/status" (unclosed IPv6 bracket).
apps/crawler/utils/links.js detectLinkIssues() — an href new URL() rejects (the catch branch), or a scheme typo.
- /issue-fixtures/uncrawlable-link#197 uncrawlable_linkverified by: crawler page check
An <a href="javascript:void(0)"> and an <a onclick> with no href.
apps/crawler/utils/links.js detectLinkIssues() — javascript: href, or onclick with a missing/empty/"#" href.
- /issue-fixtures/empty-anchor-text#198 empty_anchor_textverified by: crawler page check
An internal <a> containing only an empty icon <span> — no text, no image, no aria-label.
apps/crawler/utils/links.js detectLinkIssues() — internal link of kind 'empty' with no accessible name.
- /issue-fixtures/image-link-missing-alt#199 image_link_missing_altverified by: crawler page check
An internal <a> whose only content is an <img> with no alt attribute.
apps/crawler/utils/links.js detectLinkIssues() — internal link of kind 'image' with no alt text and no accessible name.
- /issue-fixtures/external-link-redirects#203 external_link_redirectsverified by: API post-crawl judge
External links to http://github.com/ (301 → https) and https://google.com/ (301 → www.google.com).
apps/api/src/page-external-link/page-external-link.service.ts raiseExternalRedirects — probe followed ≥ 1 redirect and ended 2xx on a different URL.
- /issue-fixtures/broken-fragment-link#205 broken_fragment_linkverified by: API post-crawl judge
Links to "#pricing-table" (no such id here) and to "/issue-fixtures/duplicate-h1-a#feature-matrix" (no such id there).
apps/api/src/page-internal-link/fragment-link.service.ts — fragment not in the target's recorded element ids (target crawled, 2xx, ids not truncated).
On-Page
- /issue-fixtures/multiple-h1-headings#175 multiple_h1_headingsverified by: crawler page check
Two <h1> elements, both with text: a logo H1 and the page title H1.
apps/crawler/toggleGroups/pageSeoBasics.js — $('h1').length > 1.
- /issue-fixtures/empty-h1-heading#176 empty_h1_headingverified by: crawler page check
<h1><span class="icon-chart"></span></h1> — no text and no image with alt.
apps/crawler/toggleGroups/pageSeoBasics.js — h1HasText() false for at least one <h1>.
- /issue-fixtures/duplicate-h1-a#177 duplicate_h1_across_pagesverified by: API post-crawl judge
H1 "Analytics Dashboards for Growing Teams" — the same text as duplicate-h1-b.
apps/api/src/page-seo-basics-data/duplicate-h1.ts findDuplicateH1s — ≥ 2 indexable 2xx pages, self-canonical, same normalised first H1.
- /issue-fixtures/duplicate-h1-b#177 duplicate_h1_across_pagesverified by: API post-crawl judge
H1 "Analytics dashboards for growing teams" — the same text as duplicate-h1-a after case folding.
duplicate-h1.ts findDuplicateH1s — H1s compared after collapsing whitespace, trimming and lower-casing.
Delivery & Trust
- /issue-fixtures/hsts-max-age-too-short#170 hsts_max_age_too_shortverified by: crawler page check
Strict-Transport-Security: max-age=86400; includeSubDomains; preload
apps/crawler/toggleGroups/httpSecurityHeaders.js — parsed HSTS policy with 0 < max-age < 31536000.
Also raises, by design: hsts_not_preload_ready (#172) — a short max-age also fails the preload list
- /issue-fixtures/hsts-missing-include-subdomains#171 hsts_missing_include_subdomainsverified by: crawler page check
Strict-Transport-Security: max-age=63072000; preload
apps/crawler/toggleGroups/httpSecurityHeaders.js — valid HSTS policy without the includeSubDomains directive.
Also raises, by design: hsts_not_preload_ready (#172) — the preload list requires includeSubDomains
- /issue-fixtures/hsts-not-preload-ready#172 hsts_not_preload_readyverified by: crawler page check
Strict-Transport-Security: max-age=63072000; includeSubDomains (no preload token).
apps/crawler/toggleGroups/httpSecurityHeaders.js — valid HSTS policy missing any of: max-age ≥ 1 year, includeSubDomains, preload.
- /issue-fixtures/html-not-compressed#173 html_not_compressedverified by: crawler page check
≥ 1 KB of HTML sent with Cache-Control: no-transform, so neither Next nor Cloudflare compresses it.
apps/crawler/toggleGroups/httpSecurityHeaders.js — 2xx text/html response of ≥ 1024 bytes whose Content-Encoding is not gzip/br/zstd.
- /issue-fixtures/render-blocking-resources#191 render_blocking_resourcesverified by: crawler page check
Two synchronous <script src> tags and three stylesheets in the head.
apps/crawler/toggleGroups/httpSecurityHeaders.js + apps/crawler/utils/pagePerformance.js — any classic head script without async/defer, or more than 2 screen stylesheets.
- /issue-fixtures/excessive-dom-size#192 excessive_dom_sizeverified by: crawler page check
A 400-row event table (5 elements per row) — about 2,000 elements.
apps/crawler/toggleGroups/httpSecurityHeaders.js + pagePerformance.js measureDomSize — $('*').length > 1500.
- /issue-fixtures/poor-fcp#174 poor_fcpverified by: Lighthouse
An inline <script> in the head busy-waits 900 ms before the parser reaches the body (≈ 3.6 s+ under mobile CPU throttling).
apps/api/src/performance/performance.service.ts + cwv-thresholds.ts — FCP (CrUX p75, else lab) > 3000 ms.
- /issue-fixtures/poor-tbt#178 poor_tbtverified by: Lighthouse
After the first paint, one 800 ms task blocks the main thread (≈ 3.2 s under mobile CPU throttling).
apps/api/src/performance/lighthouse-checks.ts — lab total-blocking-time > 600 ms.
- /issue-fixtures/slow-speed-index#179 slow_speed_indexverified by: Lighthouse
An inline head script busy-waits 1,700 ms, so nothing is visible for ≈ 7 s under mobile CPU throttling.
apps/api/src/performance/lighthouse-checks.ts — lab speed-index > 5800 ms.
Also raises, by design: poor_fcp (#174) — nothing paints until the script finishes
- /issue-fixtures/unminified-css-js#180 unminified_css_jsverified by: Lighthouse
/fixture-assets/unminified.js — 36 KB, mostly comments and indentation, every function called.
apps/api/src/performance/lighthouse-checks.ts — unminified-css + unminified-javascript wastedBytes ≥ 2048.
- /issue-fixtures/unused-javascript#181 unused_javascriptverified by: Lighthouse
/fixture-assets/unused.js — 175 KB (≈ 90 KB gzipped) of 1,200 functions, none of them called.
apps/api/src/performance/lighthouse-checks.ts — unused-javascript wastedBytes ≥ 20480.
- /issue-fixtures/unused-css#182 unused_cssverified by: Lighthouse
/fixture-assets/unused.css — 60 KB (≈ 23 KB gzipped) of 900 rules whose class selectors match nothing.
apps/api/src/performance/lighthouse-checks.ts — unused-css-rules wastedBytes ≥ 10240.
- /issue-fixtures/short-static-asset-cache#183 short_static_asset_cacheverified by: Lighthouse
A 300 KB photo from /fixture-assets/short-ttl/ served with Cache-Control: public, max-age=600 (ten minutes).
apps/api/src/performance/lighthouse-checks.ts — uses-long-cache-ttl (else cache-insight) wastedBytes ≥ 28672.
- /issue-fixtures/offscreen-images-not-deferred#184 offscreen_images_not_deferredverified by: Lighthouse
A 1 MB photo placed 4,000 px down the page with no loading attribute, downloading while the page's script runs a 120 ms task.
apps/api/src/performance/lighthouse-checks.ts — offscreen-images wastedBytes ≥ 2048. Lighthouse 13 removed that audit and the API has no fallback, so on 13.x this code cannot fire.
- /issue-fixtures/oversized-images#185 oversized_imagesverified by: Lighthouse
The 800×534 photo displayed as a 120×80 thumbnail.
apps/api/src/performance/lighthouse-checks.ts — uses-responsive-images (else image-delivery-insight) wastedBytes ≥ 4096.
- /issue-fixtures/font-display-missing#186 font_display_missingverified by: Lighthouse
Google Fonts CSS requested without &display=…, so its @font-face rules carry no font-display.
apps/api/src/performance/lighthouse-checks.ts — font-display (else font-display-insight) lists at least one font.
- /issue-fixtures/missing-preconnect#187 missing_preconnectverified by: Lighthouse
Google Fonts (display=block) stylesheet and font files loaded with no <link rel=preconnect> to fonts.googleapis.com or fonts.gstatic.com.
apps/api/src/performance/lighthouse-checks.ts — uses-rel-preconnect (else network-dependency-tree-insight) lists at least one origin.
Also raises, by design: font_display_missing (#186) — Lighthouse 13 flags font-display: block too
- /issue-fixtures/lcp-image-lazy-loaded#188 lcp_image_lazy_loadedverified by: Lighthouse
The full-width hero photo — the largest element on screen — carries loading="lazy".
apps/api/src/performance/lighthouse-checks.ts — lcp-lazy-loaded fails (else lcp-discovery-insight eagerlyLoaded is false).
- /issue-fixtures/heavy-third-party-code#189 heavy_third_party_codeverified by: Lighthouse
@babel/standalone from cdn.jsdelivr.net compiles the page's <script type="text/babel"> playground source in the browser, on its own DOMContentLoaded listener.
apps/api/src/performance/lighthouse-checks.ts — third-party-summary blockingTime ≥ 250 ms (else third-parties-insight mainThreadTime ≥ 250 and TBT ≥ 250).
Also raises, by design: poor_tbt (#178) — the compile runs as one long task; unused_javascript (#181) — most of Babel is never called
- /issue-fixtures/page-weight-over-budget#190 page_weight_over_budgetverified by: Lighthouse
Four distinct 1 MB PNG requests (query strings defeat the cache) — ≈ 4 MB in total.
apps/api/src/performance/lighthouse-checks.ts — total-byte-weight > 3 MiB (3,145,728 bytes).
Also raises, by design: oversized_images (#185) — the photos are displayed at 380 px; offscreen_images_not_deferred (#184) — the lower photos are below the fold
Supporting pages
- /issue-fixtures/targets/gone — Answers 404. Canonical target of #159.
- /issue-fixtures/targets/moved — Answers 301 to the hub. Canonical target of #160.
- /issue-fixtures/targets/noindex — 200, noindex, self-canonical. Canonical target of #161.
- /issue-fixtures/targets/chain-middle — Canonicalises onward to chain-final. Middle hop of #162.
- /issue-fixtures/targets/chain-final — Self-canonical. End of the #162 chain.
- /issue-fixtures/paginated-canonical-to-first-page — Page 1 of the #163 series. Self-canonical.
Codes with no page fixture, and why
#169 ssl_certificate_expiring_soon— Read from the leaf TLS certificate of the host (apps/crawler/utils/ssl.js getCertificateInfo), which Cloudflare issues and renews for *.workers.dev. It fires when the certificate has ≤ min(30 days, a third of its lifetime) left. No page, header or file can change the certificate. Isolated fixture: scripts/tls-fixture/server.mjs serves https://localhost:3990/ with a fresh certificate issued 80 days ago that expires in 10.#204 outdated_tls_accepted— A TLS 1.0/1.1 handshake against the seed host on port 443 (apps/crawler/seo-audit-checks.js checkOutdatedTlsAccepted). That is decided by Cloudflare's edge, not by the site. The workers.dev edge accepted TLS 1.1 when probed on 2026-10-02, so this most likely already fires on every deployment of this site — through no page in particular. Isolated fixture: scripts/tls-fixture/server.mjs serves https://localhost:3991/ accepting TLS 1.0/1.1.
Index: /crawl-audit-fixtures.