Crawl Issues — Redirects, Depth & Canonicals
22 Intentional IssuesThis page demonstrates redirect, crawl depth, canonical, and link-related SEO issues. Broken internal links and redirect chains are intentionally present.
Robots.txt & Sitemap Configuration#3Issue #3: XML sitemap missing or invalid#5Issue #5: robots.txt missing or returning server error#164Issue #164: Sitemap directive not declared in robots.txt#165Issue #165: Crawl-delay directive missing#166Issue #166: Search results pages not blocked
The /robots.txt file intentionally lacks directives for site search blocking, sitemaps, and crawl delays. Additionally, AI bots are not explicitly configured.
AI Bot Issues:
- Issue #187: GPTBot (OpenAI) not explicitly allowed/disallowed
- Issue #188: ClaudeBot (Anthropic) not explicitly allowed/disallowed
- Issue #189: PerplexityBot not explicitly allowed/disallowed
- Issue #190: Google-Extended not explicitly allowed/disallowed
- Issue #191: CCBot (Common Crawl) not explicitly allowed/disallowed
Redirect Configuration & Loop Detection#8Issue #8: Redirect types (301/302) not correctly configured#11Issue #11: Redirect chains and loops present — hurt SEO performance#220Issue #220: Moved pages should return 301 for at least 1 year, not instant 404
This site uses 302 (temporary) redirects where 301 (permanent) redirects should be used. It also contains redirect chains and potential redirect loops (e.g., A → B → A or A → A) that waste crawl budget and cause browser errors.
Redirect Loop Checker Tool:
Run locally: node redirect-tracker.js --test
Tracked redirect loop examples detected programmatically: A → B → A, A → A.
⚠️ SEO ISSUE #11: Real HTTP Redirect Loop
These links create a real redirect loop: A → B → C → A (infinite)
Redirect Loop Test Endpoints#219Issue #219: Self redirect loop — /redirect-loop-self redirects to itself#220Issue #220: Two-page redirect loop — A↔B infinite loop#221Issue #221: Multi-step redirect loop — 1→2→3→1 cycle
The following URLs implement real HTTP-level redirect loops via Next.js Middleware. SEO crawlers following 3xx redirect responses will reliably detect these as loop conditions. All loops use 302 (temporary) redirects so no cache layer absorbs them.
SEO Issue #219 — Self Redirect Loop
GET /redirect-loop-self → 302 /redirect-loop-self → 302 /redirect-loop-self → ∞
/redirect-loop-self ↗SEO Issue #220 — Two-Page Redirect Loop (A ↔ B)
GET /redirect-loop-a → 302 /redirect-loop-b → 302 /redirect-loop-a → ∞
SEO Issue #221 — Multi-Step Redirect Loop (1 → 2 → 3 → 1)
GET /redirect-loop-1 → 302 /redirect-loop-2 → 302 /redirect-loop-3 → 302 /redirect-loop-1 → ∞
Verify with curl:
curl -I http://localhost:3000/redirect-loop-self
curl -I http://localhost:3000/redirect-loop-a
curl -I http://localhost:3000/redirect-loop-1
Expected: HTTP/1.1 302 Found → Location: [same or next loop URL]
www vs non-www#10Issue #10: www vs non-www redirects not correctly set up
Both www.example.com and example.com serve content without a canonical redirect, creating duplicate content.
Crawl Depth#14Issue #14: Pages more than 3 clicks from homepage — too deep for crawlers
Some pages on this site are buried more than 3 clicks from the homepage, making them harder for search engine crawlers to discover within their crawl budget.
Orphan Pages#17Issue #17: Pages with no internal links pointing to them
Several pages on this site have zero inbound internal links, making them invisible to search engine crawlers following links.
External Link Security#50Issue #50: External links missing rel=noopener noreferrer
External link without noopener noreferrer (intentional) →Canonical Tag Issues#172Issue #172: Canonical URL does not match URL in XML sitemap#175Issue #175: Canonical page blocked by robots.txt
SEO Issue #172 — Canonical vs Sitemap Mismatch:
This page canonical: https://acmeanalytics.example.com/crawl-issues
Sitemap URL: http://acmeanalytics.example.com/crawl-issues
Protocol mismatch (https vs http) — Google cannot confirm canonical
SEO Issue #175 — Canonical Blocked by robots.txt:
Canonical href: https://acmeanalytics.example.com/blocked-canonical/about
robots.txt rule: Disallow: /blocked-canonical/
Google can see the canonical tag but cannot fetch or validate the canonical URL
Image & Video Sitemaps#157Issue #157: No image sitemap entries in XML sitemap#180Issue #180: No video sitemap with required video:video entries
The sitemap declares image: and video: namespaces but contains ZERO <image:image> or <video:video> entries, reducing image/video discovery by Google.
AI Discoverability Files Missing#186Issue #186: /llms.txt file missing — AI crawlers cannot discover site content manifest#199Issue #199: /llms-full.txt file missing — AI systems cannot access full content
Files intentionally absent (returns 404):
GET /llms.txt → 404 Not Found (Issue #186)
GET /llms-full.txt → 404 Not Found (Issue #199)
llms.txt (emerging standard) tells LLMs what site content is available for citation. llms-full.txt provides extended full-content access for AI ingestion. Neither file exists on this site — intentional test case.
Page Indexability Audit#227Issue #227: Critical pages not confirmed indexable via audit
No indexability audit has been performed. Some pages may inadvertently carry a noindex directive or be blocked in robots.txt.