⚠️ DEV TOOL — This website intentionally contains SEO issues for testing. Not for public use.Issues Index →

Crawl Issues — Redirects, Depth & Canonicals

22 Intentional Issues

This page demonstrates redirect, crawl depth, canonical, and link-related SEO issues. Broken internal links and redirect chains are intentionally present.

Robots.txt & Sitemap Configuration#3Issue #3: XML sitemap missing or invalid#5Issue #5: robots.txt missing or returning server error#164Issue #164: Sitemap directive not declared in robots.txt#165Issue #165: Crawl-delay directive missing#166Issue #166: Search results pages not blocked

The /robots.txt file intentionally lacks directives for site search blocking, sitemaps, and crawl delays. Additionally, AI bots are not explicitly configured.

AI Bot Issues:

  • Issue #187: GPTBot (OpenAI) not explicitly allowed/disallowed
  • Issue #188: ClaudeBot (Anthropic) not explicitly allowed/disallowed
  • Issue #189: PerplexityBot not explicitly allowed/disallowed
  • Issue #190: Google-Extended not explicitly allowed/disallowed
  • Issue #191: CCBot (Common Crawl) not explicitly allowed/disallowed

Redirect Configuration & Loop Detection#8Issue #8: Redirect types (301/302) not correctly configured#11Issue #11: Redirect chains and loops present — hurt SEO performance#220Issue #220: Moved pages should return 301 for at least 1 year, not instant 404

This site uses 302 (temporary) redirects where 301 (permanent) redirects should be used. It also contains redirect chains and potential redirect loops (e.g., A → B → A or A → A) that waste crawl budget and cause browser errors.

Redirect Loop Checker Tool:

Run locally: node redirect-tracker.js --test

Tracked redirect loop examples detected programmatically: A → B → A, A → A.

⚠️ SEO ISSUE #11: Real HTTP Redirect Loop

These links create a real redirect loop: A → B → C → A (infinite)

Broken internal link (404) — Issue #9

Redirect Loop Test Endpoints#219Issue #219: Self redirect loop — /redirect-loop-self redirects to itself#220Issue #220: Two-page redirect loop — A↔B infinite loop#221Issue #221: Multi-step redirect loop — 1→2→3→1 cycle

The following URLs implement real HTTP-level redirect loops via Next.js Middleware. SEO crawlers following 3xx redirect responses will reliably detect these as loop conditions. All loops use 302 (temporary) redirects so no cache layer absorbs them.

SEO Issue #219 — Self Redirect Loop

GET /redirect-loop-self → 302 /redirect-loop-self → 302 /redirect-loop-self → ∞

/redirect-loop-self ↗

SEO Issue #220 — Two-Page Redirect Loop (A ↔ B)

GET /redirect-loop-a → 302 /redirect-loop-b → 302 /redirect-loop-a → ∞

SEO Issue #221 — Multi-Step Redirect Loop (1 → 2 → 3 → 1)

GET /redirect-loop-1 → 302 /redirect-loop-2 → 302 /redirect-loop-3 → 302 /redirect-loop-1 → ∞

Verify with curl:

curl -I http://localhost:3000/redirect-loop-self

curl -I http://localhost:3000/redirect-loop-a

curl -I http://localhost:3000/redirect-loop-1

Expected: HTTP/1.1 302 Found → Location: [same or next loop URL]

www vs non-www#10Issue #10: www vs non-www redirects not correctly set up

Both www.example.com and example.com serve content without a canonical redirect, creating duplicate content.

Crawl Depth#14Issue #14: Pages more than 3 clicks from homepage — too deep for crawlers

Some pages on this site are buried more than 3 clicks from the homepage, making them harder for search engine crawlers to discover within their crawl budget.

Orphan Pages#17Issue #17: Pages with no internal links pointing to them

Several pages on this site have zero inbound internal links, making them invisible to search engine crawlers following links.

External Link Security#50Issue #50: External links missing rel=noopener noreferrer

External link without noopener noreferrer (intentional) →

Canonical Tag Issues#172Issue #172: Canonical URL does not match URL in XML sitemap#175Issue #175: Canonical page blocked by robots.txt

SEO Issue #172 — Canonical vs Sitemap Mismatch:

This page canonical: https://acmeanalytics.example.com/crawl-issues

Sitemap URL: http://acmeanalytics.example.com/crawl-issues

Protocol mismatch (https vs http) — Google cannot confirm canonical

SEO Issue #175 — Canonical Blocked by robots.txt:

Canonical href: https://acmeanalytics.example.com/blocked-canonical/about

robots.txt rule: Disallow: /blocked-canonical/

Google can see the canonical tag but cannot fetch or validate the canonical URL

Image & Video Sitemaps#157Issue #157: No image sitemap entries in XML sitemap#180Issue #180: No video sitemap with required video:video entries

The sitemap declares image: and video: namespaces but contains ZERO <image:image> or <video:video> entries, reducing image/video discovery by Google.

AI Discoverability Files Missing#186Issue #186: /llms.txt file missing — AI crawlers cannot discover site content manifest#199Issue #199: /llms-full.txt file missing — AI systems cannot access full content

Files intentionally absent (returns 404):

GET /llms.txt → 404 Not Found (Issue #186)

GET /llms-full.txt → 404 Not Found (Issue #199)

llms.txt (emerging standard) tells LLMs what site content is available for citation. llms-full.txt provides extended full-content access for AI ingestion. Neither file exists on this site — intentional test case.

Page Indexability Audit#227Issue #227: Critical pages not confirmed indexable via audit

No indexability audit has been performed. Some pages may inadvertently carry a noindex directive or be blocked in robots.txt.