Personal Lab ToolsApify Actors · workflow guides

Guide · Pre-crawl URL status hygiene

Bulk broken-link and URL status checks before you crawl or RAG-ingest

A sitemap or spreadsheet can list thousands of URLs that look fine in a cell. Some return 404. Some redirect three hops to a different host. Some never resolve. If you feed that list straight into a crawler, PDF converter, or embedding job, you pay for failures, pollute logs, and may index soft landings that are not the document you thought you had.

This guide is the pre-crawl hygiene pass: probe status at scale, classify failures, inventory redirects, and decide what enters the expensive stage. You can do it with HEAD/GET scripts; an Actor is optional.

The short answer

  1. Take a URL list (paste, CSV, or the dataset from a sitemap discovery run).
  2. Probe each URL with HEAD; fall back to GET when the server blocks HEAD — read headers, discard the body.
  3. Record status code, final URL after redirects, content-type, and an error class (dns, timeout, ssl, http, network, or none).
  4. Summarize by status class (2xx / 3xx / 4xx / 5xx / error). Optionally keep only broken rows in the detail dataset while the summary still covers every check.
  5. Drop or quarantine failures before crawl, conversion, or RAG ingest. Redirects are not “broken” by default — decide whether to follow the final URL or treat the hop as a content move.

Why status hygiene before crawl saves money and noise

Crawlers retry. Converters download. Embedders bill by token once text arrives. A list with even a modest failure rate wastes all three.

Suppose you have 1,000 document URLs from a sitemap. If about 6% fail DNS, SSL, or 4xx/5xx, that is roughly sixty wasted downloads — and sixty chances to log scary errors that look like pipeline bugs. Filtering first keeps the expensive stage boring.

Status hygiene also improves RAG quality. A 301 from an old policy PDF to a marketing homepage is worse than a clean 404: you might ingest the wrong page under the old URL’s identity if you follow redirects blindly. Recording finalUrl and redirect count lets you decide per case.

Workflow: bulk URL status before ingest

Step 1: Collect the list

Sources that work well:

  • Sitemap discovery dataset (one url field per row)
  • A DOC_TO_MARKDOWN_INPUT-style record ({"urls":[{"url":"..."},...]})
  • Hand-pasted URLs from a spreadsheet

Deduplicate. Cap maxUrls for a pilot. Same-host politeness matters: concurrency plus a per-host limit reduces 429s.

Step 2: Choose the HTTP method strategy

HEAD then GET is the practical default. Many CDNs and WAFs answer HEAD with 403 or 405 even when GET would return 200. On HEAD failure of that kind, retry with a streamed GET, keep headers, discard the body. Pure HEAD is faster when you already know the host allows it. Pure GET is fine but heavier on the wire if bodies are large — streaming and discarding avoids loading PDFs into memory just to learn they are 200.

Step 3: Classify every outcome

Field Use
httpStatus Exact code when an HTTP response arrived
statusClass Bucket: 2xx, 3xx, 4xx, 5xx, or error
ok Whether you treat the URL as usable for the next stage
finalUrl Landing URL after redirects
redirectCount How many hops
contentType / length Sanity-check expected PDF vs HTML
errorClass dns, timeout, ssl, http, network, or none
methodUsed HEAD or GET — useful when diagnosing WAF behavior

DNS: host does not resolve — drop from crawl; optionally enrich the domain with WHOIS/DNS tools later.

Timeout: raise timeout, reduce concurrency, or retry later; do not silently mark as 404.

SSL: certificate or handshake failure — pause HTTPS ingest for that host.

HTTP: 4xx/5xx with a response — classic broken or server-error links.

Network: connection resets and similar — often transient or IP-block related.

Step 4: Decide what “broken” means for your pipeline

Defaults that work for most RAG prep:

  • Keep 2xx with expected content-type
  • Review 3xx: either adopt finalUrl as the canonical source or exclude if the landing page is wrong
  • Exclude 4xx / 5xx / error classes from ingest
  • Optionally write only broken (and maybe 3xx) rows to the detail dataset so operators get a short fix list, while SUMMARY still counts every probe

Redirects are not broken by themselves. Use a redirects-focused filter when you are auditing moves after a CMS migration.

Step 5: Chain from sitemap discovery

A clean chain looks like: discover URLs from sitemaps → status-check the list (filter broken) → convert or crawl only survivors. Pass the sitemap run’s dataset id or document-input KV record into the status step so you do not re-paste thousands of URLs by hand.

When HEAD is blocked

Illustrative pattern: suppose a CDN returns 403 on HEAD for every asset URL but 200 on GET with application/pdf. A HEAD-only checker would mark the whole set broken. HEAD-then-GET records methodUsed: GET and the real 200. If both fail, classify by status or error class and move on — do not retry forever.

Some hosts rate-limit or block data-center IPs. Retries with backoff on 429/5xx help; a proxy input helps when the block is IP-based. Soft-404s (HTML error pages that return HTTP 200) are out of scope for a pure status probe — they need content heuristics later.

Measured batch behavior (Actor documentation / own runs)

Figures below come from a bulk URL status Actor’s README and measured runs (local + private cloud), not third-party averages:

  • Local: 5/5 URLs checked in about 0.3 s (mix of 2xx, 4xx, and DNS error)
  • Cloud: 100 URLs in about 8.9 s, peak about 86 MB
  • Cloud: 1,000 URLs in about 83 s, peak about 94 MB — ok 938 / issues 62
  • Default memory often 256 MB; no browser

Use them to size timeouts and memory for similar public URL mixes. Your list’s hosts, geography, and rate limits will differ.

Manual alternative

curl -sI / curl -sI -L in a shell loop, or any HTTP client that follows redirects and prints final URL and status, implements the same method. Export CSV columns for statusClass and errorClass. The Actor path is optional when you want dataset chaining, SUMMARY aggregations, and HEAD-then-GET fallback without maintaining the script.

Try the Actor or Lab tools

This guide works without buying anything. If you want the same workflow as a structured Apify Actor run, see URL Status Checker on Apify Store. Related browser tools: Lab /tools. Soft link only — no purchase required to use the checklist above.