Guide · Pre-crawl URL status hygiene
Bulk broken-link and URL status checks before you crawl or RAG-ingest
A sitemap or spreadsheet can list thousands of URLs that look fine in a cell. Some return 404. Some redirect three hops to a different host. Some never resolve. If you feed that list straight into a crawler, PDF converter, or embedding job, you pay for failures, pollute logs, and may index soft landings that are not the document you thought you had.
This guide is the pre-crawl hygiene pass: probe status at scale, classify failures, inventory redirects, and decide what enters the expensive stage. You can do it with HEAD/GET scripts; an Actor is optional.
The short answer
- Take a URL list (paste, CSV, or the dataset from a sitemap discovery run).
- Probe each URL with HEAD; fall back to GET when the server blocks HEAD — read headers, discard the body.
- Record status code, final URL after redirects, content-type, and an error class (
dns,timeout,ssl,http,network, ornone). - Summarize by status class (2xx / 3xx / 4xx / 5xx / error). Optionally keep only broken rows in the detail dataset while the summary still covers every check.
- Drop or quarantine failures before crawl, conversion, or RAG ingest. Redirects are not “broken” by default — decide whether to follow the final URL or treat the hop as a content move.
Why status hygiene before crawl saves money and noise
Crawlers retry. Converters download. Embedders bill by token once text arrives. A list with even a modest failure rate wastes all three.
Suppose you have 1,000 document URLs from a sitemap. If about 6% fail DNS, SSL, or 4xx/5xx, that is roughly sixty wasted downloads — and sixty chances to log scary errors that look like pipeline bugs. Filtering first keeps the expensive stage boring.
Status hygiene also improves RAG quality. A 301 from an old policy PDF to a marketing homepage is worse than a clean 404: you might ingest the wrong page under the old URL’s identity if you follow redirects blindly. Recording finalUrl and redirect count lets you decide per case.
Workflow: bulk URL status before ingest
Step 1: Collect the list
Sources that work well:
- Sitemap discovery dataset (one
urlfield per row) - A
DOC_TO_MARKDOWN_INPUT-style record ({"urls":[{"url":"..."},...]}) - Hand-pasted URLs from a spreadsheet
Deduplicate. Cap maxUrls for a pilot. Same-host politeness matters: concurrency plus a per-host limit reduces 429s.
Step 2: Choose the HTTP method strategy
HEAD then GET is the practical default. Many CDNs and WAFs answer HEAD with 403 or 405 even when GET would return 200. On HEAD failure of that kind, retry with a streamed GET, keep headers, discard the body. Pure HEAD is faster when you already know the host allows it. Pure GET is fine but heavier on the wire if bodies are large — streaming and discarding avoids loading PDFs into memory just to learn they are 200.
Step 3: Classify every outcome
| Field | Use |
|---|---|
httpStatus |
Exact code when an HTTP response arrived |
statusClass |
Bucket: 2xx, 3xx, 4xx, 5xx, or error |
ok |
Whether you treat the URL as usable for the next stage |
finalUrl |
Landing URL after redirects |
redirectCount |
How many hops |
contentType / length |
Sanity-check expected PDF vs HTML |
errorClass |
dns, timeout, ssl, http, network, or none |
methodUsed |
HEAD or GET — useful when diagnosing WAF behavior |
DNS: host does not resolve — drop from crawl; optionally enrich the domain with WHOIS/DNS tools later.
Timeout: raise timeout, reduce concurrency, or retry later; do not silently mark as 404.
SSL: certificate or handshake failure — pause HTTPS ingest for that host.
HTTP: 4xx/5xx with a response — classic broken or server-error links.
Network: connection resets and similar — often transient or IP-block related.
Step 4: Decide what “broken” means for your pipeline
Defaults that work for most RAG prep:
- Keep 2xx with expected content-type
- Review 3xx: either adopt
finalUrlas the canonical source or exclude if the landing page is wrong - Exclude 4xx / 5xx / error classes from ingest
- Optionally write only broken (and maybe 3xx) rows to the detail dataset so operators get a short fix list, while SUMMARY still counts every probe
Redirects are not broken by themselves. Use a redirects-focused filter when you are auditing moves after a CMS migration.
Step 5: Chain from sitemap discovery
A clean chain looks like: discover URLs from sitemaps → status-check the list (filter broken) → convert or crawl only survivors. Pass the sitemap run’s dataset id or document-input KV record into the status step so you do not re-paste thousands of URLs by hand.
When HEAD is blocked
Illustrative pattern: suppose a CDN returns 403 on HEAD for every asset URL but 200 on GET with application/pdf. A HEAD-only checker would mark the whole set broken. HEAD-then-GET records methodUsed: GET and the real 200. If both fail, classify by status or error class and move on — do not retry forever.
Some hosts rate-limit or block data-center IPs. Retries with backoff on 429/5xx help; a proxy input helps when the block is IP-based. Soft-404s (HTML error pages that return HTTP 200) are out of scope for a pure status probe — they need content heuristics later.
Measured batch behavior (Actor documentation / own runs)
Figures below come from a bulk URL status Actor’s README and measured runs (local + private cloud), not third-party averages:
- Local: 5/5 URLs checked in about 0.3 s (mix of 2xx, 4xx, and DNS error)
- Cloud: 100 URLs in about 8.9 s, peak about 86 MB
- Cloud: 1,000 URLs in about 83 s, peak about 94 MB — ok 938 / issues 62
- Default memory often 256 MB; no browser
Use them to size timeouts and memory for similar public URL mixes. Your list’s hosts, geography, and rate limits will differ.
Manual alternative
curl -sI / curl -sI -L in a shell loop, or any HTTP client that follows redirects and prints final URL and status, implements the same method. Export CSV columns for statusClass and errorClass. The Actor path is optional when you want dataset chaining, SUMMARY aggregations, and HEAD-then-GET fallback without maintaining the script.
Related guides
Try the Actor or Lab tools
This guide works without buying anything. If you want the same workflow as a structured Apify Actor run, see URL Status Checker on Apify Store. Related browser tools: Lab /tools. Soft link only — no purchase required to use the checklist above.