Personal Lab ToolsApify Actors · workflow guides

Guide · Sitemap inventory for document / RAG pipelines

Find every PDF and DOCX a site publishes via sitemaps (for document and RAG pipelines)

Link crawlers follow HTML. Many organizations publish reports, forms, manuals, and policy PDFs that appear in XML sitemaps — and barely appear as in-page links. If you are building a document pipeline or RAG corpus from a government, university, or corporate site, starting from “crawl the homepage” often under-counts the files that matter.

This guide shows how to inventory PDF, DOCX, and related document URLs from sitemaps: discovery via robots.txt and common paths, nested indexes, filters, and what to do when one sitemap is broken. The method works with ordinary HTTP tools; an Actor is optional automation at the end.

The short answer

  1. Discover sitemaps from robots.txt Sitemap: lines and common paths such as /sitemap.xml and /sitemap_index.xml.
  2. Follow nested sitemap indexes; accept gzip and plain-text formats; dedupe URLs.
  3. Tag each URL by file type from the extension (pdf, docx, xlsx, …); optionally confirm with an HTTP HEAD Content-Type check.
  4. Filter to the types you need, cap max URLs, and treat per-sitemap failures as report rows — not a failed whole job.
  5. Hand the PDF/DOCX list to your converter or chunker. Do not assume a link crawl will find the same set.

HTML crawl seeds miss documents when:

  • Files live only in a dedicated PDF sitemap (common on large public sites)
  • Links are behind login, JS menus, or download handlers the crawler never reaches
  • Older reports remain in the sitemap for search engines but were removed from nav
  • The site uses sitemap indexes with dozens of children; a shallow crawl never touches them

Sitemaps exist so machines can learn the URL set the publisher wants indexed. For document pipelines, that inventory is usually a better seed than BFS from the homepage — with the caveat that sitemaps only list what the site chose to publish. No sitemap means this method returns nothing; then you fall back to crawl or a known file index.

Workflow: sitemap inventory for document pipelines

Step 1: Start from the homepage (or paste known sitemap URLs)

For a site such as <https://www.example.gov,> fetch /robots.txt and collect every Sitemap: URL. Also probe common paths. Many CMS setups expose /sitemap.xml as an index that points at child sitemaps. If you already know a PDF-only sitemap URL, paste it directly and skip discovery for that host.

Step 2: Recurse indexes carefully

Sitemap indexes list other sitemaps. Follow children up to a depth you choose. Skip loops and duplicate child URLs so the same 21 Yoast-style children are not read twice. Record each robots.txt and sitemap with its own status: ok, not_found, error, or skipped. Probing /sitemap.xml when the site does not have one should be not_found, not a fatal error.

Step 3: Parse formats you will actually meet

Expect:

  • XML urlset and sitemapindex
  • .xml.gz or gzip transfer encoding
  • Plain-text sitemaps (one URL per line)
  • Occasionally RSS/Atom feeds used as URL lists

Tolerant parsing helps when XML is slightly malformed. Cap size per sitemap so a multi-hundred-MB file cannot blow memory. HTTP only is enough; you do not need a browser to read sitemaps.

Step 4: Tag file types and filter

From the URL extension, tag pdf, docx, doc, xlsx, xls, pptx, ppt, csv, html, image, or other. For a document pipeline, set filters to pdf and docx (add spreadsheet/slide types if your converter supports them).

Optional HTTP HEAD: confirm Content-Type (for example application/pdf). Useful when download handlers hide the type in the path. HEAD adds request volume; leave it off for a first inventory pass, then confirm a sample.

Also useful filters: lastmod date (“changed on or after”), include/exclude regex, same-domain only, max URLs total and per site. Caps keep a 20k-URL gov site from flooding a small test run.

Step 5: Export for the next stage

Save one row per URL with lastmod, source sitemap, host, and file type. For RAG, the valuable artifact is often a ready-made list of PDF/DOCX URLs you can pass to a PDF/DOCX-to-Markdown (or OCR) step. Keep a per-sitemap report so operators can see which child failed without re-running the whole site.

Worked examples (measured / illustrative)

Measured (Actor documentation / own runs, not industry averages): On fec.gov, robots.txt listed multiple sitemaps including a PDF-oriented one; a run tagged 69 PDF URLs, and HEAD checks returned HTTP 200 with application/pdf for all 69/69. On gov.uk, a large child sitemap was inventoried at about 20,000 URLs in roughly 11 seconds on the Apify platform at 256 MB default memory (peak about 107 MB). These are small tests on specific sites; unusual setups can still fail — when they do, the reason belongs in a per-sitemap report.

Illustrative: Suppose a university publishes theses only under /sitemap-pdf.xml and never links them from the library homepage. A link crawler seeded at the homepage might find dozens of HTML pages and zero theses. A sitemap pass filtered to pdf returns the thesis set for conversion.

Illustrative failure handling: Suppose one of five child sitemaps returns 404 and another times out. The run should still save URLs from the three healthy children and mark the bad ones in the report. Broken sitemap ≠ failed whole job.

Filters and honesty limits

  • Extension-based tags mislabel /download?id=123 until you HEAD or open the response.
  • Image/video extensions inside sitemap image/video namespaces may not appear as separate rows depending on the parser.
  • Sites that block data-center IPs may need a proxy.
  • This method does not crawl HTML for in-page links; it only reads sitemaps and feeds.

After inventory, status-check the document URLs before you pay for conversion (404s and redirects waste converter quota). That is a separate bulk URL status step.

Optional next step: convert PDF/DOCX to Markdown

Once you have a clean PDF/DOCX URL list, you can convert files to Markdown or RAG chunks. One optional chain is the PDF & DOCX to Markdown Actor using a DOC_TO_MARKDOWN_INPUT-shaped payload ({"urls":[{"url":"..."},...]}). Mind that Actor’s own pricing and page limits before sending hundreds of files. Conversion is not required to get value from the sitemap inventory itself.

Try the Actor or Lab tools

This guide works without buying anything. If you want the same workflow as a structured Apify Actor run, see Sitemap URL Discovery on Apify Store. Related browser tools: Lab /tools. Soft link only — no purchase required to use the checklist above.