Personal Lab ToolsApify Actors · workflow guides

Guide · Domain enrichment for RAG / crawl prep

Batch WHOIS, DNS, and SSL lookup before you enrich a RAG knowledge base

You have a spreadsheet of domains — partners, vendors, archived microsites, hosts from an old crawl. Before you spend crawl budget and embedding tokens on a knowledge base or RAG index, you need to know which hosts are still live, whether the certificate is about to expire, and whether the nameservers look like a real site or a parked placeholder.

This guide is the operator checklist: what to look up, why each field matters for KB hygiene, how to spot expired or misconfigured domains, and a batch workflow with public tools or a light Actor. You do not need to buy anything to follow the method.

The short answer

Before you index hosts into a knowledge base, enrich each domain with three cheap signals:

  1. WHOIS — registrar, creation date, and expiry (when the public server returns them). Expiry and thin/parked patterns help you drop dead names.
  2. DNS — A/AAAA, MX, NS, and TXT. Missing A records, odd NS sets, or parking-style TXT tell you the host may not serve real content.
  3. TLS leaf certificate — issuer, validity window, and days left. Short remaining life or handshake failures are a reason to pause crawl or flag the host.

Do this on a unique domain list (not every URL). Failed lookups should not block the batch. Attribute benches to your own runs or Actor documentation — not to industry averages.

Why registrar, expiry, NS, and SSL days-left matter before indexing

A RAG pipeline treats a host as a source of truth once pages are fetched and chunked. If the domain expired last month, resolves to a parking page, or presents a broken certificate, you still pay for the fetch — and you may ingest junk that looks authoritative in retrieval.

Registrar and expiry. Public WHOIS is often privacy-redacted, but registrar, created, and expires still arrive for many TLDs. Suppose you have 200 partner domains and twenty expire within 30 days. Those twenty are higher risk for sudden NXDOMAIN or registrar parking. Flag them; do not silently index them next to stable sources.

Nameservers (NS). A healthy site usually has a small, consistent NS set. Abrupt switches to parking or aftermarket NS hosts often mean the domain changed hands or lapsed. Outliers deserve a manual check before crawl.

A/AAAA and MX. No A or AAAA means nothing useful to crawl on usual web ports. MX alone does not prove a website exists. TXT with only SPF “-all” and no web records is common on reserved names — fine for demos, useless as KB sources.

SSL days left. TLS failures waste crawler retries. A leaf cert with few days left is a scheduling signal: renew or deprioritize before a large ingest. Issuer and SAN help confirm you are talking to the host you expect.

None of this replaces content review. It is hygiene so you do not feed dead or risky hosts into the index.

Workflow: enrich a domain list without a browser

Step 1: Deduplicate to hostnames

Strip paths, ports, and schemes. Lowercase. Drop IP literals if your tooling expects hostnames only. One row per unique host keeps cost linear with domains, not with every page URL.

Step 2: Decide which modules you need

For crawl prep on a small list, enable WHOIS, DNS, and SSL together. For a large list, stage it: DNS first (fastest NXDOMAIN signal), SSL on hosts that resolve, WHOIS with a polite delay. Public WHOIS servers rate-limit — keep concurrency low.

Step 3: Classify each row

Outcome Meaning Typical next step
ok All requested modules returned usable data Eligible for crawl / index queue
partial Some modules worked, others failed Keep with flags; inspect failed module
error Lookup failed completely (e.g. NXDOMAIN) Exclude from crawl

Partial is common: for example, DNS NXDOMAIN with WHOIS still returning expiry, or SSL handshake fail while DNS resolves. Do not treat partial as “delete” by default — treat it as “needs a human flag.”

Step 4: Pre-index checklist

For each domain marked ok or partial:

  1. Is whoisExpires (when present) more than N days away? (Pick N for your risk — for example 30 or 60.)
  2. Do NS names look consistent with a real operator, not a parking pattern you already know?
  3. Is there at least one A or AAAA if you plan an HTTPS crawl?
  4. Is sslDaysLeft above your threshold, and did the handshake succeed?
  5. Are errors empty or limited to modules you intentionally skipped?

Hosts that fail go to a review sheet, not into tonight’s KB ingest.

Step 5: Store structured metadata

Keep registrar, created, expires, NS, A/AAAA, SSL issuer, validTo, and daysLeft next to the host. Later debugging is easier when you can see “this source expires in 12 days” beside a weird answer.

Spotting expired, parked, and misconfigured domains

Expired or about to expire. WHOIS expires in the past or within days. DNS may still resolve briefly after expiry; do not trust A records alone.

Parked. Resolves, but NS or HTTP later shows a marketplace holding page. WHOIS may show a recent registrar change. Sample a URL with a status check after enrichment.

Misconfigured DNS. Empty A/AAAA with MX still present; NS that do not answer. These burn crawl time.

SSL problems. Handshake errors, very short daysLeft, or SAN that does not match the domain. Pause HTTPS crawl until fixed.

Illustrative example: suppose your list includes partner-old.example with WHOIS expiry yesterday, NS at a parking provider, and no A record. Leave it out of tonight’s RAG ingest. Re-check next quarter if the partnership returns.

Measured batch behavior (Actor documentation / own runs)

If you use a bulk WHOIS/DNS/SSL Actor that looks up public WHOIS, DNS (A/AAAA/MX/NS/TXT), and the TLS leaf certificate without a browser or paid WHOIS API keys, figures from that Actor’s README and measured runs include:

  • Local smoke: 2/2 domains ok in about 0.7 s, peak about 80 MB
  • Cloud: 50 domains — 46 ok / 4 partial, peak about 94 MB
  • Cloud: 100 domains — 92 ok / 8 partial, peak about 95 MB
  • Default memory often 256 MB; failed domains free by default (only ok/partial charged unless you opt in)

Use those only as sizing hints. Public WHOIS rate limits and thin TLD responses will change wall time and partial rates on your list.

Manual alternative (no Actor)

You can approximate the same hygiene with whois, dig, and openssl s_client in a loop. The method above still applies: unique hosts, polite WHOIS pacing, classify ok/partial/error, checklist before crawl. The Actor path is optional automation at list scale.

Try the Actor or Lab tools

This guide works without buying anything. If you want the same workflow as a structured Apify Actor run, see WHOIS DNS SSL Lookup on Apify Store. Related browser tools: Lab /tools. Soft link only — no purchase required to use the checklist above.