Threat Protection is an enterprise feature and is gated per organization. Contact your Firecrawl account team to have it enabled for your account.
Modes
Threat Protection has three modes, set at the organization level:- Off (default) — no checks are performed.
- Normal — URLs are checked against Google Web Risk, which flags pages and sites associated with malware, social engineering (phishing), and unwanted software. +2 credits per URL scanned.
- Zscaler — URLs are checked against your organization’s own Zscaler Internet Access (ZIA) tenant: the Zscaler-defined URL categories you choose to block, plus your custom URL categories and custom URL lists. See Zscaler mode below. No scan fees — classification runs against your own tenant.
Policy controls
Beyond the classifier, a policy can include:- Custom blacklist — exact domains or globs (e.g.
*.example.com) that are always blocked, without a classifier call. - Custom whitelist — exact domains or globs that are always allowed. The whitelist wins over every other rule, so a domain you trust is never blocked.
- Blocked TLDs — top-level domains to block outright (e.g.
zip), matched on label boundaries. - Risk score threshold — the normalized score (0–100) at or above which a classifier verdict is treated as a block. Lower is stricter. The default is
75. Applies to Normal mode; Zscaler mode blocks by category instead. - Failure policy — what to do when the classifier can’t be reached: block (
closed, the default and recommended for a security control) or allow (open).
Zscaler mode
Zscaler mode lets an organization that already maintains URL policy in ZIA enforce it on Firecrawl traffic, without maintaining a parallel taxonomy. Two planes work together:- Inline classification — URLs are classified through your tenant’s URL Lookup API into Zscaler-defined categories, and URLs in a category you have denied are blocked.
- Synced custom rules — your custom URL categories (URL lists and keywords) and your additions to Zscaler-defined categories are synced from the tenant on a schedule and evaluated by Firecrawl directly, because the ZIA lookup API does not return custom classifications. Entries removed in ZIA disappear on the next sync; a manual Sync now is available in the dashboard.
Connecting your tenant
Team admins connect the tenant from the dashboard (see below) with a Zidentity OAuth client: client ID, client secret, and your Zidentity vanity domain. Use a least-privilege API role scoped to URL Categories. The Test connection button verifies three things separately — the credentials, taxonomy access, and URL Lookup access — so a role that can read categories but not classify URLs is surfaced at setup rather than on your first scrape. The client secret is write-only and encrypted at rest; Zscaler’s support for two active secrets lets you rotate with no downtime. After connecting, pick the categories to block from your tenant’s own taxonomy — Zscaler-defined and custom categories both appear in the picker. Classification traffic to your tenant originates from a dedicated static IP for allowlisting; ask your account team for the address.Capacity and behavior
ZIA’s URL Lookup API allows 1 request per second and 400 requests per hour per tenant. Firecrawl batches lookups (up to 100 URLs per request) and enforces these limits tenant-wide, which gives a sustained classification ceiling of roughly 11 URLs per second. Your scrape throughput is never throttled: when demand outruns the budget, or the hourly budget is exhausted, affected requests resolve through your failure policy immediately instead of queueing indefinitely. Two endpoint-level differences from Normal mode:- Map results are evaluated against local rules only (your lists plus synced custom rules) rather than classified inline — one map can return thousands of URLs, and classifying them would burn the hourly budget on links that may never be fetched. Every URL still gets the full check when a scrape of it starts.
- Search results are classified inline and blocked results are removed, same as Normal mode; each unique result draws on the hourly lookup budget.
Configuring the policy
Team admins configure Threat Protection from Enterprise Controls → Threat Protection in the dashboard:- Open Enterprise Controls → Threat Protection.
- Choose a mode, set your risk score threshold, and add any blacklist, whitelist, or blocked-TLD entries.
- For Zscaler mode: enter the tenant connection, run the connection test, pick the categories to block, and set the sync interval.
- Choose whether to allow per-request overrides, and set the failure policy.
- Save. Changes take effect immediately — the next request is evaluated against the new policy.
Per-request overrides
Every endpoint that accepts URLs also accepts an optionalthreatProtection object, so an individual request can tighten (or, if your organization allows it, adjust) the policy for that call:
threatProtection object is rejected with a 403 — this lets an administrator guarantee that the organization policy is the floor for every request.
An override may select "mode": "zscaler" only when the organization has a Zscaler connection configured; the connection itself and the denied-category selection are organization-level and cannot be set per request.
If Threat Protection is enforced for your team, an override may still tighten the policy but may not include "mode": "off" — a request that tries is rejected with a 403.
When a URL is blocked
A blocked request fails with a403 and a stable error code:
- Scrape, batch scrape, extract, agent — a blocked target returns the
unsafe_domain_blockederror for that URL. - Crawl — a blocked seed URL fails the request; blocked links discovered mid-crawl are skipped and the crawl continues.
- Search, map — blocked URLs are removed from the returned results rather than surfaced and refused.
Billing
In Normal mode a scan costs +2 credits per URL scanned, on top of the base cost of the request. Zscaler mode has no scan fees: classification runs against your own ZIA tenant, with your credentials and your API quota, so the scan-fee details below apply to Normal mode only. A few details:- Decisions made entirely from your own policy (blacklist, whitelist, or blocked-TLD matches) do not call the classifier and are not charged a scan fee.
- A request that is blocked is still charged for the scan that produced the verdict.
- Scans are deduplicated within a single scrape: a redirect re-check that resolves to the same URL shares the original scan, while a redirect that lands on a different URL is a second scan.
- Crawls and batch scrapes check every page independently. Verdicts are never reused across pages — nothing about your traffic is stored (see above) — so in Normal mode expect +2 credits per scraped page. A link that is discovered mid-crawl and blocked bills its scan once per crawl, no matter how many pages link to it.
- Search and map scan each unique URL in the result set once per request, so their scan fees scale with the number of results scanned — which can slightly exceed the number returned when results are trimmed to your
limit.
Error reference
Notes
- The policy is organization-wide: it applies to every API key and every endpoint automatically.
- The whitelist always wins, so a URL on an explicitly trusted domain is never blocked by the classifier or a TLD rule.
- The error code
unsafe_domain_blockedis kept stable for compatibility even though checks are URL-level. - With the failure policy set to
closed(the default), a classifier outage causes affected requests to be blocked rather than silently allowed. - With SIEM Audit Logging configured, every decision is visible in your audit trail: events carry the deciding rule, the classifier consulted, the threat categories, and — for Zscaler-classified URLs — a
security_alertflag when the URL carried a security-alert classification.

