Firecrawl, Tavily, Exa, and Bright Data-class APIs are engineered to maximize coverage,
including getting through anti-bot defenses. That is their design goal, and they are
good at it. Plenty of agent products work because of it.
Firecrawl shipped Lockdown Mode in April 2026, marketed on prompt-injection defense and
approval workflows,5 so
the contrast here is not that scraper-tier vendors ignore security.
The contrast is which posture is the default. Clearfetch is governance-first end to end:
deterministic channel elimination, a verdict on every response, provenance, declared
license status, and immutable per-key audit logs. We do not circumvent bot protections,
so no part of the business pulls against that default. Pick by which default you want to
explain to a security team.
For bot-walled and JavaScript-heavy targets, a scraper will retrieve more than we do, by
design, because we do not circumvent. If your workload depends on that content, we are
not a full replacement. Some teams will run both, with Clearfetch as the governed
default path and scraper output quarantined for higher-risk review. We would rather say
that before you integrate than after.
The supply-chain version of this question got sharper in July 2026, when Reddit's
anti-circumvention claims survived dismissal against SerpApi, a scraping vendor in the
pipeline, rather than against the AI company alone.6 Survived dismissal is the whole of it. The claims proceed and nothing has been decided
on the merits. Read that as a fact about the supply chain, not as an accusation against
any tool you run today.
Other vendors use “retrieval firewall” for other products, including data-loss
prevention in front of vector databases. We did not coin the phrase. We use it for the
open-web ingress path, meaning what your agents read.