In July 2026 a court gave final approval to the $1.5B Anthropic settlement, the largest
copyright settlement in US history, over how training data was obtained.1 It settled, so there is no finding of liability in it. The number is what one
unresolved question cost one company.
The cases squarely about real-time retrieval keep surviving dismissal. Dow Jones v.
Perplexity was pleaded over both the RAG index and the outputs (August 2025). In Advance
Local Media v. Cohere, claims over substitutive RAG summaries were allowed to proceed
(November 2025). In Reddit v. Perplexity and SerpApi, DMCA claims survived against the
scraping vendor in the supply chain rather than the AI company alone (July 2026).2
Surviving dismissal means the claims proceed. It is not a finding of infringement, and
no court has ruled that retrieval is unlawful. The other half of the picture is just as
unsettled: no US court has held that retrieval-time copying for grounding is fair use
either. The question is undecided rather than decided against anyone.
Vendor indemnities do not close that gap. Microsoft's Customer Copyright Commitment,
OpenAI's Copyright Shield, and Anthropic's equivalent protections cover outputs, and
they are conditioned on you having the rights to your inputs.3
Meanwhile the open web is narrowing. Cloudflare has blocked AI crawlers by default since
July 2025, and from 15 September 2026 it extends default blocking of mixed-use crawlers
on ad-supported pages.4 Licensing infrastructure arrived in the same window: the RSL 1.0 standard published in
December 2025, and Microsoft's Publisher Content Marketplace pays publishers per use for
grounding content.
robots.txt and RSL declarations are publisher signals rather than binding law. They
state a preference a retrieval layer can follow, and following them is a posture you can
describe to counsel rather than a defense you can rely on in court.