Training data
When are consent, licensing, attribution, or compensation required for copyrighted and personal material?
Courts began answering in 2025, partially crediting training as transformative learning while rejecting it for pirated corpora. The supply-chain framing maps who owes whom at each stage from scraping to output.
View on the map → · Open in Browse →
What changed
01
The money and the precedents both arrived: final approval of the ~$1.5B Bartz settlement paid authors about $3,100 a title while releasing only past conduct, the first federal appellate test of fair use for AI training reached oral argument, Europe's first major ruling treated memorization in model weights as reproduction — and the UK abandoned its opt-out proposal after creator backlash.
Recent thinking
Chloe Veltman · NPR · 27 Jul 2026 news
Authors have mixed feelings about the $1.5B Anthropic copyright infringement rulingThe algorithm is being used to essentially try to put us out of a job
Reports final approval of the Bartz v. Anthropic settlement (~$3,100 per title to over 300,000 authors) and documents why many authors see the compensation as validating training on their work rather than vindicating them.
Dave Hansen · Kluwer Copyright Blog · 10 Nov 2025 essay
The Bartz v. Anthropic Settlement: Understanding America's Largest Copyright SettlementThis settlement only releases Anthropic from liability for past conduct—specifically, its acquisition, retention, and use of the identified pirated works before August 25, 2025.
The best analytical unpacking of what the settlement does and does not resolve: class certification plus statutory damages created the leverage, outputs and future conduct remain open, and a 'shadow library strategy' template now exists for plaintiffs.
Bob Ambrogi · LawSites · 26 Jun 2026 news
At 3rd Circuit, Judges Press ROSS and Thomson Reuters on Fair Use, AI Training and Market HarmThe question has legal language in it, and what we have taught this machine now is how to think like a lawyer.
The first federal appellate test of fair use for AI training reached oral argument in June 2026; the panel's questioning on transformativeness and market harm previews the first binding appellate precedent.
Alexander Fewtrell & Angel Skyers · Lewis Silkin · 24 Mar 2026 essay
Opt-out cop-out? UK Government rethinks its position on copyright and AIWe propose to gather further evidence on how copyright laws are impacting the development and deployment of AI
Analyzes the UK government's March 2026 report abandoning its preferred opt-out text-and-data-mining exception after creator backlash — a major national policy reversal on the consent-and-licensing question.
Giancarlo Frosio · Kluwer Copyright Blog · 10 Dec 2025 essay
Copyright in Formaldehyde: How GEMA v OpenAI Freezes Doctrine and Chills AImodel parameters sit in a space that copyright doctrine has always treated cautiously: they are closer to facts, statistics and functional logic than to expressive form.
The leading scholarly critique of the Munich court's GEMA v. OpenAI ruling — Europe's first major decision treating memorization in model weights as reproduction — arguing it mischaracterizes lossy compression as storage.
Additional relevant discussion (2)
Foundational reading (3)
Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply ChainLee, Cooper & Grimmelmann · 2023Resigning from Stability AI over the claim that training on copyrighted work is “fair use”Ed Newton-Rex (@ednewtonrex) on X · 2023Generative AI and Copyright Law (covering the 2025 fair-use rulings)Congressional Research Service · 2025