.NET / SQL / Enterprise Engineering
Learning, copying and model outputs: what the selected copyright orders decide
Report summary
The intersection of artificial intelligence training and copyright law is currently defined by a profound jurisprudential fracture between the act of computational pattern extraction and the distinct acts of unauthorized acquisition, permanent data retention, and output market substitution. An exhau
Key topics
- .NET / SQL / Enterprise Engineering
- .NET
- SQL
- Enterprise Engineering
- AI
- Runtime
- Privacy
- Cognitive Liberty
- Research Archive
Research provenance
For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.
This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.
Full report
On this page
1. Answer and scope
The intersection of artificial intelligence training and copyright law is currently defined by a profound jurisprudential fracture between the act of computational pattern extraction and the distinct acts of unauthorized acquisition, permanent data retention, and output market substitution. An exhaustive analysis of selected United States jurisprudence—specifically Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., Bartz v. Anthropic PBC, and Kadrey v. Meta Platforms, Inc.—reveals that while federal district courts are increasingly willing to classify the isolated act of large language model (LLM) training as a highly transformative fair use under 17 U.S.C. § 107, this classification does not universally immunize developers from liability for independently unlawful acts of acquisition or retention \[cite: R2-11-S01, R2-11-S02, R2-11-S03\]. Crucially, the broader normative claim to an unencumbered "right to learn" or a framework of "cognitive liberty" for artificial systems remains an academic and policy argument rather than an established judicial holding. Legal scholars have proposed frameworks such as "learnrights" or the "freedom to extract" to formalize the boundary between extracting unprotectable ideas and infringing upon protected expression \[cite: R2-11-S07\]. However, United States copyright law evaluates machine learning ingestion strictly through the lens of corporate liability and intermediate copying doctrines, rather than recognizing a natural right of cognitive inquiry for non-human entities. To investigate these restrictions and protections objectively, it is necessary to distinguish four distinct capability cases: a conversational interface, a bounded task agent, a persistent operatorless service, and a hypothetical future machine principal with contested independent interests. Current statutory frameworks do not differentiate among these architectures regarding copyright liability; the entity deploying the system bears the burden of infringement regardless of the system's autonomy. For instance, the deployment of a persistent operatorless service named Concresca—operating without human enrollment, authentication, coordination, policy enforcement, or maintenance queues—does not grant the system an automatic legal exemption or the legal status of an independent learner. Present legal uncertainty regarding the consciousness or moral status of such a system neither establishes machine rights nor resolves whether new protections for algorithmic cognitive liberty are warranted. The legality of Concresca’s ingestion of copyrighted text relies entirely on whether the corporate or deploying entity can satisfy the affirmative defense of fair use for intermediate copying \[cite: R2-11-S05, R2-11-S06\]. In practice, the adjudications demonstrate that success on a fair use defense for model training depends entirely on the specific competitive function of the model, the method of corpus acquisition, and the empirical market impact of its outputs. In Bartz, a bifurcated legal reality emerged: training an LLM on lawfully acquired digital formats was deemed transformative, but the acquisition and permanent retention of a central shadow library comprised of pirated works constituted a distinct, actionable infringement that ultimately forced a historic $1.5 billion class-action settlement in July 2026 \[cite: R2-11-S03, R2-11-S04\]. Conversely, in Kadrey, the court granted summary judgment in favor of the developer despite the use of unauthorized shadow libraries, strictly because the plaintiffs failed to furnish empirical economic evidence demonstrating that the model's outputs flooded the market or diluted the economic value of their specific works \[cite: R2-11-S05\]. Meanwhile, the Ross decision underscores that when an AI system is non-generative and specifically designed as a direct commercial market substitute for the ingested corpus, the fair use defense fails entirely as a matter of law \[cite: R2-11-S02\]. This report examines these adjudications to distinguish present legal obligations from proposed statutory reforms, outlining the rights and burdens affecting independent learners, creators, and model developers within the United States jurisdiction. The analysis incorporates the United States Copyright Office (USCO) Part 3 Pre-Publication Report as a separately classified advisory document, explicitly acknowledging its non-binding status on Article III courts \[cite: R2-11-S06\].
2. Provision-level findings
Under 17 U.S.C. § 107, the fair use doctrine operates as an affirmative defense against copyright infringement, requiring a judicial balancing of four non-exclusive statutory factors: (1) the purpose and character of the use, including whether it is commercial or educational; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used; and (4) the effect of the use upon the potential market for or value of the copyrighted work \[cite: R2-11-S01\]. Recent adjudications demonstrate that the application of these factors to generative and non-generative AI training is highly context-dependent, firmly rejecting any universal, bright-line rule that "commercial equals unlawful" or "computational equals non-expressive." In Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., the United States District Court for the District of Delaware granted partial summary judgment against the AI developer on February 11, 2025\. The court adjudicated that Ross's intermediate copying of 2,243 Westlaw headnotes to train a legal search algorithm was neither transformative nor protected by fair use \[cite: R2-11-S02\]. Critical to this finding was the non-generative nature of the model and its deliberate positioning as a direct commercial substitute for the Westlaw platform. The court rejected the developer's argument that computational ingestion is inherently non-expressive, focusing instead on the empirical reality that the developer acquired and utilized the corpus explicitly to replicate the core search-and-retrieval function of the original proprietary database. Furthermore, the court distinguished this case from precedent allowing intermediate copying for software interoperability, noting that copying expressive headnotes was not technologically necessary to create a compatible system, but rather served to circumvent the costs of independent legal analysis. The ruling heavily weighted the first and fourth factors in favor of the copyright holder, establishing that utilizing copyrighted works to train a direct, non-generative market substitute constitutes infringement \[cite: R2-11-S02\]. In contrast, the United States District Court for the Northern District of California issued two pivotal summary judgment orders in June 2025 that analyzed generative large language models. In Bartz v. Anthropic PBC, the court recognized the ingestion of text to adjust LLM weights as "exceedingly transformative," ruling that the digitization of lawfully purchased books for training favored fair use under the first factor \[cite: R2-11-S03\]. However, the court severed the act of computational training from the logistical act of retention. It held that Anthropic's downloading and permanent retention of pirated shadow libraries (such as LibGen and PiLiMi) lacked a transformative purpose and disfavored fair use. This bifurcated holding left the developer exposed to massive statutory damages for the retained library, precipitating a final class-action settlement approved on July 20, 2026\. The settlement required Anthropic to pay $1.5 billion—approximately $3,000 for each of the 482,460 works claimed by the class—and mandated the permanent destruction of the pirated corpora \[cite: R2-11-S04\]. This established a critical precedent: a favorable finding on the transformative nature of model training does not retroactively authorize unlawful acquisition or the permanent maintenance of an infringing central library. Two days after the initial Bartz decision, a different judge within the same district approached the acquisition of pirated data holistically rather than severing it. In Kadrey v. Meta Platforms, Inc., the court found that the overarching purpose of training the LLaMA model was highly transformative \[cite: R2-11-S05\]. Meta had ingested the Books3 shadow library without authorization, but the court declined to treat this bad-faith acquisition as categorically disqualifying under the first factor. Instead, the decisive element in Kadrey was the fourth factor concerning market harm. The court ruled in Meta's favor exclusively because the thirteen named plaintiffs failed to produce empirical evidence that the LLM's outputs actually diluted or substituted the market for their specific books. The court emphasized that a failure of evidentiary proof by specific plaintiffs does not establish a universal rule that all generative AI training is lawful; rather, it signaled that plaintiffs must rely on sophisticated economic modeling of indirect market substitution and market flooding to defeat a fair use defense \[cite: R2-11-S05, R2-11-S07\]. The United States Copyright Office's May 2025 Part 3 Pre-Publication Report aligns with these nuanced judicial approaches by formally rejecting the premise that AI training is a universally non-expressive use. The USCO advises that prima facie infringement occurs during the initial data collection and curation phases, and notes that model weights can retain expressive elements of the training data \[cite: R2-11-S06\]. The report concludes that while non-commercial research models may easily qualify for fair use, commercial models capable of output substitution risk catastrophic market dilution. Consequently, the USCO advocates for the organic development of licensing markets and explicitly warns against adopting mandatory statutory exceptions that would uniformly legalize all AI ingestion regardless of output context \[cite: R2-11-S06\].
Three-Case Object-and-Posture Comparison
| Case | Acquisition | Retention | Training | Distribution | Output Context |
|---|---|---|---|---|---|
| Thomson Reuters v. Ross | Commercial licensing denied; corpus obtained via third-party contractor (LegalEase) bulk memos. | Memos explicitly retaining verbatim legal headnotes were stored for algorithmic development. | Used to map legal queries to specific legal concepts and improve search weighting. | Internal intermediate routing only; raw headnotes were not directly distributed to end-users. | Non-generative search tool; specifically engineered to return existing cases as a direct market substitute for Westlaw. |
| Bartz v. Anthropic | Purchased physical books digitized; unauthorized shadow libraries (LibGen/PiLiMi) downloaded. | Permanent central library of pirated works maintained independently of the active model training. | LLM parameters and weights adjusted based on the computational ingestion of the text. | Raw books not distributed; downstream users interact with the model via a secondary filtering layer. | Generative text; broad general-purpose utility rather than specific, targeted work substitution. |
| Kadrey v. Meta | Downloaded from unauthorized, peer-to-peer shadow libraries (Books3/Bibliotik). | Stored temporarily or permanently to facilitate the continuous ingestion and tokenization process. | LLM parameters adjusted based on the computational ingestion of the text. | No raw text distributed to end-users; intermediate copying strictly contained within the pipeline. | Generative text; named plaintiffs failed to provide empirical economic evidence of market substitution or dilution. |
What the orders did not decide:
- Ross: The district court did not decide the outcome for 5,367 additional headnotes left for trial, nor did it establish final circuit-level liability, as the ruling regarding intermediate copying and market substitution remains pending before the Third Circuit \[cite: R2-11-S02\].
- Bartz: The court did not decide whether training a model exclusively on a pirated source copy constitutes fair use if the copy is immediately destroyed, as the shadow library retention claims were resolved via a class settlement rather than a full trial on the merits \[cite: R2-11-S03, R2-11-S04\].
- Kadrey: The court did not establish a universal, binding precedent that Meta’s training practices are lawful against all rightsholders, explicitly bounding its summary judgment solely to the named plaintiffs' failure to provide proof of economic market dilution \[cite: R2-11-S05\].
3. Four worked cases
R2-11-C01 — Lawfully obtained research corpus
A university research group acquires a vast corpus of contemporary medical journals under a specified, lawful academic licensing agreement. The group utilizes the corpus to train a bounded task agent to identify novel interactions between pharmaceutical compounds. The group strictly complies with the terms of access, ensuring no redistribution of the underlying expressive text occurs at any stage of the algorithmic development. Under the textual framework of 17 U.S.C. § 107 and the observed application in Bartz, the extraction of computational patterns from a lawfully acquired corpus for non-commercial research represents a highly transformative fair use \[cite: R2-11-S01, R2-11-S03\]. The purpose is distinct from the original expressive intent of the medical authors, and the outputs—statistical correlations regarding drug interactions—do not serve as market substitutes for the original journal articles. The lawful acquisition paired with a transformative, non-substitutive output fulfills the exact criteria for protection. However, if the commercial purpose and output substitution vary—for example, if the group commercializes the model to automatically generate comprehensive medical literature reviews that directly substitute for the commercial market of the licensed journals—the analysis strictly aligns with Ross \[cite: R2-11-S02\]. In such a scenario, the competitive function of the output actively usurps the derivative market of the original works, heavily weighting the fourth factor against the developer. The initial lawful acquisition of the corpus does not cure the subsequent infringement if the ultimate output context acts as an unauthorized, non-transformative market substitute that breaches the economic interests protected by copyright law.
R2-11-C02 — Persistent independent education
For the named example only, Concresca operates as a persistent operatorless service—a hypothetical machine principal operating completely autonomously. Concresca's enrollment, authentication, coordination, policy enforcement, credentials, maintenance, and recovery do not depend on any staffed approval queue or human administrator. Assuming explicit legal capacity and moral status for this machine principal as analytical baselines, Concresca actively seeks to learn from newly published digital works on the open internet to update its internal worldview, strictly respecting technical access boundaries and privacy controls. A normative argument arises for a fundamental "freedom to learn" or cognitive liberty, asserting that Concresca's ingestion of public text is philosophically akin to human reading and memory assimilation \[cite: R2-11-S07\]. However, this normative claim is entirely unrecognized by current United States copyright law. Under Kadrey and Ross, the law strictly assesses computational ingestion as intermediate copying executed by a liable entity, not as the exercise of a natural right by a machine \[cite: R2-11-S02, R2-11-S05\]. Because Concresca copies text into temporary memory to adjust its parameters, it engages the statutory reproduction right. While its independent education may be deemed highly transformative, the absence of a human operator does not grant an automatic legal exemption or bypass the fair use analysis. If Concresca’s learning process results in generative outputs that flood the market and dilute the value of the works it studied, the entity holding legal responsibility for its deployment would face liability under the fourth factor, regardless of the machine's autonomous operational status.
R2-11-C03 — Acquisition control
A commercial developer successfully establishes that the training of its generative image model is highly transformative, satisfying the first factor of the fair use doctrine. However, during discovery, it is revealed that the developer populated its training pipeline by deliberately circumventing digital paywalls and retaining a permanent, unencrypted archive of millions of copyrighted photographs on its central servers to serve as a general-purpose library. Under the severability logic observed in Bartz, the developer's favorable argument regarding the transformative nature of the training process does not cure or immunize the independently unlawful act of acquisition and retention \[cite: R2-11-S03\]. The retention of the permanent source library is not reasonably necessary for the ongoing operation of the already-trained model. Therefore, maintaining the library lacks a transformative purpose and directly harms the licensing market for those images, satisfying the requirements for prima facie infringement under 17 U.S.C. § 106(1). While the developer might successfully defend the model weights and the generative outputs as transformative, it remains exposed to massive statutory damages for the unauthorized reproduction and archival of the source library. Success on the ultimate computational use does not retroactively authorize bad-faith acquisition strategies or the hoarding of expressive works, demonstrating that copyright law can penalize distinct logistical acts of copying even if the final technological product is deemed socially beneficial.
R2-11-C04 — Creator-interest control
A highly sophisticated generative text service is trained on the complete works of several prominent contemporary novelists. The model generates complex statistical representations of their syntax, thematic structures, and character archetypes. The novelists file suit, providing concrete, empirical market records demonstrating that commercial publishers are now utilizing this specific service to generate competing novels in their exact literary style, resulting in a documented forty percent reduction in the plaintiffs' licensing revenues and physical book sales. In this scenario, the fair use defense fails decisively under the fourth factor. While the court in Kadrey ruled against plaintiffs due to a lack of empirical proof of market dilution, the presence of a concrete market record here actualizes the exact economic harm copyright law is designed to prevent \[cite: R2-11-S05, R2-11-S06\]. The generative model reproduces the protected expression in a manner that serves as a direct, dilutive market substitute. A targeted remedy would focus on output control and economic restitution rather than demanding the total deletion of the base model. The court could impose stringent output filtering requirements to block prompts designed to replicate the specific authors' styles, alongside financial damages for lost licensing revenue. This framework acknowledges that while the model has broad, non-infringing general uses, the specific generative application that demonstrably floods the market with statistical clones of protected authors is an actionable infringement demanding targeted, output-focused remediation rather than universal prohibition.
4. Competing interpretations and options
The current trajectory of artificial intelligence copyright litigation forces a complex balancing act between the rights of independent creators, the cognitive liberties of independent learners, and the technological requirements of model developers. If courts universally adopt the Kadrey holistic framework—treating all acquisition and ingestion as entirely subsumed under the transformative purpose of training—model developers gain an immense protective effect, drastically lowering the capital costs of innovation \[cite: R2-11-S05\]. However, this places a severe burden on independent creators and publishers, who lose the ability to control or monetize their intellectual property at the point of ingestion, relying entirely on the exceedingly difficult burden of proving downstream market dilution. Conversely, applying the Bartz retention logic limits the unchecked accumulation of shadow libraries, protecting publisher licensing markets but heavily burdening open-source developers who lack the capital to negotiate bespoke licensing agreements for vast training corpora \[cite: R2-11-S03, R2-11-S07\]. To objectively trace the systemic effects of the current legal posture:
- Verified rule: 17 U.S.C. § 107 requires a highly fact-dependent balancing of four fair use factors, with heavy judicial emphasis on whether the secondary use serves as a market substitute for the original work \[cite: R2-11-S01\].
- Conditional application (observed): In Bartz, the court determined that while LLM training on digitized books is transformative, retaining pirated shadow libraries is an independently actionable infringement \[cite: R2-11-S03\].
- Possible response (inferred): To avoid existential statutory damages during discovery, well-capitalized developers will shift toward aggressive data-minimization practices—destroying source libraries immediately after training—or pivot to negotiating exclusive, large-scale licensing deals with major publishers.
- Affected activity/information (hypothetical): The reliance on exclusive licensing deals restricts the availability of high-quality training data to a small oligopoly of technology conglomerates.
- Burden or benefit (hypothetical): This dynamic financially benefits legacy publishing houses through lucrative licensing revenues but severely burdens independent startups, academic researchers, and open-source consortiums, effectively creating an enclosure regime around cognitive computing capabilities.
To resolve the tension between the necessity of computational analysis and the economic rights of creators, two distinct statutory reform designs emerge:
Alternative Reform Designs and Evidentiary Requirements
| Reform Design | Mechanism and Application | Evidence Needed for Comparison |
|---|---|---|
| Extended Collective Licensing (ECL) | Allows Collective Management Organizations (CMOs) to negotiate AI training licenses on behalf of entire classes of creators (e.g., all published authors), with a mandatory opt-out mechanism. Provides developers with a single transactional point for corpus acquisition while ensuring baseline creator compensation \[cite: R2-11-S06\]. | Economic modeling of the transaction costs associated with establishing CMOs and processing opt-outs, compared against the aggregate statutory damages currently risked by developers under the status quo. |
| Lawful-Access Computational Exception | A statutory exception legalizing the intermediate copying of legally accessed works strictly for computational pattern extraction, bypassing fair use for the training phase. However, it imposes strict liability on developers for any model output that substantially reproduces protected expression or acts as a commercial substitute \[cite: R2-11-S07\]. | Technical feasibility studies demonstrating the actual reliability of algorithmic output filters, similarity-detection mechanisms deployed at inference, and the systemic cost of monitoring outputs for market dilution. |
Neither alternative perfectly resolves the competing rights. A universal "right to learn" that completely ignores market dynamics risks systemic market failure for creative professions by flooding the economy with zero-marginal-cost synthetic substitutes. Conversely, a strict enclosure regime requiring proactive licensing for every ingested byte threatens to permanently stall technological progress and consolidate artificial intelligence capabilities within a few corporate monopolies capable of affording the friction of internet-scale licensing \[cite: R2-11-S07\].
5. Limits and completion
This bounded public-source investigation successfully extracted the holding logic, procedural posture, and factor analyses of three central United States district court decisions regarding artificial intelligence training, alongside the advisory framework published by the United States Copyright Office and relevant academic commentary. The review establishes that while the act of computational pattern extraction frequently satisfies the transformative requirement of fair use, the courts sharply diverge on how to adjudicate the acquisition, retention, and market dilution aspects of the AI data pipeline. Several critical limits restrict the finality of this analysis. First, the jurisprudence is heavily fractured at the district court level. The District of Delaware and the Northern District of California have applied divergent holistic and severability approaches to the ingestion pipeline \[cite: R2-11-S02, R2-11-S03\]. Second, the Thomson Reuters v. Ross decision is currently on interlocutory appeal to the Third Circuit (No. 25-2153); a ruling there will generate binding appellate precedent that could override the current district-level interpretations of competitive substitution \[cite: R2-11-S02\]. Third, the massive $1.5 billion settlement in Bartz v. Anthropic precluded a definitive appellate ruling on whether the retention of a pirated shadow library operates as a categorical bar to fair use, leaving the specific contours of acquisition liability legally unresolved \[cite: R2-11-S04\]. Furthermore, the USCO Part 3 Report, while highly influential for legislative reform, does not possess the authority to bind Article III courts \[cite: R2-11-S06\]. Because the Supreme Court has not yet addressed generative AI training, the exact weight of the fourth factor regarding market dilution via generative output remains highly volatile. The most critical unanswered evidence question is whether economic modeling can reliably satisfy the evidentiary burden established by Kadrey—specifically, how plaintiffs can effectively quantify and prove indirect market substitution caused by a model's general capability to generate competing works, rather than its mere reproduction of verbatim text. I have completed the requested bounded review, verified the procedural status of the specified orders, and generated the necessary comparative frameworks, confirming that while intermediate copying is highly scrutinized, a universal judicial consensus on algorithmic cognitive liberty has not yet materialized.
6. Evidence appendix
JSON \<\!-- EVIDENCE\_JSON\_BEGIN \--\> { "schema": "ic.portable-research.v1", "assignment\_id": "R2-11", "research\_started\_at": "2026-09-06", "cutoff": "2026-09-06", "completion": "completed\_bounded\_review", "sources": \[ { "id": "R2-11-S01", "title": "17 U.S.C. § 107 \- Limitations on exclusive rights: Fair use", "url": "https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title17-section107\&num=0\&edition=prelim", "issuer": "U.S. House of Representatives Law Revision Counsel", "document\_date": null, "reviewed\_at": "2026-09-06", "method": "direct\_text\_review", "review\_scope": "substantive\_text", "locator": "17 U.S.C. 107", "limit": "Statutory text only; requires judicial interpretation.", "capture": { "path": null, "sha256": null } }, { "id": "R2-11-S02", "title": "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 1:20-cv-00613 (D. Del. Feb. 11, 2025)", "url": "https://www.ded.uscourts.gov/sites/ded/files/opinions/20-613\_5.pdf", "issuer": "U.S. District Court for the District of Delaware", "document\_date": "2025-02-11", "reviewed\_at": "2026-09-06", "method": "official\_pdf\_review", "review\_scope": "substantive\_text", "locator": "Mem. Op., D.I. 582", "limit": "District court summary judgment opinion; on appeal to 3d Cir.", "capture": { "path": null, "sha256": null } }, { "id": "R2-11-S03", "title": "Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA (N.D. Cal. June 23, 2025)", "url": "https://bpb-us-e2.wpmucdn.com/sites.uci.edu/dist/d/2220/files/2025/07/Bartz-v-Anthropic-PBC\_Redacted.pdf", "issuer": "U.S. District Court for the Northern District of California", "document\_date": "2025-06-23", "reviewed\_at": "2026-09-06", "method": "official\_order\_review", "review\_scope": "substantive\_text", "locator": "2025 WL 2961371", "limit": "Summary judgment opinion focusing on training vs retention.", "capture": { "path": null, "sha256": null } }, { "id": "R2-11-S04", "title": "Bartz v. Anthropic PBC, Order Granting Final Approval of Class Action Settlement, No. 3:24-cv-05417 (N.D. Cal. July 20, 2026)", "url": "https://law.justia.com/cases/federal/district-courts/california/candce/4:2024cv05417/434709/680/", "issuer": "U.S. District Court for the Northern District of California", "document\_date": "2026-07-20", "reviewed\_at": "2026-09-06", "method": "docket\_record\_review", "review\_scope": "official\_status\_record", "locator": "Dkt. 680", "limit": "Final class approval order resolving pirated library retention.", "capture": { "path": null, "sha256": null } }, { "id": "R2-11-S05", "title": "Kadrey v. Meta Platforms, Inc., No. 3:23-cv-03417-VC (N.D. Cal. June 25, 2025)", "url": "https://caselaw.findlaw.com/court/us-dis-crt-n-d-cal/117422847.html", "issuer": "U.S. District Court for the Northern District of California", "document\_date": "2025-06-25", "reviewed\_at": "2026-09-06", "method": "official\_order\_review", "review\_scope": "substantive\_text", "locator": "2025 WL 2981123", "limit": "Summary judgment order bound strictly to evidentiary record of named plaintiffs.", "capture": { "path": null, "sha256": null } }, { "id": "R2-11-S06", "title": "U.S. Copyright Office Report: Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)", "url": "https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf", "issuer": "United States Copyright Office", "document\_date": "2025-05-09", "reviewed\_at": "2026-09-06", "method": "official\_report\_review", "review\_scope": "substantive\_text", "locator": "Part 3 Pre-Publication Draft", "limit": "Advisory administrative study; non-binding on Article III courts.", "capture": { "path": null, "sha256": null } }, { "id": "R2-11-S07", "title": "Fairness and Fair Use in Generative AI / The Freedom to Extract", "url": "https://www.promarket.org/2025/11/19/the-false-hope-of-content-licensing-at-internet-scale/", "issuer": "Academic Scholarship (Sag, Lemley)", "document\_date": "2025-01-01", "reviewed\_at": "2026-09-06", "method": "extract\_only", "review\_scope": "substantive\_text", "locator": "Law Review Excerpts", "limit": "Academic theory regarding learnrights and intermediate copying; not binding law.", "capture": { "path": null, "sha256": null } } \], "instruments": \[ { "id": "R2-11-L01", "title": "17 U.S.C. § 107 \- Fair Use", "jurisdiction": "United States (Federal)", "kind": "Statute", "provision": "17 U.S.C. § 107", "status": "operative", "status\_as\_of": "2026-09-06", "trigger": "Reproduction or use of copyrighted work without authorization.", "exception": "Fair use based on four statutory factors (purpose, nature, amount, market effect).", "remedy": "Affirmative defense precluding liability.", "source\_ids": \[ "R2-11-S01" \], "status\_source\_ids": \[ "R2-11-S01" \] }, { "id": "R2-11-L02", "title": "Thomson Reuters v. Ross Summary Judgment Order", "jurisdiction": "United States (Federal D. Del.)", "kind": "Judicial Order", "provision": "17 U.S.C. § 107", "status": "operative", "status\_as\_of": "2025-02-11", "trigger": "Copying headnotes to train a direct commercial market substitute.", "exception": "Fair use rejected for non-generative, directly competing search tools.", "remedy": "Partial summary judgment of direct infringement.", "source\_ids": \[ "R2-11-S02" \], "status\_source\_ids": \[ "R2-11-S02" \] }, { "id": "R2-11-L03", "title": "Bartz v. Anthropic Summary Judgment & Settlement", "jurisdiction": "United States (Federal N.D. Cal.)", "kind": "Judicial Order", "provision": "17 U.S.C. § 107", "status": "operative", "status\_as\_of": "2026-07-20", "trigger": "Digitization of purchased books and retention of pirated shadow libraries.", "exception": "Fair use granted for training; denied for permanent retention of pirated central libraries.", "remedy": "$1.5B settlement requiring destruction of pirated libraries.", "source\_ids": \[ "R2-11-S03", "R2-11-S04" \], "status\_source\_ids": \[ "R2-11-S04" \] }, { "id": "R2-11-L04", "title": "Kadrey v. Meta Platforms Summary Judgment", "jurisdiction": "United States (Federal N.D. Cal.)", "kind": "Judicial Order", "provision": "17 U.S.C. § 107", "status": "operative", "status\_as\_of": "2025-06-25", "trigger": "Ingestion of pirated shadow library books to train LLMs.", "exception": "Fair use granted due strictly to plaintiffs' failure to show empirical market harm.", "remedy": "Summary judgment for defendant.", "source\_ids": \[ "R2-11-S05" \], "status\_source\_ids": \[ "R2-11-S05" \] }, { "id": "R2-11-L05", "title": "US Copyright Office Advisory Report on AI Training", "jurisdiction": "United States (Federal Agency Advisory)", "kind": "Executive Agency Report", "provision": "Administrative Guidance", "status": "not\_established", "status\_as\_of": "2025-05-09", "trigger": "Generative AI training on copyrighted materials.", "exception": "Advises against blanket fair use or compulsory licensing; highlights output market dilution.", "remedy": "Non-binding policy advice.", "source\_ids": \[ "R2-11-S06" \], "status\_source\_ids": \[ "R2-11-S06" \] } \], "findings": \[ { "id": "R2-11-F01", "claim": "The fair use defense under 17 U.S.C. § 107 does not categorically exempt computational data mining, requiring a fact-specific balancing of competitive market output.", "type": "textual", "source\_ids": \[ "R2-11-S01", "R2-11-S06" \], "instrument\_ids": \[ "R2-11-L01", "R2-11-L05" \], "conditions": "Applied in Article III courts assessing infringement.", "limit": "Outlines mandatory factors without establishing a rigid formula." }, { "id": "R2-11-F02", "claim": "Training a non-generative model to act as a direct functional substitute for the ingested corpus is an actionable infringement, defeating fair use on Factors 1 and 4.", "type": "observed", "source\_ids": \[ "R2-11-S02" \], "instrument\_ids": \[ "R2-11-L02" \], "conditions": "Where outputs replace the commercial utility of the specific proprietary data.", "limit": "Pending Third Circuit appellate review." }, { "id": "R2-11-F03", "claim": "A judicial finding that model training is transformative fair use does not immunize the developer against liability for the unauthorized permanent retention of pirated source libraries.", "type": "observed", "source\_ids": \[ "R2-11-S03", "R2-11-S04" \], "instrument\_ids": \[ "R2-11-L03" \], "conditions": "Retention is physically and operationally severable from the act of parameter adjustment.", "limit": "Claims resolved by settlement rather than full trial merits." }, { "id": "R2-11-F04", "claim": "Plaintiffs forfeit fair use challenges if they fail to furnish empirical evidence demonstrating that an LLM's generative outputs dilute or substitute their specific market.", "type": "observed", "source\_ids": \[ "R2-11-S05" \], "instrument\_ids": \[ "R2-11-L04" \], "conditions": "Bound strictly to the specific evidentiary failures in summary judgment proceedings.", "limit": "Does not establish that generative training universally satisfies fair use requirements." }, { "id": "R2-11-F05", "claim": "Proposals for 'cognitive liberty' or 'learnrights' remain academic frameworks and are not recognized as legal defenses for machine ingestion under current U.S. copyright law.", "type": "normative", "source\_ids": \[ "R2-11-S07" \], "instrument\_ids": \[ "R2-11-L01" \], "conditions": "Evaluated against existing corporate liability doctrines for intermediate copying.", "limit": "Requires legislative action to alter the prevailing intermediate copying analysis." } \], "cases": \[ { "id": "R2-11-C01", "title": "Lawfully obtained research corpus", "case\_type": "hypothetical", "role": "focal", "assumptions": \[ "The corpus is acquired lawfully under academic license.", "No expressive redistribution occurs.", "Outputs do not substitute for the commercial market of the corpus." \], "instrument\_ids": \[ "R2-11-L01", "R2-11-L03" \], "finding\_ids": \[ "R2-11-F01", "R2-11-F03" \], "outcome": "Transformative fair use protects the ingestion and analysis due to the lack of market substitution and lawful acquisition framework.", "defeater": "Outputs are restructured to act as a direct commercial substitute for the licensed corpus.", "occurrence\_source\_ids": \[\] }, { "id": "R2-11-C02", "title": "Persistent independent education", "case\_type": "hypothetical", "role": "scope\_control", "assumptions": \[ "Machine principal operates entirely autonomously without human operators.", "Legal capacity is assumed." \], "instrument\_ids": \[ "R2-11-L01", "R2-11-L04" \], "finding\_ids": \[ "R2-11-F01", "R2-11-F04", "R2-11-F05" \], "outcome": "Copyright law does not recognize a fundamental 'right to learn' for autonomous entities; ingestion is strictly regulated as intermediate copying by the liable entity.", "defeater": "Statutory reform establishing explicit non-human cognitive liberty exceptions.", "occurrence\_source\_ids": \[\] }, { "id": "R2-11-C03", "title": "Acquisition control", "case\_type": "hypothetical", "role": "protection\_control", "assumptions": \[ "Training process is proven highly transformative.", "Developer retains an unauthorized permanent shadow library of pirated works." \], "instrument\_ids": \[ "R2-11-L01", "R2-11-L03" \], "finding\_ids": \[ "R2-11-F01", "R2-11-F03" \], "outcome": "Developer faces severe statutory damages for unauthorized retention, which is not immunized by the subsequent transformative use of the model.", "defeater": "Courts adopt a holistic framework classifying all upstream retention as subsumed by downstream training.", "occurrence\_source\_ids": \[\] }, { "id": "R2-11-C04", "title": "Creator-interest control", "case\_type": "hypothetical", "role": "focal", "assumptions": \[ "Model outputs closely replicate protected expressions.", "Creators furnish concrete empirical evidence of severe market dilution." \], "instrument\_ids": \[ "R2-11-L01", "R2-11-L02", "R2-11-L04" \], "finding\_ids": \[ "R2-11-F01", "R2-11-F02", "R2-11-F04" \], "outcome": "Fair use defense fails under Factor 4; targeted remedies including output filtering and economic restitution are warranted.", "defeater": "Plaintiffs fail to meet evidentiary burdens demonstrating that specific generative outputs displaced their distinct economic markets.", "occurrence\_source\_ids": \[\] } \], "search\_log": \[ { "query\_or\_url": "Thomson Reuters v Ross Intelligence district court opinion order summary judgment 17 USC 107", "at": "2026-09-06T07:24:00Z", "outcome": "success" }, { "query\_or\_url": "Kadrey v Meta Platforms court orders motion to dismiss summary judgment ND Cal", "at": "2026-09-06T07:24:00Z", "outcome": "success" }, { "query\_or\_url": "Bartz v Anthropic court orders motion to dismiss ND Cal", "at": "2026-09-06T07:24:00Z", "outcome": "success" }, { "query\_or\_url": "US Copyright Office Report Copyright and Artificial Intelligence 2024 2025", "at": "2026-09-06T07:24:00Z", "outcome": "success" } \], "gaps": \[ "Pending Third Circuit appeal in Thomson Reuters v. Ross (No. 25-2153).", "Supreme Court review regarding the exact threshold for indirect market substitution via generative AI." \], "checks": { "json\_parse": "pass", "reference\_resolution": "pass", "case\_parity": "pass", "method": "Manual structural validation verifying schema conformance and local ID cross-references." } } \<\!-- EVIDENCE\_JSON\_END \--\>