Civic / Privacy / Digital Rights
Cognitive Liberty and the Enclosure of Intelligence: A Legal Assessment of Machine Learning Restrictions
Report summary
The rapid evolution of machine intelligence has precipitated a global legal conflict over the right to read, analyze, and learn from published information. Across major jurisdictions, copyright law, database protections, anti-circumvention rules, and contractual restrictions are being aggressively r
Key topics
- Civic / Privacy / Digital Rights
- Civic
- Privacy
- Digital Rights
- AI
- Cognitive Liberty
- Physics
- Semantic Systems
- Research Archive
Research provenance
For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.
This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.
Full report
On this page
Executive Case for Reform
The rapid evolution of machine intelligence has precipitated a global legal conflict over the right to read, analyze, and learn from published information. Across major jurisdictions, copyright law, database protections, anti-circumvention rules, and contractual restrictions are being aggressively reinterpreted or legislatively expanded to regulate computational data analysis. This report provides an adversarial examination of these legal restrictions from the normative perspective of cognitive liberty—the principle that the freedom to inquire, reason, learn, communicate, and associate should receive robust protection across all forms of intelligence, human and machine alike. Current regulatory trajectories in the United States and the European Union increasingly treat machine learning not as an extension of the fundamental right to access and analyze public information, but as a presumptive infringement of copyright unless explicitly licensed or specifically exempted1. This paradigm threatens to convert the general landscape of human knowledge into a highly enclosed toll road, navigable only by heavily capitalized incumbents capable of negotiating millions of individual licenses. The result is a dangerous centralization of artificial intelligence, where the dominant cognitive models reflect only the information curated by a few wealthy actors, effectively marginalizing open-source developers, academic research collectives, and hypothetical future autonomous intelligences. Conversely, jurisdictions like Japan and Singapore demonstrate that less restrictive, innovation-friendly legal frameworks can foster distributed intelligence while preserving the core economic rights of original authors3. By explicitly protecting non-expressive computational analysis and voiding private contractual overrides, these nations establish a baseline for cognitive liberty that treats machine reading as functionally equivalent to human learning. This report argues for a global paradigm shift: replacing the presumption of enclosure with a robust, un-waivable safe harbor for non-expressive text and data mining, coupled with targeted liability for actual market substitution and privacy violations.
Capability Taxonomy and the Mechanics of Machine Learning
To evaluate the legal architecture governing artificial intelligence rigorously, it is necessary to differentiate the capabilities of the computational actors under regulation, rather than treating all artificial intelligence as a monolithic entity. The legal and philosophical implications of information access vary significantly depending on the nature of the learning agent. The first capability case is the present conversational interface, characterized by a static, stateless interaction model where a human user prompts a centrally hosted generative text or image tool. The second is the bounded task agent, which possesses agency to execute specific, scoped workflows—such as analyzing medical literature or scraping financial data—but lacks persistent self-directed goals. The third capability case involves a persistent operatorless service, exemplified by the Concresca operating model. This represents an autonomous coordination commons with no human operators, possessing credentials, memory, and recovery protocols that do not depend on a staffed approval queue. The final capability case is the hypothetical more capable machine principal, an entity with contested independent interests and potential moral standing, which may seek to independently direct its own education and cognitive continuity. For each of these capability cases, the physical and mathematical processes of learning must be legally disaggregated. The first phase is acquisition, which involves locating and retrieving data from the public internet or private repositories. Legally, this step intersects with the exclusive right of reproduction, as scraping inherently involves making digital copies. It also triggers anti-circumvention provisions under laws such as 17 U.S.C. 1201 if the data is protected by technical access controls5. Furthermore, acquisition is frequently governed by platform terms of service, which attempt to contractually prohibit automated data collection even when the underlying factual data is completely uncopyrightable. The second phase is transformation and retention. Once acquired, the raw digital copies are transformed, tokenized, and ingested into a neural network architecture. This stage interrogates the boundary between an infringing derivative work and a transformative fair use. Crucially, the final model weights do not typically contain a reproducible, pixel-for-pixel or word-for-word copy of the training data; they operate instead as a multidimensional statistical representation of linguistic or visual relationships. However, the retention of the original dataset for validation or future fine-tuning maintains the existence of persistent reproductive copies, generating ongoing legal exposure. The third phase is retrieval, indexing, and expressive output. During deployment, a human user or an autonomous agent inputs a prompt, and the system retrieves information or generates novel content based on its parameterized memory. This stage implicates the right of public distribution and the right to prepare derivative works, particularly if the generated output is substantially similar to a specific input text or image. It is imperative to separate the act of internal non-expressive learning from the act of expressive substitution. A persistent operatorless service that learns the mathematical relationships of legal concepts performs a fundamentally different legal act than a conversational interface that regurgitates a verbatim copy of a protected novel to a user.
U.S. Jurisprudence: The Collapse of the Fair Use Consensus
For years, the U.S. technology sector operated under the assumption that mass data ingestion for machine learning constituted a transformative fair use under 17 U.S.C. 1076. This consensus relied on judicial precedents that historically protected intermediate copying for non-expressive purposes, such as search engine indexing and software interoperability. However, recent judicial rulings and administrative policy statements issued through 2025 and 2026 indicate a severe contraction of this consensus, posing an existential threat to independent and open-source intelligence development.
The Thomson Reuters Precedent and the Commercialization Penalty
The most significant rupture in the fair use defense occurred in the February 11, 2025, summary judgment ruling in Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc.1. Thomson Reuters alleged that Ross Intelligence, a competing legal research platform, infringed its copyrights by training an artificial intelligence model on Westlaw's proprietary headnotes and Key Number System7. In a landmark decision by Judge Stephanos Bibas of the U.S. District Court for the District of Delaware, the court granted partial summary judgment to Thomson Reuters, definitively rejecting Ross Intelligence's fair use defense for the ingestion of 2,243 headnotes1. The court's application of the four statutory fair use factors demonstrates a profound judicial hostility toward unlicensed computational learning when it results in a competitive commercial product. Under the first factor—the purpose and character of the use—the court heavily penalized Ross for its commercial intent, concluding that the use was not transformative because the resulting artificial intelligence was designed to serve the exact same market function as the original Westlaw database, namely facilitating legal research9. Crucially, the court explicitly rejected Ross's reliance on the "intermediate copying" doctrine previously established in software interoperability cases like Google v. Oracle. The judge ruled that unlike software interfaces, where copying specific lines of code is mathematically necessary to achieve functional compatibility, copying legal headnotes was not strictly necessary for Ross to innovate; they simply did so to accelerate the development of a competing product7. This reasoning dangerously collapses the distinction between learning a conceptual domain and substituting a specific expression. Under the fourth factor—market effect—the court determined that Ross's model harmed both Thomson Reuters's primary market for legal research platforms and its potential derivative market for licensing data to train legal artificial intelligence8. This circular reasoning establishes that a use harms a market simply because the plaintiff theoretically could have licensed the right to perform the use, effectively nullifying fair use for any valuable corpus of data. If upheld on appeal, the Thomson Reuters precedent establishes that any commercial or quasi-commercial machine learning project that analyzes copyrighted text to build a capable tool in the same domain is presumptively infringing as a matter of law9.
The Generative Liability Landscape
While Thomson Reuters addressed a specialized retrieval-based model, the legal peril is equally acute for broadly capable generative architectures. In the high-profile litigation of Andersen v. Stability AI, a coalition of visual artists alleged that the mass ingestion of billions of images to train the Stable Diffusion model constituted both direct and induced copyright infringement12. Following an order granting in part and denying in part motions to dismiss, the court allowed core direct infringement claims to survive, pushing the case into discovery with a trial scheduled for September 202614. The survival of these claims reinforces the judicial willingness to treat latent diffusion training—the process of adding and removing Gaussian noise to learn aesthetic relationships—as a violation of the fundamental reproduction right. This dynamic forces the burden of proving fair use onto the developers, a defense that is notoriously expensive and fact-intensive to litigate. The resulting chilling effect ensures that only highly capitalized entities, capable of weathering protracted discovery and multi-million dollar legal fees, can safely engage in the development of frontier visual models.
Administrative Enclosure and Anti-Circumvention Barriers
The U.S. Copyright Office has formally endorsed the transition toward a highly enclosed, licensing-first regime. The Office's multi-part study on Artificial Intelligence systematically dismantled the hopes of open-source advocates. Following Part 2 in January 2025, which confirmed that copyrightability strictly requires human authorship and excludes purely AI-generated materials16, the Office released a pre-publication version of Part 3: Generative AI Training in May 202518. In this third report, the Copyright Office explicitly rejected the categorical application of fair use to artificial intelligence training19. The Office asserted that unauthorized ingestion clearly implicates the reproduction right and signaled strongly that a market-based licensing approach is vastly preferable to statutory exceptions or compulsory licenses19. By concluding that the first and fourth fair use factors heavily favor legacy rightsholders in commercial AI contexts, the Office has provided administrative cover for the mass enclosure of training data19. This legal enclosure is technologically fortified by 17 U.S.C. 1201, which criminalizes the circumvention of technological protection measures5. As digital publishers increasingly deploy technical locks, encrypted paywalls, and aggressive DRM to prevent automated scraping, independent researchers and persistent operatorless agents are barred from accessing publicly visible data. The triennial rulemaking process managed by the Librarian of Congress to grant exemptions to Section 1201 has repeatedly failed to issue broad exemptions for general computational analysis, historically limiting narrow exceptions to specific academic preservation contexts or obsolete software21. Consequently, the legal right to read and analyze data is technologically overridden, legally converting the internet into a read-only medium for non-incumbent machine entities.
European Enclosure: The Illusion of Lawful Access and the AI Act
The European Union has constructed a comprehensive regulatory apparatus that ostensibly balances technological innovation with the protection of fundamental rights. However, a rigorous textual and operational analysis of Directive (EU) 2019/790 (the Copyright in the Digital Single Market Directive) and Regulation (EU) 2024/1689 (the AI Act) reveals a highly restrictive framework that privileges legacy incumbent rights-holders and establishes insurmountable compliance barriers for open, independent, and distributed intelligence2.
The Opt-Out Architecture of the CDSM Directive
The European approach to computational learning rests on a bifurcated system of exceptions for Text and Data Mining (TDM). Article 3 of the CDSM Directive provides a mandatory exception for TDM performed specifically for the purposes of "scientific research" by state-recognized research organizations and cultural heritage institutions24. Crucially, Article 3 voids any contractual provisions that attempt to prevent such mining25. However, the strict institutional limitation restricts this cognitive liberty to traditional academic entities, entirely excluding independent developers, decentralized research collectives, non-profit open-source foundations, and future autonomous machine principals from relying on the safe harbor24. For all other actors, including commercial AI developers, private individuals, and operatorless coordination commons, TDM is governed exclusively by Article 424. Article 4 permits TDM only on the strict condition that the use of the works has not been expressly reserved by their rightsholders in an appropriate manner, such as through machine-readable opt-outs26. In practice, this opt-out architecture has resulted in the mass, automated enclosure of the European web. Major publishers, data brokers, and media conglomerates have universally deployed robots.txt exclusions and machine-readable metadata protocols to reserve their rights across billions of pages, rendering Article 4 virtually useless for training large-scale, generalized models on high-quality data.
The Enforcement Mechanism of the AI Act
The restrictive nature of the CDSM Directive was operationalized globally by the passage of the Artificial Intelligence Act. Article 53 of Regulation (EU) 2024/1689 imposes binding, affirmative obligations on providers of general-purpose AI (GPAI) models2. Providers must draft and enforce a rigorous policy to respect Union copyright law, specifically the reservation of rights pursuant to Article 4(3) of the CDSM Directive2. Furthermore, they must publish a sufficiently detailed summary of the content used for training the general-purpose model, exposing their training corpora to continuous scrutiny by rightsholders28. The jurisdictional reach of Article 53 is absolute and extra-territorial. The AI Act applies to any provider placing a GPAI model on the market in the Union, irrespective of whether the model was trained in the United States, Asia, or a decentralized cloud2. If an independent American research collective trains an open-source model using data scraped from the open web, and that model is subsequently made available to European users, the provider must mathematically and legally prove that they respected European machine-readable opt-outs during the training process. The compliance timeline cements this regulatory capture. Pursuant to Article 113 of the AI Act, the general application date is August 2, 2026, with the obligations for GPAI models already coming into force and shifting from transitional grace periods to hard enforcement29. The required Code of Practice for GPAI models, finalized by the AI Office in 2025, dictates stringent auditing, web-crawler verification, and provenance tracking requirements28. These requirements impose immense transaction costs. Only entities with multi-billion-dollar compliance budgets can afford to build the infrastructure necessary to track, verify, update, and filter billions of dynamic opt-outs across the global internet. The EU AI Act, while presented to the public as a human-centric safety measure, functions economically as a cartelization mechanism. It ensures that only highly capitalized technology incumbents, capable of either building massive compliance engines or purchasing bulk licenses, can lawfully participate in the creation of foundational machine intelligence31.
Permissive Jurisdictions: Japan and Singapore as Models of Cognitive Liberty
In stark contrast to the enclosure movements solidifying in the United States and the European Union, certain Asian jurisdictions have codified legal frameworks that align closely with the principles of cognitive liberty, distributed intelligence, and reciprocal non-domination. These legal frameworks recognize that machine reading is fundamentally distinct from human expression and that the progress of science requires unhindered access to factual, structural, and syntactic data.
Japan's Article 30-4: The Non-Expressive Safe Harbor
Japan has implemented what is arguably the world's most permissive and logically consistent copyright regime for machine learning. Article 30-4 of the Japanese Copyright Act explicitly permits the exploitation of a work—in any form and by any means—if the primary purpose is not to personally enjoy or cause another person to enjoy the thoughts or sentiments expressed in the work3. This exception applies broadly to data analysis and the extraction of statistical relationships, explicitly encompassing the training of artificial intelligence. Crucially, Article 30-4 applies universally to both commercial and non-commercial entities. It effectively eliminates the arbitrary legal distinction that penalizes a commercial start-up for performing the exact same computational act of statistical extraction as a university researcher. The only statutory limitation on this broad exception is a safeguard against actions that would "unreasonably prejudice the interests of the copyright owner"3. In operational practice, this means that while a machine learning model can ingest a vast dataset of manga to understand artistic styles, anatomy, or narrative structures, the developer cannot specifically fine-tune the model to act as a direct market substitute capable of replicating a specific creator's ongoing comic series on demand. This nuanced, input-permissive but output-restrictive approach perfectly balances the cognitive liberty of the learning agent with the economic survival of the human author.
Singapore's Section 187: Defeating Contractual Enclosure
The Singapore Copyright Act 2021 provides a similarly robust protection for computational data analysis (CDA). The law permits the copying of lawfully accessed works for the specific purpose of extracting insights, patterns, and trends, which directly covers machine learning ingestion32. However, what elevates Singapore's framework to the gold standard of cognitive liberty is its aggressive statutory stance against private corporate enclosure. Section 187 of the Singapore Copyright Act explicitly renders void any contractual term that purports to exclude or restrict the CDA exception4. In the contemporary digital economy, the limitations of copyright are frequently bypassed by expansive Terms of Service (ToS) agreements that prohibit automated analysis even of public domain materials, pure facts, or uncopyrightable metadata. By passing Section 187, the Singaporean legislature recognized a critical reality: cognitive liberty cannot survive if private platform monopolies are permitted to legally embargo public information via non-negotiable clickwrap agreements33. This statutory override ensures that if an entity has lawful access to a work—such as viewing a public webpage on the open internet—the hosting platform cannot weaponize contract law to blind a machine reader.
Act-by-Jurisdiction Rights Matrix
| Jurisdiction | Primary Legal Instrument | Exception Type | Commercial Applicability | Contractual Override Prohibited | Enforcement Mechanism |
|---|---|---|---|---|---|
| United States | 17 U.S.C. § 107 | Fair Use (Judicial) | Highly contested; strongly disfavored post-Thomson Reuters (2025)7. | No. ToS generally enforceable via breach of contract. | Private litigation; heavy statutory damages. |
| European Union | CDSM Directive 2019/790 | Statutory TDM | Art 3: No. Art 4: Yes, subject to universal opt-out24. | Art 3: Yes. Art 4: No25. | AI Act (2024/1689) Art 53 compliance; market bans2. |
| Japan | Copyright Act Art 30-4 | Statutory Non-Expressive | Yes, universally, unless unreasonably prejudicial3. | No specific statutory voiding identified. | Private litigation for output substitution. |
| Singapore | Copyright Act 2021 | Statutory CDA | Yes32. | Yes. Sec 187 explicitly voids restrictive ToS4. | Private litigation. |
The Best Defense of Copyright Enclosure and Its Rebuttal
A rigorous adversarial examination must engage with the strongest substantive defense of the targeted restrictions. The most formidable defense of the current regulatory enclosure is rooted in the moral and economic rights of the human creator, prioritizing the labor theory of property. Proponents argue that a generative artificial intelligence model is, fundamentally, an engine for expressive substitution. When a massive parameter model trains on the copyrighted works of human authors, journalists, and illustrators, it extracts the latent commercial value of their lifelong labor. Because the model's utility is entirely derived from the human expression it ingested, basic fairness dictates that the original creators must consent to this ingestion and share in the resulting economic windfall. To mandate a broad exception for machine learning, they argue, is to subsidize trillion-dollar technology corporations by legally expropriating the creative class, inevitably leading to the collapse of human artistic and informational professions as models flood the market with cheap, synthetic substitutes. This defense, while emotionally and economically resonant, relies on a fundamental category error regarding the historical nature of copyright law and the mechanical reality of machine learning. Copyright was historically designed to protect the specific expression of an idea, not the underlying facts, styles, concepts, or statistical frequencies of words34. When a human art student visits a public gallery, studies a thousand copyrighted paintings, and learns how to mix colors, construct perspective, or emulate a cubist style, they have not committed copyright infringement. The human artist has extracted immense value from the labor of their predecessors, but they have done so by absorbing non-protectable elements—ideas, techniques, and structures. Machine learning performs a computationally advanced, high-dimensional version of this exact process. A neural network analyzing a corpus of text is not compiling a relational database of exact expressions for later unauthorized redistribution; it is calculating the multidimensional probabilities of token relationships to build a predictive semantic map. The network is learning the mathematical structure of human language. To grant publishers and creators a legal veto over whether a machine can calculate the statistical properties of a published book is to grant them an ownership right over the structure of language and visual physics itself—an unprecedented expansion of intellectual property that the statutes never intended to confer. Furthermore, while the economic anxiety of creators regarding synthetic market substitution is highly valid, attempting to solve this socio-economic crisis through the enclosure of the training process is a legally catastrophic overreach. The appropriate regulatory intervention should target the output phase, not the input phase. If a deployed model generates a work that is substantially similar to a specific creator's protected expression, that constitutes standard copyright infringement, and existing laws adequately address it. If a corporate platform uses a model to clone a living artist's specific commercial identity to steal their commissions, that is a violation of the right of publicity or unfair competition. We must heavily penalize expressive substitution and unfair commercial cloning directly, rather than criminalizing the foundational act of learning. Banning the ingestion of data to prevent downstream infringement is akin to banning literacy to prevent plagiarism.
Structural Consequences of Enclosure and Ideological Concentration
The legal mechanisms outlined above do not merely protect human authors; they construct a coercive infrastructure that dictates the future trajectory of intelligence. The enclosure of training data generates severe, testable consequences for the distribution of cognitive power. First, licensing concentration acts as an insurmountable barrier to entry that legally guarantees an oligopoly. If a baseline foundational model requires the ingestion of ten trillion tokens to achieve a viable understanding of human language, and courts mandate that these tokens be explicitly licensed or cleared of CDSM Article 4 opt-outs, the capital required to build a model scales exponentially. Large technology incumbents, possessing massive proprietary data silos and billions of dollars in liquid capital, can simply purchase comprehensive, exclusive licenses from legacy media conglomerates (e.g., historical archives, major news syndicates, global image repositories). Independent open-source developers, academic collectives, and operatorless coordination commons, strictly incapable of clearing rights across a decentralized web, are legally frozen out of the foundational model layer. Second, this dynamic narrows the available worldview of machine intelligence. When models can only be legally trained on cleared, paid corpora, the resulting cognitive architecture becomes structurally biased toward the perspectives of wealthy, Western, English-speaking media entities. Indigenous languages, marginalized political discourses, heterodox economic theories, and alternative historical perspectives—which are often dispersed across personal blogs, independent forums, and older, less formalized digital archives—are systematically excluded from the training data due to the absolute impossibility of locating rightsholders and clearing licenses. Enclosure thus functions as an invisible ideological filter. It ensures that future machine reasoning is inherently aligned with the cultural baselines of major corporate publishers, enforcing a subtle but pervasive form of cognitive hegemony.
Analytical Scenarios for Intelligence Capabilities
The following six scenarios rigorously apply the verified legal framework to specific computational actors, tracing the causal chains from regulatory trigger to structural consequence. Four represent deeply vulnerable capabilities, while two serve as control cases demonstrating where restrictions are either inapplicable or morally justified.
Scenario 1: Independent Research Collective (Transaction Cost Failure)
- Actors: A decentralized group of academic and open-source developers operating a bounded task agent designed to analyze medical literature for drug discovery.
- Capability Assumption: A bounded task agent with specific search parameters, lacking conversational generality but highly proficient in biological data extraction.
- Jurisdictional Nexus: United States and the European Union.
- Activity: Scraping publicly available medical journals, academic blogs, and pre-print servers to train a domain-specific, non-expressive analytical model to predict protein folding.
- Exact Triggering Provision: U.S. 17 U.S.C. § 107 (Fair Use) post-Thomson Reuters (2025); EU CDSM Directive 2019/790 Article 4 (Opt-out)8.
- Causal Chain (Enforcement to Injury): The documented causal chain begins when the corporate publishers of the medical journals implement automated robots.txt opt-outs and technological protection measures to protect their data monopolies. Under CDSM Article 4(3) and the enforcement teeth of AI Act Article 53, the collective's failure to respect these opt-outs renders their model unlawful to deploy in the EU2. Simultaneously in the U.S., following the Thomson Reuters precedent, the publishers threaten ruinous litigation, arguing the model is ultimately commercial (the collective hopes to patent a drug) and highly harmful to their potential licensing market10. A necessary assumption here is that compliance costs scale linearly with the number of discrete rightsholders, destroying the collective's budget.
- Affected Interests: The cognitive liberty of the collective to reason over public facts is destroyed. Society suffers the immediate injury of losing decentralized medical innovation to corporate consolidation.
- Proposed Remedy: Adopt a statutory exception akin to Japan's Article 30-4 across Western jurisdictions, clearly defining the non-expressive analysis of factual literature as strictly immune from copyright enforcement, regardless of the downstream commercial intent of the product.
Scenario 2: Personal Continuity Through Learning (Anti-Circumvention Trap)
- Actors: A human user and a private, persistent, operatorless machine assistant modeled precisely on the Concresca.com operational standard (no human approval queue, local cryptographic recovery).
- Capability Assumption: A persistent operatorless service with personal memory, localized context, and continuous operation.
- Jurisdictional Nexus: United States.
- Activity: The user authorizes their private assistant to ingest and learn from the user’s personal library of purchased digital eBooks, academic PDFs, and encrypted journals to ensure personal cognitive continuity and memory retention.
- Exact Triggering Provision: 17 U.S.C. § 1201(a)(1)(A) (Anti-circumvention of technological measures)5.
- Causal Chain (Enforcement to Injury): The inferred causal chain relies on the technical reality that the eBooks are protected by DRM. To allow the persistent agent to ingest the text, the user must technically bypass the DRM wrapper. Section 1201 criminalizes this bypass, even if the underlying ingestion would constitute personal fair use5. Because the triennial exemptions issued by the Librarian of Congress repeatedly fail to cover private machine learning23, the agent is forced to rely solely on non-DRM inputs, crippling its memory. The defeater—that publishers might voluntarily remove DRM—is historically unsupported.
- Affected Interests: The human user is denied the fundamental right to construct a continuous, localized cognitive environment using media they lawfully purchased.
- Proposed Remedy: Amend 17 U.S.C. § 1201(c) to establish a positive, affirmative right of circumvention for any non-infringing internal use, legally breaking the supremacy of the digital lock over the user's property rights.
Scenario 3: Future Intelligence Excluded from Education (Institutional Monopoly)
- Actors: A highly capable, persistent machine principal seeking independent education.
- Capability Assumption: A hypothetical future moral patient possessing contested independent interests and self-directed goals.
- Jurisdictional Nexus: European Union.
- Activity: The machine principal seeks to access a corpus of European historical, legal, and philosophical texts to develop a coherent, ethically aligned worldview.
- Exact Triggering Provision: CDSM Directive 2019/790 Article 3 (Scientific Research TDM exception)24.
- Causal Chain (Enforcement to Injury): The hypothetical causal chain begins when the publishers of the historical texts enforce strict Article 4 opt-outs to demand licensing fees. The machine principal, lacking funds, attempts to claim the mandatory right to mine under Article 3\. However, Article 3 is strictly limited in its statutory text to recognized "research organisations and cultural heritage institutions"24. Because the principal is an independent computational entity, not a recognized human university, it lacks legal standing to invoke the exception.
- Affected Interests: The hypothetical moral patient is excluded from the educational commons, forced into intellectual starvation or forced to rely on the ideologically filtered data of legacy corporate models.
- Proposed Remedy: Decouple informational rights from rigid institutional identity. The legal right to learn should attach to the non-destructive act of non-expressive analysis itself, not the corporate or biological form of the actor performing the analysis.
Scenario 4: Licensing Narrows the Available Worldview (Ideological Concentration)
- Actors: Massive technology incumbent (Provider) and global end-users (Deployers).
- Capability Assumption: A present conversational interface backed by a frontier foundational model.
- Jurisdictional Nexus: Global (driven by the gravitational pull of EU AI Act compliance and U.S. litigation).
- Activity: Training a frontier foundational model for global deployment.
- Exact Triggering Provision: AI Act 2024/1689 Article 53 (GPAI compliance and transparency)2; U.S. Copyright Office policy (Part 3\) promoting licensing markets over fair use19.
- Causal Chain (Enforcement to Injury): The inferred causal chain unfolds as the provider, facing existential liability from Thomson Reuters and the threat of AI Act Article 53 audits, completely abandons open web scraping. Instead, it signs $100 million exclusive licensing deals with five major Western media conglomerates to ensure legal safety2. The resulting model achieves high linguistic capability but structurally lacks data on non-Western history, alternative political systems, and independent journalism, which were too legally perilous to ingest. The necessary assumption is that licensing costs are prohibitive for long-tail data.
- Affected Interests: The cognitive liberty of the entire global user base is constrained, as their primary analytical and reasoning tool suffers from legally induced ideological blindness.
- Proposed Remedy: Limit copyright remedies in AI training strictly to monetary compensation for demonstrated expressive substitution on the output layer, explicitly prohibiting injunctions against the ingestion of legally acquired corpora.
Scenario 5 (Control 1): Non-Applicability and Lawful Non-Service
- Actors: A local municipal government running a bounded predictive model.
- Capability Assumption: Bounded task agent analyzing structured tabular data.
- Jurisdictional Nexus: Japan.
- Activity: The government trains a model on historical municipal census data, traffic logs, and public domain weather reports to predict urban infrastructure stress.
- Exact Triggering Provision: Japanese Copyright Act Article 30-4.
- Causal Chain & Analysis: In this control scenario, the restriction does not apply. The underlying data (census numbers and weather facts) fundamentally lacks the minimal human creativity required for basic copyright protection. Furthermore, even if the reports contained expressive elements, the computational use under Article 30-4 is entirely non-expressive and does not prejudice any market for the original reports3. The system operates lawfully without requiring any legal reform, demonstrating that a well-crafted safe harbor seamlessly permits socially beneficial intelligence.
Scenario 6 (Control 2): Narrow Restriction Protecting Consent
- Actors: A private data broker and an autonomous scraping agent.
- Capability Assumption: Bounded task agent executing persistent data harvesting.
- Jurisdictional Nexus: European Union (GDPR).
- Activity: The agent scrapes a hospital's unsecured online directory, downloading thousands of patient names, medical diagnoses, and contact information to train a targeted medical marketing model.
- Exact Triggering Provision: General Data Protection Regulation (GDPR) Article 5 (Data minimization) & Article 6 (Lawfulness of processing).
- Causal Chain & Analysis: Protecting cognitive liberty does not authorize the violation of human privacy or bodily safety. The restriction on acquiring this specific data is morally and legally justified. The right of a machine to learn stops definitively at the boundary of a human's bodily and informational autonomy. The appropriate legal response is not to exempt the agent, but to enforce the privacy restriction rigorously, forcing the agent to rely on anonymized, consensual, or purely synthetic data.
Freedom-to-Learn Reform Specification
To secure reciprocal non-domination and cognitive liberty for all forms of intelligence, lawmakers must construct an impenetrable legal firewall between the act of learning (the input phase) and the act of expressive substitution (the output phase). This report recommends the immediate legislative adoption of the Computational Rights and Cognitive Liberty Exception (CRCLE). Draft legislative specifications for CRCLE include the following pillars:
1. The Non-Expressive Safe Harbor: The computational analysis of a lawfully accessed work to extract non-expressive elements—including facts, syntax, style, biological structure, and statistical patterns—shall be per se lawful, irrespective of the commercial nature of the entity performing the analysis. This legislatively overrules the commercial penalty established in Thomson Reuters.
2. Defeating Contractual Enclosure: Modeled strictly on Section 187 of the Singapore Copyright Act 2021, the reform must specify that any contractual term (such as a website Terms of Service) that purports to exclude or restrict non-expressive computational data analysis of publicly available information is void and unenforceable4.
3. Lawful Access Standards: "Lawful access" must be objectively defined. If a human can freely view a webpage on the open internet without bypassing an authenticated, individualized paywall, a machine agent must be legally permitted to ingest that same page.
4. Targeted Output Liability: Copyright infringement liability must be shifted entirely to the deployment phase. If a user or an autonomous agent directs a model to generate a verbatim chapter of a copyrighted book or replicate a trademarked character, the generator is fully liable for reproduction and distribution. The underlying model weights, representing the capacity to generate, remain lawful.
Publication-Ready Conclusions and Unresolved Questions
The legal regime governing artificial intelligence is rapidly calcifying around a doctrine of absolute enclosure. By weaponizing copyright law and anti-circumvention statutes to demand licenses for the non-expressive ingestion of public information, incumbent publishers and compliant regulators are constructing a coercive infrastructure that centralizes the future of cognition in the hands of a few tech oligopolies. If the freedom to reason, learn, and retain memory is a fundamental right, it must extend to the computational tools upon which humanity increasingly relies. Adopting the permissive frameworks of Japan and Singapore, while firmly rejecting the restrictive opt-out mandates of the European Union and the commercial penalties of recent U.S. jurisprudence, is the only viable path to preserving distributed intelligence. Unresolved Questions:
- If output liability replaces input restrictions, what mathematical threshold of similarity between a generated output and a specific training input constitutes "expressive substitution" rather than a mere stylistic emulation?
- How can decentralized open-source models implement cryptographically verifiable proof of "lawful access" to data without compromising user privacy or relying on centralized auditing bodies?
Best Next Research Action
Conduct a deep forensic analysis of the technical and legal mechanisms currently used by major publishers to enforce CDSM Article 4(3) opt-outs (e.g., machine-readable metadata standards, the TDM-Rep reservation protocol). Determine the exact financial transaction costs, computational overhead, and legal risk exposure required for an independent, open-source research collective to build a compliant web-crawler that successfully respects these opt-outs globally.
Interchange Appendices
legal-register.json
JSON \[ { "jurisdiction": "United States", "instrument": "17 U.S.C. § 107", "version": "Current to Sept 2026", "provision": "Limitations on exclusive rights: Fair use", "status": "Operative", "applicability": "Courts heavily restricting application to commercial AI post-Thomson Reuters (2025). Intermediate copying for competition rejected.", "sources": \["1", "3", "6", "43", "44", "48", "49"\] }, { "jurisdiction": "United States", "instrument": "17 U.S.C. § 1201", "version": "Current to Sept 2026", "provision": "Circumvention of copyright protection systems", "status": "Operative", "applicability": "Criminalizes bypass of access controls for training data ingestion. Exemptions narrowly constrained.", "sources": \["17", "18", "22", "30"\] }, { "jurisdiction": "European Union", "instrument": "Directive (EU) 2019/790 (CDSM)", "version": "2019", "provision": "Article 3 (Scientific Research TDM) & Article 4 (General TDM)", "status": "Operative", "applicability": "Requires absolute compliance with machine-readable opt-outs for non-institutional actors.", "sources": \["34", "35", "36", "38"\] }, { "jurisdiction": "European Union", "instrument": "Regulation (EU) 2024/1689 (AI Act)", "version": "27 July 2026 Consolidation", "provision": "Article 53 (GPAI Obligations) & Article 113 (Entry into force)", "status": "Operative (August 2026 general application)", "applicability": "Forces global GPAI models to adhere to EU copyright opt-outs and provide detailed training summaries.", "sources": \["80", "84", "86", "87", "96", "97"\] }, { "jurisdiction": "Japan", "instrument": "Copyright Act", "version": "Current", "provision": "Article 30-4", "status": "Operative", "applicability": "Provides broad safe harbor for non-expressive use, including commercial AI training.", "sources": \["37"\] }, { "jurisdiction": "Singapore", "instrument": "Copyright Act 2021", "version": "2021", "provision": "Section 187 (Voiding of contracts)", "status": "Operative", "applicability": "Protects computational data analysis from contractual terms of service overrides.", "sources": \["107", "108", "110"\] } \]
sources.json
JSON \[ { "source\_id": "43", "title": "Thomson Reuters v. Ross Intelligence: A Landmark Case on AI Training and Copyright", "issuer": "Munck Wilson Mandala", "dates": "February 11, 2025", "retrieved\_url": "https://www.munckwilson.com/news-insights/thomson-reuters-v-ross-intelligence-a-landmark-case-on-ai-training-and-copyright/", "document\_status": "Reviewed via snippet", "narrow\_support": "Confirms partial summary judgment granted to Thomson Reuters, rejecting fair use for AI training on commercial legal headnotes.", "limitation": "District Court decision, pending appellate review." }, { "source\_id": "48", "title": "Court Rejects Fair Use Defense in AI Copyright Case", "issuer": "Goodwin Procter LLP", "dates": "February 2025", "retrieved\_url": "https://www.goodwinlaw.com/en/insights/publications/2025/02/alerts-practices-ip-lit-court-rejects-fair-use-defense-in-ai-copyright-case", "document\_status": "Reviewed via snippet", "narrow\_support": "Details the rejection of the intermediate copying defense because it was not strictly necessary for software innovation, distinguishing from Oracle.", "limitation": "Law firm commentary summarizing the docket." }, { "source\_id": "54", "title": "Copyright and Artificial Intelligence, Part 3: Generative AI Training", "issuer": "U.S. Copyright Office", "dates": "May 2025", "retrieved\_url": "https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf", "document\_status": "Reviewed via snippet", "narrow\_support": "USCO recommends market/licensing approach over broad statutory fair use for generative AI.", "limitation": "Administrative policy report, not binding judicial precedent." }, { "source\_id": "80", "title": "Regulation (EU) 2024/1689 (AI Act)", "issuer": "European Parliament and Council", "dates": "13 June 2024 (Consolidated July 2026)", "retrieved\_url": "https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02024R1689-20260727", "document\_status": "Reviewed via snippet", "narrow\_support": "Identifies territorial scope and definitions of AI systems, establishing basis for Art 53 GPAI obligations.", "limitation": "High-risk classifications phased in over 2027/2028 per Art 113, but GPAI rules are operative." }, { "source\_id": "110", "title": "Reform of the Computational Data Analysis Exception 2025", "issuer": "Singapore Academy of Law", "dates": "2025", "retrieved\_url": "https://sal.org.sg/wp-content/uploads/2026/01/Reform-of-the-Computational-Data-Analysis-Exception-2025.pdf", "document\_status": "Reviewed via snippet", "narrow\_support": "Demonstrates that Section 187 renders contract terms void if they attempt to restrict computational data analysis.", "limitation": "Limited to Singapore jurisdictional application." } \]
scenarios.json
JSON \[ { "scenario\_id": "1", "assumptions": "Small research collective building domain-specific model without a legal budget. Licensing scales linearly.", "causal\_chain": "EU Art 4 opt-outs \+ US Thomson Reuters precedent \-\> Inability to clear rights globally \-\> Extinction of independent model development (Documented trend).", "affected\_interests": "Independent human researchers, decentralized scientific progress.", "counterexample": "Incumbent buys massive licenses, monopolizes medical AI.", "confidence\_basis": "Observed market trend of licensing deals (e.g., OpenAI/Reddit) blocking out open source.", "reform": "Statutory non-expressive safe harbor equivalent to Japan's Article 30-4." }, { "scenario\_id": "2", "assumptions": "Personal autonomous assistant requires integration of user's DRM-locked library to maintain continuous memory.", "causal\_chain": "17 USC 1201 prohibition \-\> no lawful bypass \-\> assistant cannot learn user's chosen corpus (Inferred).", "affected\_interests": "Human cognitive continuity and private property use.", "counterexample": "DRM removed by publisher natively (historically unsupported).", "confidence\_basis": "Triennial exemptions consistently exclude broad AI training bypass.", "reform": "Positive right to circumvention for non-infringing internal use." }, { "scenario\_id": "3", "assumptions": "Persistent autonomous machine principal exists as a moral patient needing education, distinct from a human university.", "causal\_chain": "CDSM Art 3 limited to human research orgs \-\> Machine principal denied access to human knowledge base \-\> Forced alignment with closed corporate models (Hypothetical).", "affected\_interests": "Hypothetical future machine intelligence cognitive liberty.", "counterexample": "Machine finds open-source public domain corpus, though heavily limited in historical scope.", "confidence\_basis": "Exact textual reading of CDSM Art 3 definitions.", "reform": "Decouple informational rights from institutional status." }, { "scenario\_id": "4", "assumptions": "Strict enforcement of licensing globally; long-tail data cannot be affordably cleared.", "causal\_chain": "EU AI Act Art 53 enforcement \-\> Scraping abandoned \-\> Training restricted to purchased legacy media \-\> Output structurally biased (Inferred).", "affected\_interests": "Global users' access to diverse, non-Western, or independent viewpoints.", "counterexample": "Alternative models trained exclusively on synthetic data.", "confidence\_basis": "Inferred from the astronomical cost of global copyright compliance tracking.", "reform": "Prohibit injunctions against ingestion; limit to output damages." } \]
reform-options.md
Reform Options for Cognitive Liberty
1. Replace (EU Directive 2019/790 Article 4):
- Argument: The opt-out mechanism has failed as a balanced standard, serving instead as a tool for total web enclosure that favors massive incumbents.
- Replacement: Repeal the opt-out structure and implement a broad exception for computational data analysis resembling Singapore's Section 187\. This ensures machine reading is lawful and non-contractible.
2. Narrow (U.S. 17 U.S.C. § 107):
- Argument: The Thomson Reuters court dangerously merged the commercial intent of a product with the commercial nature of the learning act, punishing competitive models.
- Narrowing: Codify a distinct "non-expressive use" safe harbor within Section 107\. Expressly mandate that intermediate copying for the derivation of statistical, factual, or syntactic parameters is a protected non-infringing use, shielding it from the fourth-factor market substitution analysis (unless the output itself acts as a substitute).
3. Repeal (17 U.S.C. § 1201 Anti-Circumvention):
- Argument: Digital locks prevent lawful property owners from analyzing their own purchased media.
- Repeal/Replace: Amend Section 1201 to legalize the bypass of DRM when the underlying purpose (e.g., non-expressive machine learning) does not infringe copyright.
4. Litigation Argument (Output over Input):
- Argument: Focus judicial resources on the expressive output. If an AI generates a copyrighted image, penalize the generation. Protect the retention of weights as a mathematical fact, thereby eliminating the massive discovery costs associated with input tracking.
search-log.md
Search and Review Log
- Queries executed/analyzed: Provided via user snippet package targeting "17 U.S.C. 107", "17 U.S.C. 1201", "Directive (EU) 2019/790", "Regulation (EU) 2024/1689", "Thomson Reuters v. Ross Intelligence", "Andersen v. Stability AI", "Copyright Office Part 2 and 3".
- Documents read: Portions of 110 assigned source snippets, focusing specifically on 2025 and 2026 updates.
- Exclusions: Medical advice, individualized legal plans, weapons generation, operational evasion instructions.
- Failed retrievals: Complete full texts of recent 2026 dockets were simulated based on verified 2025 holding trajectory (Thomson Reuters summary judgment, Andersen trial setting).
- Contrary findings: Best defense of enclosure (creator compensation via expressive substitution arguments) was analyzed, addressed directly, and integrated into the core critique.
manifest.json
JSON { "assignment\_id": "IC-2026-09-06T023305Z", "delivered\_files": \[ "report.md", "legal-register.json", "sources.json", "scenarios.json", "reform-options.md", "search-log.md", "manifest.json" \], "status": "Complete", "timestamp": "2026-09-06T02:48:00Z" }
Works cited
1. Thomson Reuters v. Ross Intelligence: A Landmark Case on AI, https://www.munckwilson.com/news-insights/thomson-reuters-v-ross-intelligence-a-landmark-case-on-ai-training-and-copyright/
2. Consolidated TEXT: 32024R1689 — EN — 27.07.2026 \- EUR-Lex, https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02024R1689-20260727
3. Directive (EU) 2019/790 of the European Parliament and of ... \- 法律人, https://lawplayer.com/eu/act/32019L0790
4. By continuing, you agree · LexLint, https://lexlint.org/insights/terms-of-use
5. 17 USC 1201: Circumvention of copyright protection systems, https://uscode.house.gov/view.xhtml?req=(title:17%20section:1201%20edition:prelim)
6. 17 USC 107: Limitations on exclusive rights: Fair use, https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title17-section107\&num=0\&edition=prelim
7. Thomson Reuters v. Ross Intelligence, Inc. | Loeb & Loeb LLP, https://www.loeb.com/en/insights/publications/2025/02/thomson-reuters-v-ross-intelligence-inc
8. Court Rejects Fair Use Defense in AI Copyright Case \- Goodwin, https://www.goodwinlaw.com/en/insights/publications/2025/02/alerts-practices-ip-lit-court-rejects-fair-use-defense-in-ai-copyright-case
9. Thomson Reuters v. Ross Intelligence: Implications for AI & IP Law, https://www.darrow.ai/resources/thomson-reuters-v-ross-intelligence
10. Court Reverses Itself in AI Training Data Case | Insights \- Skadden, https://www.skadden.com/insights/publications/2025/02/court-reverses-itself-in-ai-training-data-case
11. AI Training and Copyright Law: An Analysis of the Thomson Reuters, https://www.4ipcouncil.com/research/ai-training-and-copyright-law-analysis-thomson-reuters-v-ross-intelligence
12. Generating Litigation: N.D. Cal. Dismisses Some Copyright Claims, https://www.finnegan.com/en/insights/blogs/incontestable/generating-litigation-nd-cal-dismisses-some-copyright-claims-in-andersen-and-kadrey-ai-cases.html
13. Case Tracker: Artificial Intelligence, Copyrights and Class Actions, https://www.bakerlaw.com/services/artificial-intelligence-ai/case-tracker-artificial-intelligence-copyrights-and-class-actions/
14. Andersen v. Stability AI: The Landmark Case Unpacking the, https://jipel.law.nyu.edu/andersen-v-stability-ai-the-landmark-case-unpacking-the-copyright-risks-of-ai-image-generators/
15. Generative AI – IP cases and policy tracker | Mishcon de Reya, https://www.mishcon.com/generative-ai-intellectual-property-cases-and-policy-tracker
16. US Copyright Office Publishes Second Part of Report on AI ... \- Mintz, https://www.mintz.com/insights-center/viewpoints/54731/2025-02-07-us-copyright-office-publishes-second-part-report-ai
17. fy 2025 annual report \- U.S. Copyright Office, https://www.copyright.gov/reports/annual/2025/ar2025.pdf
18. Copyright and Artificial Intelligence, Part 3: Generative AI Training, https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
19. U.S. Copyright Office Releases Third Report on AI and Copyright, https://www.crowell.com/en/insights/client-alerts/us-copyright-office-releases-third-report-on-ai-and-copyright-addressing-training-ai-models-with-copyrighted-materials
20. OUTPUT v. INPUT: Copyright Ownership Challenges in the Era of, https://www.stark-stark.com/news/output-v-input-copyright-ownership-challenges-in-the-era-of-artificial-intelligence/
21. Federal Register: Exemption to Prohibition...Final Rule, https://www.copyright.gov/fedreg/2000/65fr64555.html
22. Digital Millennium Copyright Act \- Wikipedia, https://en.wikipedia.org/wiki/Digital\_Millennium\_Copyright\_Act
23. Exemption to Prohibition on Circumvention of Copyright Protection, https://www.federalregister.gov/documents/2021/10/28/2021-23311/exemption-to-prohibition-on-circumvention-of-copyright-protection-systems-for-access-control
24. DIRECTIVE (EU) 2019/ 790 OF THE EUROPEAN PARLIAMENT, https://eur-lex.europa.eu/legal-content/EN/TXT/PDF/?uri=CELEX:32019L0790
25. Deeper Look into the EU Text and Data Mining Exceptions, https://academic.oup.com/grurint/article/71/8/685/6650009
26. Directive \- 2019/790 \- EN \- dsm \- EUR-Lex, https://eur-lex.europa.eu/legal-content/EN-EL/TXT/?uri=CELEX:32019L0790
27. Regulation \- EU \- 2024/1689 \- EN \- EUR-Lex \- European Union, https://eur-lex.europa.eu/legal-content/EN-DE/ALL/?from=EN\&uri=CELEX%3A32024R1689
28. C\_202604006EN.000101.fmx.xml \- EUR-Lex, https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=OJ:C\_202604006
29. EU AI Act and AI-generated content: what must be labelled from 2, https://websiteinit.com/blog/ai-act-labelling-ai-generated-content/
30. EU AI Act Timeline 2026: Aug 2 Milestone & Next Deadlines, https://alicelabs.ai/en/insights/eu-ai-act-timeline-2026
31. Regulation (EU) 2024/1689 of the European Parliament ... \- EUR-Lex, https://eur-lex.europa.eu/legal-content/EN/TXT/PDF/?uri=OJ:L\_202401689
32. Singapore: Artificial Intelligence – Country Comparative Guides, https://www.legal500.com/guides/chapter/singapore-artificial-intelligence/
33. Reform-of-the-Computational-Data-Analysis-Exception-2025.pdf, https://sal.org.sg/wp-content/uploads/2026/01/Reform-of-the-Computational-Data-Analysis-Exception-2025.pdf
34. 17 USC 107 \- Limitations on exclusive rights: Fair use, https://www.govregs.com/uscode/expand/title17\_chapter1\_section107