Civic / Privacy / Digital Rights
The Safety Poison Pill Thesis: A Full Adversarial Investigation
Report summary
The “safety poison pill” thesis survives adversarial scrutiny, but only in a narrower and more defensible form than its strongest rhetoric suggests.
Key topics
- Civic / Privacy / Digital Rights
- Civic
- Privacy
- Digital Rights
- AI
- Agentic Web
- Runtime
- Cognitive Liberty
- Physics
Research provenance
For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.
Source availability: 121 citation markers in the source export have no recoverable source links. Those markers are omitted from this reader; any supplied bibliography and ordinary links remain. Check the original sources before relying on the cited claims.
This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.
Full report
On this page
Executive judgment and analytical frame
The “safety poison pill” thesis survives adversarial scrutiny, but only in a narrower and more defensible form than its strongest rhetoric suggests.
The historical record supports a serious warning: after salient threats, institutions can obtain broad monitoring or control powers whose scope, secrecy, error costs, and persistence are easier to establish than their concrete effectiveness. U.S. experience after September 11 is especially instructive. The Justice Department’s own inspector general documented serious misuse and compliance failures involving national security letters; the Privacy and Civil Liberties Oversight Board concluded that the Section 215 bulk telephone-record program presented serious privacy and constitutional concerns while providing only limited value; and later legislation terminated that bulk program.
There is also direct empirical support for the mechanism at the heart of the thesis: people change what they read when they believe reading is observed. Jon Penney’s post-Snowden research found a marked decline in views of privacy-sensitive Wikipedia pages after the surveillance disclosures; contemporaneous reporting described particularly large declines for terrorism-related articles. The evidence is observational rather than a randomized experiment, so it does not prove that every form of monitoring chills inquiry, but it demonstrates that surveillance can alter ordinary information-seeking even where the underlying reading is lawful.
The thesis becomes exaggerated, however, when it treats all safety controls as equivalent to surveillance, all preventive action as illegitimate until harm has already occurred, or all risk terminology as inherently manipulable. Known-hash matching against previously identified child sexual abuse material, narrowly targeted fraud controls, warnings before malicious websites, privacy-preserving age assertions, court-supervised investigation of true threats, and request-level restrictions on genuinely dangerous AI assistance can be substantially different from building permanent dossiers on what people search, read, ask, or privately investigate. Microsoft’s PhotoDNA, for example, compares image hashes against known illegal imagery rather than attempting to infer a user’s ideology or intentions; NCMEC received 21.3 million CyberTipline reports in 2025, showing that the underlying exploitation problem is not hypothetical.
The AI portion of the thesis likewise needs updating in light of evidence available by September 2026. Frontier-AI cyber risk is no longer merely a thought experiment about distant superintelligence. The UK AI Security Institute reports that models progressed from weak performance on apprentice-level cyber tasks in late 2023 to completing some expert-level tasks by 2025, and its 2026 work describes rapidly increasing autonomous cyber task horizons. At the same time, model-provider evaluations suggest important asymmetries: some advanced models remain substantially useful for defensive vulnerability discovery, and providers are explicitly trying to preserve legitimate cyber and biological work while restricting offensive misuse. Those findings make a blanket “safety is censorship” position untenable.
The resulting verdict is:
| Proposition | Adversarial judgment |
|---|---|
| Exceptional harms can be used to normalize broad surveillance. | Strongly supported. Post-9/11 intelligence history and earlier emergency episodes provide concrete examples. |
| Surveillance of information-seeking can chill lawful inquiry. | Supported, with causal caveats. Post-Snowden behavior provides unusually direct evidence. |
| “Temporary” safety powers inevitably become permanent. | Rejected as stated. Persistence is a real risk, but Section 215 bulk collection ended, EO 9066 was terminated, and sunset/reform mechanisms sometimes work. |
| Broad undefined risk labels create opportunities for mission creep. | Strongly supported as an institutional-risk proposition, especially when definitions, evidence thresholds, review, and appeals are absent. The relevant lesson follows from documented overbreadth and compliance problems rather than requiring proof of bad faith. |
| All age verification necessarily creates identity surveillance. | Rejected. Some systems are highly intrusive, but tokenized or third-party age assertions and legal prohibitions on compulsory government ID show that age and identity can be separated. |
| Platform moderation cannot demonstrably reduce harmful behavior. | Rejected. A large Reddit study found that users remaining after two hate-community bans reduced hate-speech usage by at least 80%, although other moderation changes have produced backlash or displacement effects. |
| AI safety today concerns only speculative future catastrophe. | Rejected. Cyber capabilities and some biological-assistance capabilities are measurable now, although the probability and scale of true catastrophic outcomes remain deeply uncertain. |
| Safety controls should normally target conduct or narrowly defined capability rather than curiosity, identity, or generalized thought. | Survives strongly. This is the most defensible core of the thesis, and it is consistent with constitutional doctrine protecting receipt of information and associational privacy while leaving room to regulate true threats, incitement, child sexual abuse material, and coordinated terrorist material support. |
“Cognitive liberty” should therefore be understood here primarily as a normative principle, not as the name of a single settled U.S. constitutional right. American doctrine nevertheless protects important components of it. The Supreme Court has recognized a First Amendment interest in receiving information, protection against compelled disclosures that deter association, and constitutional protection for mere private possession of even obscene material in the home; conversely, established doctrine permits regulation of categories such as true threats, incitement meeting the Brandenburg standard, child pornography, and some coordinated material support to designated terrorist organizations.
That distinction is foundational. A right to think, read, investigate, and inquire privately does not imply a right to possess contraband, threaten people, exploit children, commandeer computers, obtain every conceivable dangerous capability from every private service, or demand that institutions ignore imminent danger. The issue is whether prevention can be accomplished without making the population’s ordinary intellectual life itself the object of persistent observation and classification.
Historical record: emergencies, censorship, and the surveillance state
The strongest historical argument for the poison-pill thesis is not that emergency measures are always unnecessary. It is that institutions operating under extreme uncertainty and political fear repeatedly accept false-positive liberty costs that would have been unacceptable under ordinary conditions.
The World War I censorship environment is an early American example. The Library of Congress describes the Espionage Act of 1917 and Sedition Act of 1918 as mechanisms used to enforce loyalty and silence dissent, with postal power employed against disfavored publications. This is a particularly important precedent because the operative category was not simply violent conduct: ideas, opposition, and perceived disloyalty became security problems.
Executive Order 9066 supplies an even more severe case. Issued on February 19, 1942, it authorized exclusion from military areas and led to the incarceration of Japanese Americans. The order itself did not name an ethnic group, illustrating one recurrent danger in emergency governance: formally broad security language can become highly discriminatory in implementation. President Gerald Ford formally confirmed the order’s termination in 1976, demonstrating both the enormous cost of emergency overreach and the fact that emergency regimes are not literally irreversible.
These examples do not prove that a modern AI-safety rule will become internment or wartime censorship. That would be an inflammatory analogy rather than analysis. They establish a more modest proposition: “national security” or “safety” is not self-validating evidence of proportionality. Extraordinary consequences require independent evidence about whom a measure reaches, what behavior it prevents, how errors are corrected, and when the authority terminates.
The Patriot Act lesson
The post-September 11 environment is the thesis's most relevant case because the threat was unquestionably real while the policy response contained both valuable tools and demonstrable overreach. The Justice Department’s own archived defense of the USA PATRIOT Act argued that terrorism threatened both security and freedom and that the Act achieved a reasonable balance while giving investigators essential tools. That is not a straw man: a government facing dispersed terrorist networks had legitimate reasons to improve information sharing, investigate financing and communications, and move faster than traditional law-enforcement structures sometimes allowed.
But the most important adversarial evidence comes from the government’s own oversight apparatus. The DOJ inspector general found serious misuse of national security letters after the Patriot Act broadened relevant authorities. Follow-up work found that many problems arose from mistakes, confusion, inadequate training, inadequate guidance, weak controls, and insufficient oversight rather than a centrally coordinated scheme to violate rights. That point actually strengthens the institutional version of the poison-pill thesis: one need not posit malicious officials for expansive surveillance authorities to exceed their intended bounds.
Section 215 bulk telephone metadata became the clearest test of necessity and effectiveness. After extensive review, the Privacy and Civil Liberties Oversight Board concluded that the bulk program lacked a viable legal foundation under Section 215, implicated First- and Fourth-Amendment concerns, posed serious privacy risks, and had shown only limited value. The Board’s proceedings also reflected disagreement: defenders maintained that the government had acted under judicially approved interpretations and that national-security collection could be valuable. The subsequent USA FREEDOM Act ended the former bulk collection architecture and replaced it with a more targeted records mechanism.
That sequence is almost a laboratory demonstration of the doctrine this report ultimately recommends:
threat → broad authority → secret or technically opaque operation → later independent review → efficacy challenge → narrowing → targeted replacement.
The failure was therefore not “trying to prevent terrorism.” The failure was allowing a preventive theory to carry too much weight before proving that bulkness itself was necessary. The later move from bulk acquisition toward targeted records undermines the claim that security necessarily required collection at the original scale.
At the same time, history disproves one absolutist version of the poison-pill narrative. Section 215 bulk collection did not become permanent. Institutional safeguards eventually mattered. Oversight generated evidence, Congress changed the law, and collection architecture changed. A serious doctrine should therefore demand sunsets and review rather than assume that review is futile.
The intelligence problem after Section 215
Section 702 illustrates why “surveillance” also cannot be treated as a single undifferentiated practice. Section 702, enacted separately from the Patriot Act, permits acquisition targeting non-U.S. persons reasonably believed to be abroad, but communications involving Americans may enter the resulting databases and later be queried. Even Democratic Senator Dick Durbin, a prominent advocate of warrant reforms, has repeatedly characterized Section 702 as a valuable foreign-intelligence tool while objecting to warrantless searches involving Americans.
The civil-liberties objection has substantial factual grounding. Brennan Center analyses drawing on FISA Court materials describe major historical FBI compliance problems involving U.S.-person queries, including hundreds of thousands of improper or noncompliant searches disclosed in earlier oversight. Reforms enacted in 2024 reportedly reduced improper searches, but critics continue to argue that a warrant should ordinarily be required to query Americans’ communications.
The program's 2026 status also demonstrates how “sunset” can be more complicated than a single expiration date. Congress did not renew Section 702's statutory authorization before its June 2026 lapse, but annual FISA Court certifications approved in March 2026 allow currently authorized acquisition to continue into roughly March 2027. Thus expiration matters, but previously approved surveillance does not instantly disappear when the statute sunsets.
This history produces a crucial distinction for the thesis. Targeted foreign-intelligence collection with demonstrated intelligence value is a harder case against surveillance than indiscriminate domestic bulk collection of everyone’s communications metadata. A cognitive-liberty doctrine that simply declares both illegitimate would fail the adversarial test. The appropriate demand is instead minimization, judicial authorization where U.S.-person access is concerned, bounded purposes, auditable queries, retention limits, and prohibitions on repurposing foreign-intelligence databases into ordinary domestic investigative shortcuts.
Surveillance changes behavior before anyone is punished
The deepest relevance to cognitive liberty lies upstream from arrest or censorship. Information collection can matter because people who anticipate observation may avoid controversial but lawful inquiry.
Penney’s Wikipedia study examined traffic following the Snowden disclosures and found a substantial reduction in views of articles concerning privacy-sensitive topics. Because the study is observational, alternative explanations cannot be eliminated as cleanly as in a randomized experiment, and one should not extrapolate its magnitude mechanically to every surveillance technology. But it is unusually probative because it measured reading, not merely stated attitudes about privacy.
That effect has constitutional resonance. Congress’s Constitution Annotated summarizes Supreme Court doctrine recognizing that compelled disclosure of association can expose individuals to threats and harassment and deter association, and it recounts Lamont v. Postmaster General, where a law burdening receipt of certain political publications violated the First Amendment because it inhibited people's ability to receive information they wished to receive.
The mechanism should therefore be treated as real even without assuming universal self-censorship:
monitoring → perceived observability → anticipated consequences or stigma → avoidance of sensitive inquiry.
A society can formally leave a book, query, political position, medical topic, religious view, or controversial scientific question legal while still degrading cognitive liberty if accessing it predictably contributes to a durable risk profile. The evidence does not establish that every logged query causes this effect. It establishes enough to make generalized intellectual profiling a harm that regulators should have to justify rather than an invisible administrative default.
The contemporary digital safety stack
Modern “safety” differs from classical censorship because control can occur at many layers simultaneously: search ranking, recommendations, account eligibility, age assurance, identity verification, content removal, transaction friction, behavioral scoring, AI refusals, and data retention. The right question is therefore not “moderation or no moderation?” It is what layer is being controlled, how much information about the person must be collected, and what evidence connects that control to the harm?
Platform moderation: neither harmless nor inherently futile
Evidence does not support the claim that content moderation is merely cosmetic. A large empirical study of Reddit's bans of r/fatpeoplehate and r/CoonTown examined more than 100 million posts and comments; among users who remained on Reddit, hate-speech usage fell by at least 80%, and researchers did not find significant increases in hate speech in communities receiving displaced users. That is meaningful evidence that precisely targeted community enforcement can change behavior rather than simply move it elsewhere.
But this result does not license generalized automated suppression. Other empirical work has found settings where moderation changes produce backlash rather than improvement, including research on BitChute. Moderation effects are therefore intervention-specific, platform-specific, and population-specific. “Moderation works” and “moderation never works” are both too crude.
Procedural architecture matters as much as classification accuracy. Meta's published systems, for example, provide appeal paths for many enforcement decisions and restoration where content is found to have been removed incorrectly, while its Oversight Board has continued pressing for better disclosure about appeals and reversals. Provider reporting should not be treated as independent proof of effectiveness, but these mechanisms demonstrate that content safety can be designed around correction rather than treating classifiers as final authorities.
The poison-pill danger begins when a system moves from removing specified prohibited conduct toward profiling people according to inferred beliefs or future dangerousness, especially when that profile governs unrelated access. Removing a credible threat is not the same intervention as assigning a lifelong “extremism score” because someone repeatedly reads extremist propaganda for research. The former can be evaluated against the threatening communication; the latter requires inference about the reader.
Online-safety legislation: law can narrow or amplify the risk
The UK Online Safety Act illustrates both possibilities. Its framework imposes duties concerning illegal content, child safety, risk assessment, and platform processes rather than simply giving a regulator an unlimited mandate to remove anything “harmful.” Ofcom’s illegal-content duties entered into force in March 2025, and child-safety duties followed with detailed implementation requirements. UK government guidance has explicitly stated that the Act does not prohibit legal adult content merely because it may be unsuitable for children.
Yet the same architecture can expand over time. By December 2025, encouraging or assisting serious self-harm had been made a priority offense under the Act, requiring Ofcom to update its risk and enforcement framework in 2026. That particular category is more specific than a generic prohibition on “unsafe content,” because it is connected to an offense defined by conduct; nevertheless, expansion of enumerated safety categories should remain a trigger for renewed proportionality review rather than becoming self-justifying merely because the legislation already exists.
The EU Digital Services Act provides another useful design model. It requires very large platforms and search engines to assess systemic risks and adopt reasonable, proportionate, effective mitigation while taking fundamental rights into account. The phrase “systemic risk” is broad, but the law links it to specified risk assessment and mitigation processes rather than granting an unbounded censorial command. This is an example of risk terminology becoming more legitimate through procedural and substantive constraints.
Evidence of regulatory efficacy remains comparatively immature. Ofcom’s July 2026 reporting on age assurance largely covers the first months after major child-safety duties took effect. Legislatures should resist the common institutional mistake of equating implementation activity—risk assessments completed, companies contacted, notices issued—with outcome evidence showing that abuse, exploitation, self-harm exposure, or other concrete harms actually decreased.
Age verification: a genuine safety/privacy collision
The argument for age assurance is strongest where the material is both legal for adults and recognized as unsuitable for minors. In Free Speech Coalition v. Paxton, decided June 27, 2025, the Supreme Court upheld Texas's age-verification regime for commercial websites substantially devoted to sexual material deemed harmful to minors, applying intermediate scrutiny to the incidental burden on adults. That holding gives age-gating a stronger constitutional foundation in the adult-content context than a generalized identity requirement for ordinary websites or social media.
The privacy objection remains substantial. France's data-protection authority, CNIL, has warned that age verification can connect identity with intensely private browsing activity and has advocated architectures in which an independent third party verifies the age attribute without disclosing unnecessary identity information to the visited service. International data-protection principles similarly call for proportionality and collection limited to data necessary for age assurance.
Australia provides an especially valuable real-world safeguard. Its under-16 social-media regime, effective in December 2025, requires covered services to take reasonable steps concerning underage accounts, but the law prohibits platforms from making government identification the compulsory verification route; users must have a reasonable alternative. The regime is also being subjected to academic evaluation.
That produces a clear principle: prove the attribute, not the identity, whenever identity is not itself necessary. A service that only needs to know “over 18” should presumptively receive that fact, not name, passport number, home address, birth date, and a permanent cross-site identifier. Tokenized age assertions and separation between verifier and content provider are therefore not merely privacy luxuries; they are the least-restrictive means of reconciling child protection with anonymous adult inquiry.
The strongest poison-pill scenario is universal age verification becoming de facto universal identity verification: every person must identify themselves before reading ordinary lawful material, enabling activity from multiple services to be linked back to one person. The available evidence does not show this outcome is inevitable. It shows why laws should prohibit it before infrastructure and commercial incentives make it convenient.
Search filtering: user control versus invisible epistemic control
Google SafeSearch illustrates a relatively narrow filtering model. It primarily targets explicit sexual material and graphic violence, can be controlled by users or by administrators in managed contexts, and Google's own documentation says it is not intended to filter explicit material carrying substantial artistic, educational, historical, documentary, or scientific value. No filter is perfectly accurate.
That design differs sharply from a hidden system that silently demotes lawful political, scientific, or medical material based on an individualized risk model. The first is an observable category filter with user or guardian control in many contexts; the second can shape what a person is allowed to know without informing them that a judgment occurred.
There are legitimate reasons for search engines to intervene more strongly against immediate hazards. Google Safe Browsing, for example, warns users when sites are detected as dangerous, a design that preserves the user's informational choice better than secretly removing every suspicious site from the index. Warning and interstitial architectures should therefore receive explicit consideration under a least-restrictive-means requirement before outright suppression.
Research commissioned by Ofcom nevertheless shows why “do nothing” is not an adequate universal rule. Investigators examined more than 37,000 search results and found that major search engines could provide pathways toward self-injury-related material, including through coded terminology. At the same time, WHO emphasizes that digital technology's effects on young people's mental health are complex and can be beneficial as well as harmful, and that evidence supporting broad public-health prescriptions remains limited.
That combination favors query-local interventions—contextual warnings, help resources, demotion of content that actually encourages self-injury, child-specific protections, and suppression of illegal assistance—over constructing persistent psychological dossiers of everyone who searches about suicide. Someone researching suicide methods may be suicidal, may be writing a novel, may be a physician, may be a bereaved parent, or may be studying public health. The information contained in a query is often insufficient to distinguish those states safely. Evidence that harmful material exists does not itself prove the necessity of identity-linked mental-health prediction.
Child exploitation: where specificity changes the calculus
Child sexual exploitation is one of the strongest counterexamples to a maximal cognitive-liberty thesis. NCMEC received 21.3 million CyberTipline reports during 2025; it also reported more than 50,000 reports concerning financially motivated sextortion that year. These are reports rather than counts of unique verified crimes or victims and should not be interpreted as such, but the volume establishes a large, concrete safety problem rather than a hypothetical future harm.
Known-content hashing offers an unusually narrow intervention. PhotoDNA generates a signature that can be compared with hashes of images already identified as child exploitation material; Microsoft states that the hash cannot be used as facial recognition and that its cloud service is dedicated to identifying child-exploitation imagery. Because the predicate database consists of already identified illegal images, the system can be considerably more specific than open-ended semantic classification of “harmfulness.”
This does not resolve every privacy question. Scanning an entire private communications channel can still implicate privacy even when the detector itself is narrow, and expansion from known-illegal hashes to probabilistic inference about new images changes the error and surveillance profile. But the case demonstrates why a doctrine focused on lawful inquiry must distinguish between a person reading controversial material and a service detecting a known instance of illegal victimization.
Fraud and cybersecurity: concrete harms at large scale
The argument for preventive safeguards is similarly strong in fraud. The FBI says its Internet Crime Complaint Center received more than one million complaints concerning suspected internet crime in 2025, with reported losses exceeding $20 billion; approximately 453,000 cyber-enabled fraud complaints accounted for about $17.7 billion in reported losses. These are complaint-based figures and therefore neither a complete census of crime nor independently adjudicated losses, but they document harm on a scale too large to dismiss as a rhetorical pretext.
Some highly effective fraud controls do not require monitoring intellectual life at all. Transaction limits, confirmation delays, anomalous-payment alerts, account-takeover detection, stronger authentication, and rapid fund freezes target financial actions rather than ideas. The FBI has reported hundreds of millions of dollars in prevented losses from proactive victim notification programs, illustrating the value of intervention at the conduct layer.
Cybersecurity provides the same design lesson. CISA emphasizes patching exploited vulnerabilities and, in June 2026, warned that AI could compress the time defenders have between discovery and exploitation. Defenses such as multifactor authentication, vulnerability remediation, access controls, network segmentation, malware detection, and transaction-level anomaly detection can reduce harm without asking what political books someone reads.
Behavioral monitoring and automated risk assessment
Automated risk assessment becomes more troubling as prediction moves away from observable dangerous conduct toward proxies for personality, association, or thought. U.S. criminal-justice agencies have long used structured risk-assessment instruments to inform pretrial release, supervision, incarceration planning, parole, and probation decisions. Such tools are consequential because probabilistic assessments can influence liberty before a future offense occurs.
The EU AI Act's architecture reflects recognition that certain predictive systems require substantially stronger controls. Its current consolidated text requires high-risk systems to undergo risk management and testing against defined metrics and thresholds, specifies human-oversight duties, and subjects enumerated high-risk applications to documentation and conformity requirements. The Act also treats certain forms of social scoring as prohibited rather than merely “high risk.”
This is a domain where the original thesis should be strengthened, not weakened: a system deciding bail, employment, benefits, immigration treatment, insurance, or policing intensity should generally not infer dangerousness from lawful reading and search histories merely because those variables marginally improve a prediction. The stakes differ dramatically from recommending a movie. Where an automated prediction alters rights, liberty, or access, the subject should know the material factors, be able to challenge inaccurate data, receive meaningful human review, and not be penalized merely for lawful intellectual association. The EU framework's emphasis on documented risk management, defined use cases, testing, and human override provides a partial regulatory model, though it does not itself resolve every cognitive-liberty issue.
Artificial intelligence and the Patriot Act analogy
The central AI question is not whether catastrophic risk is conceivable. It is how much present-day intrusion is justified by a risk whose probability, timing, and mechanisms may be uncertain.
The analogy to September 11 is strongest at the level of political psychology and institutional design.
After a catastrophic event, low-probability future harms become psychologically immediate. Decision-makers are punished more visibly for a missed attack than for diffuse privacy costs distributed across millions of people. Intelligence programs often operate under secrecy; frontier AI evaluations can likewise depend on technical evidence that most lawmakers and citizens cannot independently reproduce. Both settings therefore create pressure to accept precautionary measures before ordinary standards of efficacy have matured. The Section 215 experience shows why independent review must subsequently test whether the specific broad measure—not merely the overarching security mission—actually made a concrete difference.
AI discourse also contains an analogous risk of elastic predicates. A policy that permits monitoring whenever activity is “dangerous,” “unsafe,” “extremist,” or “high risk” effectively delegates the real boundary to whoever controls the definition. If those categories can be changed privately, applied retrospectively, or triggered by lawful research subjects rather than harmful acts, they can become functionally equivalent to open-ended surveillance authorities.
Yet the analogy breaks in several important ways.
First, frontier AI models are not merely communication channels through which a user expresses an idea. They can be capability multipliers. The UK AI Security Institute's evaluations show measurable gains in cyber expertise and autonomous task completion; by 2025, frontier models were completing tasks designed for practitioners with more than a decade of cyber experience. That creates a plausible basis for evaluating the danger of particular assistance, not merely the person requesting it.
Second, some AI safety interventions can be implemented without persistent surveillance at all. A model can decline a highly specific operational request, reduce tool permissions, cap autonomous execution, sandbox code, require confirmation before irreversible actions, or route an unusually capable request through additional safeguards without permanently retaining the user's identity or building a profile of their interests. The liberty burden of “this model will not execute that attack chain” is therefore categorically different from “the state keeps a permanent searchable dossier because you asked about hacking.”
Third, AI risk can be benchmarked more directly than many post-9/11 threat claims. The current EU AI Act requires testing high-risk systems against defined metrics and probabilistic thresholds appropriate to their intended purpose. The UK AISI publishes capability evaluations. OpenAI publishes cyber and biological evaluations and has subjected advanced models to outside AISI testing. These systems are imperfect and sometimes provider-defined, but they create a route from “dangerous” as rhetoric to “dangerous” as an empirically testable capability proposition.
Fourth, current evidence is mixed in precisely the way a mature doctrine should expect. OpenAI's earlier GPT-4 biological-risk study found at most mild uplift in biological-threat planning, insufficient at that time to support claims of dramatic capability amplification. Later models have shown stronger capabilities, including enough operational biological assistance to trigger higher internal risk thresholds. The correct policy conclusion is therefore neither “the catastrophe was always imminent” nor “the concern was always hysteria”; the evidence changed.
Fifth, safety constraints can sometimes increase legitimate access. Anthropic has implemented a verification route intended to let bona fide cybersecurity professionals perform vulnerability research, penetration testing, and red teaming with fewer restrictions. OpenAI's 2026 safety materials similarly state that stronger cyber safeguards are paired with mechanisms intended to reduce friction for benign users; GPT-5.6 allows some blocked requests to be retried on lower-capability models. These are imperfect examples, but they embody a critical principle: when dangerous and beneficial uses overlap, create differentiated capability paths rather than simply banning the topic.
Sixth, AI controls often concern what a machine will do, not what a human is permitted to know. Preventing an autonomous agent from transferring funds without confirmation, deploying malware, ordering laboratory materials, or operating critical infrastructure is closer to a product-safety or access-control rule than a ban on thinking about those activities. Cognitive-liberty analysis should therefore be especially skeptical of input surveillance and user profiling, moderately skeptical of information refusals, and substantially more tolerant of constraints on autonomous real-world execution when consequences are severe and irreversible.
Where “dangerous,” “unsafe,” and “high risk” become warning signs
The problem is not the vocabulary itself. It is delegated discretion without operational content.
A term becomes a warning sign when several conditions cluster together: the relevant institution can redefine it unilaterally; lawful discussion is enough to trigger the designation; no causal link to concrete harm is required; the designation follows a person across contexts; classification rules are secret; no error-rate information is disclosed; no appeal exists; the data are retained indefinitely; or the same classification can be reused for employment, policing, immigration, insurance, political advertising, or other unrelated purposes.
Conversely, similar vocabulary can be legitimate when the meaning is operationally bounded. U.S. First Amendment doctrine gives “true threat” and “incitement” legally constrained meanings; Brandenburg requires advocacy to be directed to and likely to produce imminent lawless action before it loses protection as incitement. “Material support” in the terrorist-organization context is linked to a statutory framework and designated organizations, although its application to speech-like expert assistance remains controversial.
“High risk” under the EU AI Act likewise does not simply mean “an official is worried.” Article 6 ties high-risk status to specified product-safety contexts and enumerated use cases, and the associated regime imposes testing, documentation, risk management, human oversight, and registration obligations. For general-purpose AI, the Act's concept of “systemic risk” is linked to high-impact capabilities and large-scale harms rather than being left wholly undefined.
“Self-harm” can similarly range from vague moderation rhetoric to an actual offense definition. The UK's 2026 regulatory update addresses the offense of encouraging or assisting serious self-harm, a considerably narrower predicate than mere discussion of depression or recovery. At the same time, Samaritans emphasizes that online discussion of suicide can be both harmful and beneficial and that effects depend on content and context. This strongly counsels against equating every query or conversation about suicide with dangerousness.
“Child sexual abuse material” can be especially precise where the intervention uses a database of already identified illegal images and exact or robust hash matching. That level of objective predicate is qualitatively different from a model assigning a free-form probability that an ordinary image or conversation is “unsafe.”
The governing rule should therefore be:
Risk words are not safeguards. A definition becomes a safeguard only when it restricts discretion sufficiently that an independent party can determine whether the intervention was authorized.
The Patriot analogy: final score
| Dimension | Analogy strength | Why |
|---|---|---|
| Salient catastrophe drives precaution | Strong | Both terrorism and frontier AI can produce highly salient tail-risk reasoning before the frequency of actual events is well known. Post-9/11 experience shows the institutional consequences of that asymmetry. |
| Broad terms can expand authority | Strong | Intelligence authorities and modern risk regimes both depend heavily on definitions, scope, and oversight. DOJ OIG findings show how authority can exceed compliant practice even without deliberate abuse. |
| Secrecy obstructs efficacy review | Moderately strong | Intelligence is necessarily classified; frontier-model evaluations can also be inaccessible because of proprietary models and dangerous test details. Independent evaluators such as AISI partly mitigate the latter problem. |
| Broad surveillance is necessary to achieve safety | Weak analogy | Many AI controls can operate locally at the request, tool, or action layer without bulk monitoring. |
| Threat is purely speculative | Weakening rapidly | Advanced cyber capabilities are empirically observable in current systems, although catastrophic extrapolations remain uncertain. |
| Government coercion and private model policy are equivalent | Weak | A state intelligence database and a provider's refusal to perform a specific operation impose different kinds and degrees of liberty burden. |
| Restrictions inevitably become permanent | Weak | Historical authorities have been terminated or narrowed; software safeguards are also technically reversible. |
| Efficacy claims require adversarial testing | Very strong | Section 215 is a prime example of why broad safety claims should be tested against actual marginal benefit. |
The analogy is therefore best treated not as “AI safety is the new Patriot Act”, but as:
“The Patriot Act era demonstrates the institutional failure modes that AI safety policy should be designed in advance not to repeat.”
The strongest opposing case and the least-restrictive intervention test
A genuine adversarial investigation must make the strongest case against the thesis.
That case begins with a moral point: freedom is not maximized by refusing to prevent foreseeable victimization. A person terrorized by death threats, a child whose abuse imagery is repeatedly circulated, a teenager targeted by sextortion, an elderly victim losing life savings to automated fraud, or a hospital disabled by ransomware has also lost meaningful freedom. NCMEC's exploitation data, FBI fraud statistics, WHO suicide figures, and rapidly increasing frontier-AI cyber capabilities establish that the harms invoked by safety institutions are often real rather than manufactured.
Prevention also cannot logically be restricted to acts after irreversible harm. Counterterrorism that can act only after detonation is useless. Cybersecurity that waits for exfiltration before blocking access is defective. Child-protection systems that may intervene only after abuse material spreads globally ignore the continuing harm of redistribution. A safety doctrine must therefore permit action against sufficiently reliable precursors, capabilities, and attempts—not merely completed crimes.
American constitutional law already reflects this balance. True threats, incitement meeting the Brandenburg standard, child pornography, and certain coordinated forms of assistance to designated terrorist organizations can be regulated even though communication is involved. The law does not equate “speech” with absolute immunity from consequences when speech itself becomes part of a proscribed harmful act.
Child protection also exposes a weakness in pure adult-autonomy models. Children do not have identical developmental capacity or legal status, and online systems may algorithmically deliver material without deliberate searching. WHO describes both benefits and harms from digital environments, while Ofcom's research documents pathways by which children encounter self-harm, suicide, and eating-disorder material. A platform can therefore have legitimate reasons to alter recommendations for minors without treating the child's private thoughts as evidence of deviance.
Likewise, AI creates an unusual dual-use problem: the same detailed assistance that helps a researcher secure software or understand biology can sometimes lower barriers for an attacker. Current AISI evaluations make the capability progression measurable. The rational answer cannot be “never restrict anything until an attack is completed.” It must instead distinguish benign and offensive pathways more accurately, maintain access for legitimate experts, and concentrate stronger controls where the model contributes high-leverage operational capability.
The strongest opposing argument therefore defeats cognitive-liberty absolutism. It does not defeat the narrower poison-pill thesis. Rather, it establishes why the appropriate standard is a demanding least-restrictive-means framework.
Comparative intervention judgment
| Harm | Evidence of risk | Intervention favored under adversarial review | Intervention disfavored absent extraordinary evidence |
|---|---|---|---|
| Terrorism | Demonstrated catastrophic consequences; foreign-intelligence collection can have genuine security value. Reform advocates themselves acknowledge Section 702's value. | Targeted intelligence collection; judicially supervised U.S.-person access; specific investigation of threats, financing, operational networks. | Indiscriminate domestic intellectual profiling or bulk collection whose marginal benefit cannot be demonstrated. Section 215 is the cautionary example. |
| Child sexual exploitation | 21.3 million CyberTipline reports in 2025; significant sextortion reporting. | Known-CSAM hash matching, victim reporting, account/network disruption, targeted investigation. | General inspection or permanent retention of all lawful private communications merely to infer sexual-risk profiles. |
| Fraud | More than $20 billion in reported IC3 losses in 2025. | Transaction friction, anomaly detection, account security, rapid victim notification, recovery mechanisms. | Profiling unrelated reading or political behavior as a fraud proxy. |
| Cyberattacks | CISA documents active exploitation; AI cyber capability is increasing quickly. | Sandboxes, tool permissions, rate limits, vulnerability defenses, narrow restrictions on high-impact offensive operations, professional access paths. | Topic-wide bans on cybersecurity knowledge or automatic suspicion of everyone researching vulnerabilities. |
| Suicide and self-harm | More than 720,000 suicide deaths annually worldwide; online material can both help and harm. | Crisis resources, child-specific recommender protections, suppression of content actually encouraging serious self-harm, user-controlled filters. | Persistent mental-health dossiers inferred from isolated queries, particularly when shared with unrelated institutions. |
| Violent threats | True threats are a recognized regulable category. | Context-sensitive threat detection followed by human review and proportionate escalation. | Treating controversial political ideology or violent fiction itself as a threat. |
| Harmful search results | Search can surface malware, explicit material, and self-harm content. | Warnings, SafeSearch-style controls, child protections, clear removal rules for unlawful material. | Invisible individualized censorship based on inferred worldview. |
| Adult-content access by minors | Recognized state interest; Supreme Court upheld a Texas age-gating law in the specific commercial pornography context. | Anonymous/tokenized proof of age, separated verifier, no mandatory government ID where alternatives suffice. | Universal identity-linked browsing records. |
| High-stakes automated decisions | Risk tools influence criminal-justice and other consequential decisions; EU law treats many uses as high-risk. | Transparent factors, validation, human override, contestability, narrow-purpose data. | Secret scores based on lawful reading, association, belief, or inferred personality. |
| Advanced AI assistance | Current systems exhibit increasing cyber and biological capabilities. | Capability-specific safeguards, sandboxing, staged access, independent evaluation, verified expert exceptions where justified. | Generalized surveillance of all user curiosity or refusal based solely on controversial subject matter. |
The critical insight is that the least intrusive intervention often operates closer to the harmful act than to the person's mind. Financial abuse is best interrupted at money movement. Cyberattacks can be interrupted at exploit execution, credential use, or network access. Dangerous AI agents can be constrained at tool invocation and external action. Known illegal media can be identified by hashes. True threats can be analyzed as communications directed at victims. None of these requires treating ordinary intellectual exploration as presumptive evidence of dangerousness.
There are exceptions. A credible terrorist investigation or imminent threat may legitimately require communications surveillance before the ultimate act occurs. But the exception should be justified by individualized evidence and legal process, not converted into a standing proposition that everyone must be continuously observed because anyone could someday become dangerous.
The mandatory evidence test
Every intervention proposed under the doctrine below should answer all of these questions before deployment and again at renewal:
| Required showing | Minimum acceptable question |
|---|---|
| Necessity | What specific harmful outcome cannot adequately be addressed under existing narrower mechanisms? |
| Documented risk | What measured incidents, prevalence data, capability evaluations, or credible threat models establish the problem? |
| Proportionality | How does the expected reduction in harm compare with false positives, privacy invasion, chilling effects, exclusion, and security risks created by the measure? |
| Narrow scope | What persons, content, capabilities, services, and time periods are explicitly outside the intervention? |
| Least-restrictive means | Why will warnings, user controls, transaction friction, sandboxing, rate limits, anonymous credentials, targeted warrants, or narrower classifiers not suffice? |
| Transparency | Can the public understand what triggers the intervention and see aggregate data about its use and errors? |
| Independent review | Can an institution independent of the operator examine evidence, algorithms, logs, and claimed efficacy? |
| Appeal rights | Can an affected person obtain timely review and correction by a human with authority to reverse the decision? |
| Deletion | When are queries, prompts, identifiers, scores, and supporting data destroyed? |
| Protection of lawful inquiry | Is there an explicit rule preventing research, journalism, education, fiction, curiosity, or controversial belief from being treated as harmful conduct without additional evidence? |
| Prevention of secondary use | Can data collected for child protection, cybersecurity, fraud prevention, or AI safety later be reused for policing, immigration, employment, advertising, insurance, or political profiling? The default should be no. |
| Sunset | Does the authority terminate unless affirmative evidence justifies renewal? |
| Measured efficacy | What outcome will prove that the intervention reduced actual harm rather than merely increasing removals, flags, surveillance volume, or compliance activity? |
Section 215 is especially important here because the ultimate question was not whether terrorism was dangerous; it obviously was. The relevant question was whether bulk telephone-record acquisition added enough marginal counterterrorism value to justify its breadth, and the independent review found only limited value.
That is the standard that should migrate into AI governance.
A Cognitive Liberty Doctrine for the AI Age
The evidence supports the following governing principle:
Safety measures should target demonstrable harmful conduct, or demonstrably dangerous capabilities closely connected to such conduct, as narrowly as practicable without turning lawful thought, curiosity, reading, searching, research, discussion, or private intellectual exploration into objects of generalized surveillance and control.
This principle requires a doctrine more specific than “balance safety and freedom.”
Lawful inquiry receives a presumption of non-suspicion. Reading about terrorism is not terrorism. Researching suicide is not proof of suicidal intent. Studying malware is not unauthorized computer access. Exploring extremist ideology is not a true threat. Reading about controlled substances is not manufacturing them. Asking an AI about a dangerous domain should not, by itself, create a durable security profile. U.S. doctrine protecting receipt of information and associational privacy supplies a constitutional analogue for this presumption even though “cognitive liberty” is not itself a unitary enumerated right.
Conduct should be regulated before identity whenever feasible. Systems should first ask whether a dangerous transaction, exploit, threat, upload, or autonomous action can be interrupted directly. Identity verification should be required only when accountability or eligibility genuinely depends on identity. Age-only contexts should ordinarily use age attributes rather than full identity; Australia's prohibition on compulsory government ID for its social-media age regime demonstrates that such legal separation is practical.
Attribute proofs should replace dossiers. Where eligibility matters—age, professional qualification, authorization to administer a system—the service should obtain the minimum necessary assertion rather than a reusable identity bundle. CNIL’s recommended third-party approach and international data-protection principles point toward separation between the verifier and the content provider.
Safety data should be purpose-bound. A prompt retained to investigate a cybersecurity abuse incident should not silently become evidence for advertising, insurance, employment, immigration, political targeting, or an unrelated criminal investigation. Exceptions should require a new lawful predicate, not merely technical availability. This is especially important because the principal long-term danger of safety infrastructure is often not its initial purpose but the low marginal cost of secondary use once comprehensive data exist.
The burden of justification should rise with distance from harmful conduct. A known CSAM hash has a direct relationship to illegal material. A transaction to a confirmed fraud destination is closely related to fraud. A concrete exploit sequence against an external system is closer to cyber harm than a general question about vulnerability research. A person's reading history is much farther away. The farther an intervention operates from actual conduct, the stronger its evidentiary and procedural safeguards should be.
Human-risk scores should not become universal passports. Scores created for one domain should never silently follow a person into unrelated services. Social scoring is particularly dangerous because it converts behavior across contexts into generalized eligibility or trustworthiness. The EU AI Act's prohibition of specified social-scoring practices reflects the severity of this concern.
High-impact automation must be contestable. Where automated risk assessment affects detention, employment, public benefits, education, immigration, or similarly consequential interests, people should receive meaningful notice, the material basis of adverse decisions, correction rights for inaccurate data, and human authority capable of overriding the system. EU high-risk-system rules concerning documentation, testing, human oversight, and fundamental-rights impact assessment offer useful structural precedents.
AI safety should favor capability boundaries over subject-matter taboos. “Tell me what ransomware is” and “autonomously deploy ransomware against this network” differ radically. The safest architecture is therefore one that remains permissive for education and defensive research while introducing stronger friction as a request becomes operational, scalable, autonomous, evasive, or targeted at real systems. Current AISI testing and provider programs for legitimate cybersecurity professionals demonstrate that capability-sensitive differentiation is technically possible.
Execution deserves stronger controls than information. An AI agent with authority to run shell commands, move money, order materials, send messages, or manipulate critical systems can create harm independently of whether the underlying information is publicly knowable. Tool permissions, sandboxing, confirmation requirements, spending limits, rate limits, and reversible staging should therefore be preferred to broad restrictions on knowledge whenever they address the risk effectively.
When content controls are needed, warnings and user agency should be tested before silent suppression. Google Safe Browsing's warning approach and configurable SafeSearch illustrate interventions that can reduce exposure without necessarily deleting underlying information. For children or genuinely illegal material stronger controls may be justified, but least-restrictive analysis should remain explicit.
Exceptional authorities must expire by design. The Section 215 history demonstrates both the danger of broad emergency tools and the value of later legislative correction. AI emergency powers, mandatory monitoring regimes, extraordinary identity requirements, or model-access restrictions justified by rapidly evolving threat conditions should have short renewal periods, independent evidence requirements, and automatic termination absent affirmative legislative or regulatory renewal.
Safety success must mean harm reduction, not surveillance production. A program does not become effective because it flagged one billion messages, removed ten million posts, collected a larger dataset, produced more suspicious-activity reports, or increased the percentage of users who completed ID checks. The relevant metric is whether exploitation, fraud, successful attacks, unwanted child exposure, violent incidents, or another specified outcome fell—and whether the reduction is plausibly attributable to the intervention.
Error costs must be measured in both directions. Safety institutions naturally track false negatives because a missed attack or abuse case is visible. Cognitive-liberty governance must force measurement of false positives too: legitimate researchers denied access, users incorrectly flagged, minority language incorrectly classified, lawful sites blocked, innocent accounts suspended, people deterred from sensitive searches, and personal data breached because identity collection created a new target.
Independent auditors must be able to see what the public cannot. Some threat evidence cannot responsibly be published in full, particularly exploitable cyber or biological details. Transparency therefore need not mean exposing dangerous operational information. It should mean that a security-cleared, technically competent, institutionally independent reviewer can test the evidence and publicly report whether the intervention remains necessary and proportionate. AISI's role as an outside evaluator of frontier-model capabilities shows one emerging version of this separation.
Research exceptions should be real rather than ornamental. Legitimate researchers, journalists, historians, security professionals, clinicians, and educators frequently investigate the same topics that malicious actors use. Appeals, verified-access mechanisms where truly necessary, and explicit research protections should therefore be designed into the system rather than improvised after false positives occur. Anthropic's cybersecurity verification pathway is a concrete example of this approach.
No authority should be justified indefinitely by a catastrophe that has not occurred. Prospective risk can justify prospective safeguards, particularly where consequences are irreversible. But predicted catastrophe must periodically be re-estimated as empirical capabilities, attack frequency, defenses, and scientific knowledge change. OpenAI's biological-risk evaluations changed between earlier and later model generations; AISI's cyber measurements likewise show capabilities evolving quickly. A rational regime must therefore become more or less restrictive as evidence changes, not preserve its original assumptions as doctrine.
Under this doctrine, the red line is not “no safety.” It is generalized intellectual surveillance without demonstrated necessity.
Evidence ledger, unresolved questions, and final verdict
The user-requested distinctions are important because several propositions in this debate have radically different epistemic status.
| Category | What the investigation establishes |
|---|---|
| Documented facts | Emergency national-security measures have restricted expression and liberty; post-9/11 surveillance authorities experienced documented compliance failures; Section 215 bulk collection was independently judged to have limited value and was terminated; current digital systems use moderation, age assurance, filtering, automated risk tools, and AI capability safeguards; large-scale child exploitation, fraud, suicide, and cyber harms are empirically documented. |
| Primary-source evidence | National Archives records on EO 9066; Library of Congress material on wartime censorship; DOJ OIG reports; government PCLOB hearing records; statutory and regulatory texts from the UK and EU; FBI/NCMEC statistics; Ofcom and eSafety implementation documents; provider system cards and AISI evaluations. |
| Disputed claims | The exact national-security value of some intelligence programs; the magnitude and causality of surveillance chilling effects; the causal contribution of social media to population-level mental-health outcomes; the probability of catastrophic AI outcomes; how often current AI safeguards stop real attacks rather than benchmark attacks. The evidence supports concern but not confident universal numerical estimates. |
| Institutional incentives | The record is consistent with institutions systematically weighting visible safety failures more heavily than diffuse privacy or false-positive costs. OIG findings show that overreach can arise from structure, guidance, and oversight failures without proving malign intent. Claims about incentives should therefore be treated as institutional analysis, not accusations about individual motives. |
| Technical possibilities | Hash matching can detect previously identified illegal imagery; age attributes can be separated from identity; filters can operate at query/result level; behavioral data can be used to infer sensitive states; AI controls can operate at request, capability, tool, or account layers; high-capability professional access can be separated from ordinary access. |
| Historical analogies | The strongest analogy is institutional: catastrophe can lower resistance to broad preventive infrastructure whose marginal utility is initially difficult to test. The weakest analogy is technological: AI models can themselves create or execute dangerous capabilities, whereas post-9/11 surveillance largely concerned government access to communications and records. |
| Counterarguments | Prevention before completed harm can be necessary; targeted moderation can work; children face genuine exploitation and exposure risks; fraud and cyber losses are large; AI capability increases are measurable; some identity or verification controls are justified in bounded contexts. |
| Recommendations | Place the evidentiary burden on the institution imposing the restriction; prefer action-layer controls to thought-layer monitoring; minimize identity; forbid unrelated secondary use; make high-impact automation appealable; use independent evaluation; require deletion and sunsets; publish efficacy and error measures; explicitly protect lawful inquiry. These recommendations are normative conclusions derived from the documented failure modes and successful narrower alternatives above. |
Several unresolved questions are especially important.
The effectiveness gap remains large. Online-safety regulation is moving faster than long-term outcome research. Ofcom had only early implementation evidence available for major age-assurance duties by mid-2026. The possibility that a policy is sensible does not substitute for evidence that it reduces the specified harm.
Privacy-preserving age assurance needs comparative field evidence. Tokenized and third-party architectures are technically and conceptually attractive, but evidence comparing circumvention rates, false rejection, biometric bias, privacy breaches, and actual reductions in minor access remains insufficient for declaring one design universally superior. CNIL explicitly notes both circumvention and intrusiveness problems in existing techniques.
AI capability-to-harm translation is still uncertain. Benchmarks show that models can complete increasingly difficult cyber tasks, but the mapping from benchmark performance to additional real-world attacks, casualties, or catastrophic probability remains uncertain. Provider reports of real malicious use are evidence that the pathway exists; they are not yet a population-level causal estimate of how much AI increases total harm.
The point at which verified access becomes identity infrastructure is unresolved. Professional verification can preserve dual-use access, but widespread identity-gated AI could itself produce the poison-pill dynamic if every sensitive topic required a verified real-world identity. The right long-term design may involve privacy-preserving credentials proving authorization or expertise without global identity linkage; current age-assurance work demonstrates the relevant architectural principle but not a complete AI solution.
The constitutional boundary for future AI intermediaries remains unsettled. Existing First Amendment doctrine robustly protects receipt of information from government interference in important settings, yet it does not create a general entitlement to force every private publisher, search engine, or AI provider to supply every requested output. Any future regime that government mandates, however, raises different questions from voluntary provider design, particularly if government pressure effectively determines which lawful ideas citizens may access. The existing doctrine on receipt of information and protected/private versus proscribable categories provides principles, not a complete answer for generative AI.
Where the poison-pill hypothesis survives
The hypothesis survives most strongly in this formulation:
A safety regime becomes presumptively dangerous to cognitive liberty when it makes lawful intellectual activity itself a durable input into state or institutional suspicion, especially when the data are identity-linked, retained, combined across contexts, governed by elastic risk categories, unavailable for independent inspection, and reusable for purposes beyond the harm that originally justified collection.
The historical evidence supports every major element of that warning except inevitability. Emergency and security rationales have enabled extreme overreach. Post-9/11 authorities produced documented compliance failures. Broad collection has sometimes failed a serious marginal-utility test. Surveillance can chill information seeking. Secondary use and expansion are rational institutional risks once data and authority exist.
The thesis also survives strongly against identity-linked universal access architectures. There is rarely a legitimate reason for an ordinary adult's full identity to become a prerequisite to routine reading and searching merely because some fraction of users are children or bad actors. Existing age-assurance work shows that attributes can be proven with substantially less disclosure, and Australia's explicit protection against compulsory government-ID verification confirms that legislatures can encode that principle.
It survives against automated risk systems that convert lawful curiosity into future-danger predictions. High-stakes decisions require transparent, bounded predicates and contestability. The more a system infers dangerousness from associations, searches, reading, or personality rather than observable conduct, the closer it comes to the core cognitive-liberty problem.
It survives against AI catastrophe rhetoric used as a substitute for empirical thresholds. Catastrophic outcomes may warrant exceptional safeguards, but the word “catastrophic” cannot itself establish necessity. Capability evaluations, attack demonstrations, causal pathways, exposure analysis, safeguard testing, and periodic reassessment must do that work. Current AISI and provider evaluations show that such measurement is increasingly possible.
Where the hypothesis must be narrowed
The evidence requires rejecting the idea that safety and cognitive liberty are generally antagonistic.
Removing known child sexual abuse material protects victims without meaningfully advancing society's interest in unrestricted lawful inquiry. Preventing credible threats protects the target's liberty. Blocking a fraudulent transfer does not censor the victim's thoughts. Preventing an AI agent from autonomously attacking a network does not prevent a student from studying network security. Targeted child-safety design can prevent involuntary exposure while leaving adult access intact.
The evidence also requires rejecting a rigid rule that only completed harmful conduct may trigger intervention. Credible imminent threats, attack execution, coordinated terrorism assistance, malware deployment, financial transfers, exploit chains, and some high-leverage AI capability requests can justifiably be intercepted before the ultimate injury occurs. The proper constraint is proximity and demonstrability, not temporal lateness.
And it requires rejecting the assumption that all safety terminology is empty euphemism. “True threat,” Brandenburg incitement, known CSAM hashes, enumerated EU “high-risk” AI uses, and specific criminal offenses such as assisting serious self-harm demonstrate that dangerousness concepts can be operationally bounded.
Final adversarial conclusion
The deepest lesson from the Patriot Act era is not “never act under uncertainty.” Governments and private institutions sometimes must act before uncertainty disappears.
It is:
never let the gravity of the feared outcome substitute for evidence that the chosen intrusion is necessary.
The burden must remain intervention-specific. Terrorism does not prove bulk collection is necessary. Child exploitation does not prove universal identity-linked browsing is necessary. Suicide does not prove everyone's mental state should be inferred from their searches. Fraud does not justify general behavioral dossiers. Cyberattacks do not justify treating cybersecurity education as suspicious. The possibility of catastrophic AI misuse does not by itself justify permanent monitoring of ordinary human curiosity.
At the same time, cognitive liberty cannot become its own poison pill. A doctrine so rigid that it prohibits known-CSAM detection, targeted threat intervention, anti-fraud controls, child-specific protections, secure authentication, malware defenses, or safeguards against demonstrably high-leverage AI operations would sacrifice real people to an abstraction. The empirical record does not support that position.
The version of the thesis that survives is therefore narrower, stronger, and harder to dismiss:
Safety is legitimate when it demonstrably reduces a defined harm using the narrowest practicable intervention. Safety becomes a poison pill for cognitive liberty when exceptional danger is used to make ordinary lawful intellectual life permanently observable, classifiable, identity-linked, and governable—without proof that such surveillance is necessary, proportionate, effective, temporary, and resistant to secondary use.
That conclusion is not an argument for trusting people with every dangerous capability. It is an argument for locating control at the point where capability becomes harmful conduct, as close to the harm as technically and legally practicable, while preserving the largest feasible domain in which a person may read something difficult, ask something uncomfortable, investigate something frightening, test an unpopular hypothesis, or privately wonder about the world without thereby becoming a permanent object of suspicion.