Civic / Privacy / Digital Rights

The Real-World “Pre-Crime” Landscape: A Comparative Inventory and Governance Framework for Predictive Law Enforcement

Report summary

This report inventories and evaluates predictive policing, algorithmic threat assessment, security watchlisting, behavioral threat assessment, biometric surveillance, traveler-risk analysis, and related systems documented through August 2, 2026 . It covers 36 programs in 15 countries , spanning fede

Status
Research archive item
Category
Civic / Privacy / Digital Rights
Length
15,414 words
Reading time
71 minutes
Report type
evaluation

Key topics

  • Civic / Privacy / Digital Rights
  • Civic
  • Privacy
  • Digital Rights
  • AI
  • Runtime
  • Research Archive
  • Strategy
  • Audit

Research provenance

Archive status
Research archive item
Content identity
sha256:96af19a1cdc883359be55654ba4b7d9da483fb711ff07a08bb6cb2fbc2b036f8

For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.

Source availability: 83 citation markers in the source export have no recoverable source links. Those markers are omitted from this reader; any supplied bibliography and ordinary links remain. Check the original sources before relying on the cited claims.

This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.

Full report

On this page

Scope, method, definitions, and principal findings

This report inventories and evaluates predictive policing, algorithmic threat assessment, security watchlisting, behavioral threat assessment, biometric surveillance, traveler-risk analysis, and related systems documented through August 2, 2026. It covers 36 programs in 15 countries, spanning federal, national, state, provincial, municipal, and county jurisdictions. The inventory is intentionally broader than products marketed as “predictive policing.” Agencies increasingly describe comparable practices as intelligence-led policing, precision policing, risk assessment, threat management, targeting, harm assessment, dynamic analysis, decision support, or resource optimization. Those labels do not determine whether a system exerts preemptive power.

The central finding is that “pre-crime” is not a single technology. It is a decision architecture linking data about past events, associations, identity, travel, geography, speech, or behavior to anticipatory state action. The relevant unit of analysis is therefore not merely an algorithm. It is the full chain:

collection → identity resolution → inference or prioritization → human interpretation → intervention → retention → review or redress.

A system that forecasts burglary locations but only informs patrol planning differs fundamentally from a secret individual watchlist that produces travel denial. A behavioral-threat team that develops a contextual case assessment differs from a model that assigns a person a probability of future violence. Facial recognition may identify a person without predicting conduct, yet becomes pre-crime infrastructure when identification automatically triggers a risk classification, checkpoint restriction, questioning, or arrest.

Principal conclusions

First, the strongest evidence supports a narrow proposition: geographically focused analysis can sometimes help allocate patrol resources, but the public evidence does not establish that proprietary crime forecasts consistently outperform transparent hotspot methods, professional analysis, or carefully designed problem-oriented policing. Even positive studies generally evaluate combined interventions rather than the algorithm alone, and reported crime is partly a measure of police deployment and detection. Systematic reviews continue to find heterogeneous, methodologically inconsistent results. [Confidence: 92/100.]

Second, person-based systems create much greater rights risks than place-based forecasting because a score can persist across encounters and influence stops, visits, surveillance, diversion eligibility, bail or custody judgments, border questioning, and information sharing. Chicago’s Strategic Subject List did not demonstrably reduce homicide and was associated with increased arrests; Los Angeles’s LASER suffered inconsistent selection and removal practices; Pasco County’s program escalated predictions into repeated visits, code enforcement, and family pressure; Australia’s STMP was found to involve serious maladministration in its treatment of children. [Confidence: 96/100.]

Third, watchlisting and traveler-risk systems are among the most consequential and least measurable programs. Their outputs can affect boarding, screening, border admission, detention, and investigative attention, but public precision, recall, calibration, and false-positive statistics are generally unavailable. Litigation and government audits document erroneous nominations, identity confusion, severe travel consequences, and historically inadequate redress. [Confidence: 97/100.]

Fourth, behavioral threat assessment should not automatically be equated with statistical prediction. The FBI and Secret Service describe threat assessment as dynamic, contextual, behavior-based inquiry rather than profiling or a claim that violence can be predicted with certainty. Properly designed multidisciplinary teams can connect a person to services and evaluate escalating conduct. Risks arise when vague indicators, protected speech, undisclosed intelligence, or school and workplace reporting are converted into durable law-enforcement suspicion without evidentiary thresholds or appeal. [Confidence: 91/100.]

Fifth, the most intense systems are not necessarily the most technically sophisticated. Pasco County’s relatively simple list-driven program and London’s former Gangs Violence Matrix could produce serious consequences despite limited machine learning. Conversely, sophisticated analytic tools may remain low-intensity if they produce aggregate planning information without identifying individuals or triggering coercive decisions. [Confidence: 94/100.]

Sixth, opacity is not merely a transparency problem; it prevents scientific evaluation. When agencies withhold the target definition, denominators, thresholds, interventions, override rates, or downstream decisions, neither effectiveness nor fairness can be determined. Vendor accuracy claims, counts of “hits,” arrests after alerts, or crimes found in forecast areas are not substitutes for prospective, independently audited outcome measures. [Confidence: 98/100.]

Seventh, rebranding is now a major governance problem. Chicago moved from the Strategic Subject List to other violence-management approaches; London discontinued the Gangs Violence Matrix and introduced the Violence Harm Assessment; Los Angeles abandoned LASER and PredPol while retaining data-driven policing; the United Kingdom’s national analytics work shifted away from an explicitly predictive serious-violence model toward broader analytics infrastructure. Functional continuity may persist despite a changed name, procurement vehicle, model, or administrative owner. [Confidence: 93/100.]

Analytical distinctions

This report separates five operations frequently conflated in policy debate:

OperationMeaningIllustrative exampleRequired evidentiary question
IdentificationDetermining who or what is presentFacial recognition or identity matchingIs the match accurate, lawfully obtained, and verified?
CorrelationFinding statistical or network relationshipsCo-travel, co-arrest, communication, or geographic clusteringDoes the relationship have operational meaning, or is it incidental?
PrioritizationOrdering cases, people, or placesA threat-assessment queue or high-risk worklistIs the ranking valid for the intended decision?
PredictionEstimating an unobserved or future outcomeProbability of violence, recidivism, victimization, or burglaryIs the target well defined, calibrated, and prospectively validated?
Automated actionExecuting or materially determining a consequenceBoard/no-board response, watchlist routing, or automated alertIs human review meaningful, timely, and empowered to reverse the result?

The report uses ten functional categories requested by the commission:

C1 place-based crime forecasting; C2 person-based offending or victimization prediction; C3 recidivism or custody risk assessment; C4 terrorism or security watchlisting; C5 behavioral threat assessment; C6 association and network analysis; C7 biometric identification; C8 traveler or border risk scoring; C9 social-media or open-source monitoring; and C10 resource allocation or patrol optimization.

A program may occupy several categories. Classification does not imply that every system uses artificial intelligence or makes a probabilistic forecast.

Research method and confidence scoring

The inventory prioritizes official policy documents, privacy-impact assessments, government evaluations, inspector-general reports, legislative records, court decisions, technical studies, procurement records, and peer-reviewed research. Investigative reporting is used when official records are incomplete, particularly for covert or discontinued programs, but is marked accordingly. Program status is based on the latest located evidence, not on historical summaries.

Program confidence scores estimate the reliability of the descriptive record:

ScoreInterpretation
90–100Official existence, authority, status, and principal operation are well documented; important gaps may remain about performance or thresholds.
75–89Existence and general operation are corroborated, but some fields such as vendor, current use, algorithm, or interventions remain incomplete.
50–74Credible evidence exists, but reliance on investigative, advocacy, vendor, or foreign-language reporting limits certainty.
Below 50Program is reported or proposed, but operational status or core design cannot be reliably verified.

“Not available” means no reliable public metric was located. It does not mean the agency lacks internal data.

The principal limitations are substantial. National-security programs remain partly classified; police often refuse to disclose thresholds or names on operational-security grounds; decentralized procurements make national coverage difficult to determine; discontinued systems may be replaced without clear public notice; non-English sources are unevenly indexed; and vendor descriptions frequently merge marketing claims with evaluation claims. China’s security architecture, Israel’s operations in the occupied Palestinian territory, India’s rapidly changing state-level projects, France’s municipal deployments, Japan’s social-media monitoring, and South Korea’s Pre-CAS have particularly important documentary gaps.

Global program database and jurisdictional comparison

Global program database

The following database records the official or best-established program name, authority, dates and status, purpose, inputs, outputs, affected decisions, automation, geographic or individual focus, vendor where known, legal basis, evaluation, oversight, consequences, categories, and descriptive confidence.

Jurisdiction and programAuthority, dates, and statusOperation: inputs, outputs, decisions, and automationVendor, legal basis, evaluation, oversight, and documented consequencesCategories; confidence
United States — Terrorist Screening Dataset, formerly Terrorist Screening DatabaseFBI Terrorist Screening Center; created after HSPD-6 in 2003. TSDS name in use since 2021. Active.Consolidates biographic and biometric identifiers for persons nominated under watchlisting standards. Records support identity matching and screening by federal, state, local, territorial, tribal, and foreign partners. Consequences may include additional screening, border action, investigative referral, or No Fly/Selectee treatment. Automated matching is combined with analyst and agency review.Government-developed federated architecture; legal authorities include HSPD-6, agency statutes, screening authorities, and classified or controlled nomination guidance. The FBI says nominations generally require reasonable suspicion and cannot rest solely on race, ethnicity, religion, or protected First Amendment activity. Aggregate error and performance metrics are not routinely published. Courts and audits document mistaken identity, erroneous nomination, travel disruption, and contested redress.C4, C6, C7, C8; 97/100
United States — TSA Secure FlightTransportation Security Administration; phased implementation during the late 2000s. Active.Receives airline passenger data and matches travelers against security lists before boarding. Produces boarding-pass instructions and enhanced-screening or denial outcomes. Identity matching is substantially automated; final resolution may involve TSA and TSC personnel. Individual focus.Federal aviation-security statutes and TSA authorities; government systems and contractors. GAO repeatedly identified implementation, data-quality, identity-resolution, redress, and performance-measurement concerns. Public precision, recall, and demographic error rates are unavailable.C4, C7, C8; 94/100
United States — Quiet Skies and Silent PartnerTSA and Federal Air Marshal Service. Silent Partner and Quiet Skies were documented in DHS privacy assessments; Quiet Skies terminated June 5, 2025, while ordinary TSA vetting continues.Used travel patterns, intelligence, watchlist relationships, and behavioral indicators to select travelers, including some not on the principal watchlists, for enhanced screening or in-flight observation. Outputs could generate federal-air-marshal surveillance. Mixed rule-based and human observation.DHS privacy authorities and aviation-security powers. Detailed rules remained controlled. DHS stated when terminating Quiet Skies that it had not stopped a terrorist attack; that statement is not a causal evaluation. Political-targeting allegations and the absence of transparent performance measures underscore governance failure.C4, C5, C8; 95/100
United States — Automated Targeting System–PassengerU.S. Customs and Border Protection. ATS dates to the 1990s; passenger modules expanded after September 11. Active.Compares API, PNR, border, immigration, law-enforcement, intelligence, travel, document, and related data against targeting rules and risk indicators. Produces referrals and risk assessments affecting inspection, questioning, detention, search, admissibility, and investigative attention. Automated decision support with officer review.Customs, immigration, border-security, and privacy authorities; CBP-developed and contractor-supported. Privacy-impact assessments describe purpose and data flows but not rules, weights, precision, recall, or disparate-impact statistics. Oversight includes DHS privacy and inspector-general mechanisms, congressional review, and courts, but operational transparency is low.C4, C6, C7, C8; 96/100
United States — FBI Behavioral Threat Assessment CenterFBI Critical Incident Response Group; national multiagency, multidisciplinary center. Active.Receives referrals concerning targeted violence and terrorism; examines behavior, context, grievances, stressors, capability, access to weapons, communications, and protective factors. Produces case consultation and threat-management recommendations. It is structured professional judgment, not an automatic violence probability.Federal investigative and prevention authorities. No single commercial vendor. FBI guidance explicitly rejects profiling and deterministic prediction. Public evaluation is mainly descriptive; no comprehensive national causal study establishes how many attacks are prevented. Oversight follows FBI, DOJ, and interagency processes.C5, C6, C9; 95/100
United States — Secret Service National Threat Assessment Center and Behavioral Threat Assessment UnitsU.S. Secret Service. NTAC established in 1998; behavioral-threat functions subsequently expanded. Active.Conducts research, training, consultation, and operational threat assessment regarding attacks on protected persons, schools, workplaces, and public targets. Inputs include communications, planning behavior, weapons access, grievances, life stressors, and third-party reports. Outputs guide management, referral, protection, and investigation. Human multidisciplinary judgment.Secret Service protective and investigative statutes. Government methodology rather than a commercial score. Case studies provide descriptive evidence of common pathways, but they do not validate prediction of rare violence. Oversight is principally DHS, congressional, inspector-general, and agency review.C5, C6, C9; 95/100
Chicago — Strategic Subject List and Crime and Victimization Risk ModelChicago Police Department; SSL versions from 2012, followed by CVRM. Decommissioned November 1, 2019. Ended.Used age, arrest and violent-incident history, weapon and narcotics arrests, victimization, trends, and in some versions gang affiliation to estimate whether a person would be a victim or offender in a shooting within approximately 18 months. Outputs were scores or tiers used in intelligence, “custom notifications,” and enforcement attention.Developed with Illinois Institute of Technology and $3.8 million in federal grants. OIG found poor documentation, unreliable scores, weak access controls, inadequate training, and adverse consequences. A RAND quasi-experiment found no homicide reduction and evidence of increased arrests.C2, C6, C10; 99/100
Los Angeles — Operation LASERLos Angeles Police Department; developed in 2009, expanded during the 2010s, discontinued in 2019. Ended.Combined “chronic offender” bulletins with location-based analysis. Person selection used arrest, parole or probation status, field interviews, gang information, and officer intelligence; outputs guided surveillance, stops, contacts, and focused enforcement. Human scoring and analytics were intertwined.Federal grant support; Palantir and other data systems assisted analysis, although LASER was a strategy rather than a single vendor product. LAPD OIG found inconsistent selection, retention, documentation, and oversight and could not isolate crime-reduction effects.C1, C2, C6, C10; 97/100
Los Angeles — PredPolLAPD; piloted and deployed during the 2010s, discontinued by 2020. Ended.Used recent crime type, location, and time to generate small geographic forecast boxes for short patrol periods. Output informed patrol deployment rather than naming a suspect. High automation in forecast generation, discretionary officer response.PredPol, later Geolitica. A randomized trial received an OJP “Promising” designation and reported lower property crime in treatment areas, but the LAPD OIG could not independently isolate system effects or verify broader claims. Demographic and feedback-loop concerns remained.C1, C10; 96/100
Pasco County, Florida — Intelligence-Led Policing, Prolific Offender, Focused Deterrence, and At-Risk Youth programsPasco County Sheriff’s Office; principal program began around 2011. The challenged targeting practices were discontinued under a 2024 settlement. Ended or materially constrained.Used criminal history, intelligence, associations, school and child data, calls, and officer observations to identify adults and children for repeated home visits and “prolific offender” or at-risk attention. Outputs drove visits, questioning, citations, code enforcement, and family contacts.Largely agency-designed, supported by federal grants and data platforms. Investigative reporting documented repeated visits and pressure on people not suspected of new crimes. A federal settlement required non-return to the challenged program and monetary payment. No valid prospective evaluation demonstrated crime prevention.C2, C6, C9, C10; 95/100
Plainfield, New Jersey — PredPol/Geolitica deploymentPlainfield Police Department; historical deployment analyzed through obtained prediction logs. Vendor ceased operations at the end of 2023. Ended.Generated place-and-time crime forecast boxes for patrol. An investigative comparison reported that fewer than 0.5 percent of tens of thousands of forecasts aligned with reported crimes under the matching method used. That statistic is not necessarily conventional model precision because patrol exposure, time windows, and target definitions matter, but it demonstrates why denominators must be disclosed.PredPol/Geolitica; assets, staff, and customers transferred to SoundThinking. No public independent causal evaluation of Plainfield’s deployment.C1, C10; 91/100
United Kingdom, Durham — Harm Assessment Risk ToolDurham Constabulary; introduced in 2016 for the Checkpoint deferred-prosecution program. Later operational status is less clear; treated here as historical or limited-use pending verification.A supervised machine-learning model trained on approximately 104,000 custody events used 34 variables to classify detained persons’ two-year risk of serious or non-serious reoffending. Output informed eligibility for Checkpoint. Human decision makers could review the category.Developed with academic support; initially used a random-forest approach. Legal basis arose from police diversion discretion, UK data-protection law, and public-law duties. Government reviews noted preliminary performance but the complete evaluation and error distributions were not published.C3; 91/100
United Kingdom — National Data Analytics Solution Most Serious Violence modelWest Midlands Police-led national project beginning in 2018 with multiple forces; the individual serious-violence prediction use case was paused, while national police analytics capability continued under successor structures.Proposed to predict which persons would commit a first serious gun or knife offense within roughly 24 months, using linked police datasets and generating a propensity score. It was intended for prevention and resource prioritization.Accenture participated in early development. A West Midlands presentation states that the project was paused because of predictive-analytics and ethical issues. Current National Data and Analytics Office functions are broader and should not be assumed to operate the original model.C2, C6, C10; 94/100
London — Gangs Violence Matrix, succeeded in part by Violence Harm AssessmentMetropolitan Police Service. GVM created in 2012, discontinued February 13, 2024; data was scheduled for destruction by February 13, 2025. VHA became operational in 2024 and remained active in 2026.GVM identified and risk-assessed alleged gang members and persons thought at risk of gang violence. VHA scores people involved or likely to be involved in violence, using harm methodologies and intelligence, for prioritization and possible support or enforcement. Human intelligence and scoring.UK policing powers, Data Protection Act 2018, law-enforcement data rules, equality and human-rights duties. ICO found serious data-protection breaches in the GVM; reviews documented extreme racial disproportionality. VHA publishes more procedural material, but independent prospective validation and complete outcome data remain unavailable.C2, C6, C10; 98/100
Germany — PRECOBSDeployed or piloted by police in Bavaria and other German states beginning in the 2010s. Current use varies by state and should not be presumed nationwide.Near-repeat burglary forecasting uses recent burglary location, time, and modus operandi to produce short-term geographic alerts. It does not name likely offenders. Outputs inform patrol allocation. Automated rule-based or statistical forecasts with human deployment decisions.LogObject. State police statutes and data-protection law provide legal framework. German evaluations recorded limited alert volumes and mixed results; effects were difficult to distinguish from operational changes, and evidence of displacement or cost effectiveness is weak.C1, C10; 89/100
Germany — RADAR-iTEFederal Criminal Police Office, with state police and partners; introduced nationally in 2017, updated version documented in 2019. Active or institutionally available.Structured assessment of police-known persons in the militant Islamist context. Analysts code risk and protective factors and place persons into prioritization categories for case management and resource allocation. Rule-based structured professional judgment, not autonomous machine learning.Developed by BKA with forensic-psychology expertise. Legal basis lies in federal and state police and counterterrorism powers. BKA documentation describes nationwide standardization; independent false-positive, calibration, and causal-effect data are not public. Religious belief is officially distinguished from observable behavior, but the population definition remains ideologically bounded.C4, C5, C6; 94/100
Hesse, Germany — hessenDATAHesse police; deployed from 2017. Automated data-analysis provisions were substantially constrained by the Federal Constitutional Court’s February 2023 judgment. A narrowed legal framework may permit revised use. Legally constrained/reconfigured.Integrates police databases and supports entity resolution, link analysis, pattern discovery, and searches across persons, places, communications, events, and objects. It can prioritize investigative leads but does not by itself establish future offending.Palantir Gotham-based. The Federal Constitutional Court held that broad Hesse and Hamburg provisions permitting automated data analysis lacked sufficiently specific thresholds and safeguards and violated informational self-determination. This is one of the clearest judicial limits on generalized police data fusion.C4, C6, C9; 97/100
Netherlands — Crime Anticipation SystemDutch National Police; developed in Amsterdam and expanded nationally during the 2010s. Reported active, though current local configuration is not fully public.Uses recorded crime, time and location, municipal and demographic variables, and geographic features to forecast weekly crime risk for small grid cells. Outputs support briefings and patrol allocation. Aggregate, automated forecasting with local discretion.Police-developed. Governed by Dutch police-data and administrative law. Research documents nationwide adoption but little independent causal evidence. Public precision, calibration, intervention dosage, and displacement measures are incomplete.C1, C10; 91/100
Amsterdam — Top600 and Top400/ProKid-related approachesAmsterdam municipality, police, prosecutors, care and youth agencies; Top600 began in 2011; Top400 followed. Active or successor practice continuing, with changing criteria.Multiagency prioritization of persons associated with repeated high-impact offending and young people considered at risk. Inputs include convictions, police contacts, school or care information, family context, and ProKid indicators. Outputs coordinate enforcement, supervision, services, and family interventions.Public multiagency system rather than a single vendor model. Legal basis is fragmented across police, municipal, youth, welfare, and data-sharing powers. Government evaluation found mixed effects; civil-society research documents notice, stigmatization, family surveillance, and data-sharing concerns. Complete technical criteria and false-positive rates remain unavailable.C2, C3, C6, C10; 88/100
China, Xinjiang — Integrated Joint Operations PlatformXinjiang authorities and public-security organs; prominently documented from the mid-2010s. The architecture’s exact current configuration is secret, but mass digital surveillance remains. Operational status: likely continued or evolved.Aggregates identity, movement, device, household, vehicle, checkpoint, religious, communication, and behavioral data. Generates person-level flags for investigation and can contribute to interrogation or detention. Protected and ordinary behavior—such as travel, religious practice, communication patterns, or use of technology—has reportedly triggered suspicion.Developed through government security contractors; exact division of vendor responsibility is incomplete. Chinese national-security, counterterrorism, cybersecurity, and regional rules are invoked. Reverse engineering and field evidence link the system to discriminatory mass surveillance and arbitrary detention. OHCHR concluded that serious violations in Xinjiang may constitute crimes against humanity.C2, C4, C6, C7, C9; 95/100
China — Police Cloud platformsMinistry of Public Security and provincial or municipal police authorities; large-scale projects documented since the 2010s. Active/evolving.Integrates police, government, commercial, communications, travel, health, family, and social data for identity resolution, network mapping, anomaly detection, investigation, and population control. Outputs may be leads, alerts, relationship graphs, or watchlists.Multiple Chinese technology vendors and local integrators. Broad security and policing authorities; limited independent judicial review or individual contestability. Public technical validation is unavailable. Human-rights reporting documents surveillance of activists, petitioners, ethnic minorities, and ordinary social networks.C2, C4, C6, C7, C9; 85/100
China — Sharp Eyes, Skynet, and related video-surveillance systemsCentral and local public-security and political-legal authorities; expanded nationally during the 2010s. Active.Dense CCTV networks integrate facial recognition, license-plate recognition, behavior detection, and command systems. Primary function is identification and tracking, but outputs can feed watchlists, predictive analysis, and preemptive intervention.Highly decentralized market involving major camera, AI, and integration firms. Procurement research identified tens of thousands of notices and billions of renminbi in spending. Accuracy and false-match rates are not publicly auditable at system level. Consequences include pervasive monitoring and expansion of state capacity to track targeted populations.C4, C6, C7, C9, C10; 88/100
India, Maharashtra — MARVEL and AI-Powered Investigation Platform pilotMaharashtra Research and Vigilance for Enhanced Law Enforcement, a state-backed company registered in 2024; Nagpur Rural pilot documented in 2026. Pilot.Intended to integrate CCTNS data and provide AI-assisted investigation, search, pattern analysis, and decision support. Public documents do not establish a validated future-crime model; classification here is association and investigative prioritization rather than proven prediction.Government-backed development; pilot budget approximately ₹2 crore. Scaling was to depend on presentation of outcomes. Legal basis includes state police powers, CCTNS governance, criminal-procedure rules, and Indian data law, but a public AIA and detailed model documentation were not located.C6, C9, C10; 85/100
India, Ghazipur — AI-SPS crime mapping and predictive policing systemGhazipur district police, Uttar Pradesh; reported in 2025. Reported operational or pilot; independent verification limited.Reportedly maps incidents, identifies patterns, and produces hotspot or patrol guidance. Public information is insufficient to establish data fields, model family, thresholds, vendor, legal authority, or whether individuals are scored.Evidence is principally official statements relayed through news reporting. No technical validation, rights assessment, procurement record, or outcome evaluation was located.C1, C10; 58/100
Australia, New South Wales — Suspect Target Management Plan, including DV-STMPNSW Police; STMP II documented from 2005 and DV-STMP from 2015. The general STMP was ended in 2023, with related risk-management practices continuing under other policies. Ended/replaced.Officers selected “high-risk” targets, including children, for proactive policing, visits, stops, searches, bail checks, and enforcement. Inputs included offending history, intelligence, associations, and officer judgment. Person-level selection was more important than automation.Police powers and internal policy rather than a commercial algorithm. BOCSAR evaluated reoffending outcomes, while the Law Enforcement Conduct Commission’s Operation Tepito identified serious maladministration and discriminatory or oppressive effects in cases involving children.C2, C3, C6, C10; 97/100
Canada — CBSA Scenario-Based Targeting and PAXIS/API-PNR assessmentCanada Border Services Agency National Targeting Centre. Active.API and PNR data are automatically screened through intelligence-derived scenarios. Matches enter a Scenario Work List for targeting-officer review. Decisions can affect pre-arrival examination, questioning, referral, seizure, immigration action, and intelligence development.Customs Act, Immigration and Refugee Protection Act and regulations, Passenger Information Regulations, privacy law, and international PNR commitments. PNR is generally retained up to 3.5 years, with progressive masking and limited exceptions. Scenarios may be simulated on depersonalized historical data. Public precision, recall, and demographic impact are unavailable.C4, C6, C8; 97/100
Vancouver — GeoDASH and predictive crime analysisVancouver Police Department; described in official budget and planning material. Reported active, configuration uncertain.GeoDASH maps incidents and supports crime analysis, CompStat, forecasting, and property-crime deployment. The public record does not establish whether a proprietary forecast model remains active or how much it affects patrol decisions.Public police analytics environment; vendor details and model specifications are not fully disclosed. No independent causal evaluation, precision measure, or impact assessment was located.C1, C10; 66/100
New Zealand — Youth Offending Risk Screening ToolNew Zealand Police; validation reports published in 2011. Current operational scope is uncertain.Structured instrument intended to assess youth-offending and recidivism risk, using individual history and psychosocial or situational factors. Output supports referrals, case planning, and prioritization.Police-developed with external validation research. Official reports examined reliability, predictive capability, variable contribution, and cultural validity, an unusually explicit validation agenda. Current recalibration, Māori impact, intervention outcomes, and present-day use need verification.C2, C3; 80/100
**New Zealand — RoC*RoI**Department of Corrections; integrated into practice from 1998. Active or institutionally embedded.Uses approximately 35 criminal-history and demographic variables to estimate risk of reconviction and reimprisonment over five years. Outputs inform correctional classification, program allocation, parole-related information, and case management.Government statistical model under corrections and parole law. It has a longer documentation history than many police tools, but current subgroup calibration, override practices, and the causal effect of score-informed interventions require fuller publication.C3, C10; 94/100
France — PAVEDFrench National Gendarmerie; developed around 2017, trialed in 2018, and reportedly paused in 2019 before broad deployment. Paused/ended.Intended to forecast burglaries or theft-related events geographically using police incident data. Output was patrol-oriented geographic risk.Public-sector development; complete model details and legal assessment were not located. Reporting indicates the national rollout was halted, leaving little independent evidence of performance or consequences.C1, C10; 76/100
France, Marseille — M-PulseMarseille municipality in partnership with public-safety actors; developed in the late 2010s. Current operational status is uncertain.Combined urban, incident, event, environmental, and municipal data to anticipate security demand and guide deployment. It appears principally place-based and resource-oriented rather than an individual criminality score.Engie participated. Municipal-security powers and French/EU data-protection law apply. Vendor and municipal claims exceed the available independent evidence; public precision, causal benefit, displacement, and cost-effectiveness data were not located.C1, C9, C10; 69/100
Switzerland — PRECOBS deploymentsCantonal police, including Zurich and other reported cantons, from around 2013. Current canton-by-canton status varies.Near-repeat burglary forecasts produce short-term geographic alerts for patrol. No named person is predicted.LogObject. Cantonal police law and Swiss data-protection requirements apply. Independent and journalistic reviews found no clear evidence of durable crime reduction; details of alerts and patrol dosage are incomplete.C1, C10; 84/100
Israel and occupied Palestinian territory — Red Wolf, Wolf Pack, and related facial-recognition systemsIsraeli military and security authorities at checkpoints and in Hebron and East Jerusalem; documented in 2023 and thereafter. Reported active/evolving.Checkpoint cameras and mobile tools scan Palestinians, match faces against databases, and return color-coded or similar guidance affecting passage, questioning, detention, or arrest. Identification, watchlisting, and movement control are integrated; the system need not predict a discrete offense to operate preemptively.Exact vendors and legal architecture are not fully disclosed. Amnesty’s investigation documented expansion through scans of people not previously in the database. There is little notice, meaningful consent, independent audit, or accessible contestation for affected Palestinians.C4, C6, C7; 86/100
South Korea — Crime Risk Prediction and Analysis System, Pre-CASKorean police; reported nationwide implementation from May 1, 2021. Active, subject to limited public documentation.Integrates crime, spatial, environmental, and other data to produce area risk analysis and support crime-prevention and patrol planning. Public evidence indicates place-level rather than named-person prediction.Police-developed or government-contracted system; exact vendor and legal documentation were not located in English. A 2026 study surveyed 106 police users and noted that empirical research on actual utilization remained limited.C1, C10; 72/100
Japan, Tokyo — AI monitoring of online “dark job” recruitmentTokyo Metropolitan Police; reported operational from 2025. Active, based on corroborated reporting but limited primary documentation.AI scans public social-media posts for language associated with recruitment into anonymous criminal groups, classifies posts by risk, and routes them to officers who review and send warnings or pursue inquiries. It targets content and accounts rather than forecasting a general crime rate.Vendor and model details are undisclosed. Legal issues include police intelligence authority, freedom of expression, platform terms, data protection, and the distinction between illegal solicitation and ambiguous speech. Reporting indicates warning volumes increased substantially; prevention effects and false-positive rates are unavailable.C5, C6, C9; 67/100
Brazil, São Paulo — Smart SampaSão Paulo municipal government and Guarda Civil Metropolitana; integrated monitoring center inaugurated in 2024. Active and expanding.Integrates municipal and connected private cameras with facial recognition, license-plate recognition, behavioral analytics, emergency response, and watchlist matching. Outputs can trigger police deployment, identity checks, and arrest of wanted persons.Multiple camera and analytics contractors under municipal procurement. Litigation initially raised racial-discrimination and fundamental-rights concerns. By 2026 the platform reportedly comprised tens of thousands of cameras, but system-level accuracy, watchlist quality, demographic false matches, and effectiveness are not publicly established.C4, C6, C7, C10; 83/100

Country and agency comparison matrix

Country or systemDominant operational modelGoverning postureBest-established benefit evidencePrincipal unresolved risk
United StatesFragmented watchlisting, traveler targeting, threat assessment, local place and person systemsConstitutional rights, sectoral statutes, Privacy Act and administrative law; no single national law governing police algorithmsSome place-based experiments; mature threat-assessment practice; limited traveler-screening output dataNational-security secrecy, weak notice, decentralized procurement, racial and associational feedback
United KingdomNational analytics infrastructure, local risk tools, police intelligence matricesData Protection Act, UK GDPR outside core law enforcement, Law Enforcement Directive-derived rules, Equality Act, Human Rights Act, ICO and policing oversightSome validation of HART; procedural learning from Gangs Matrix and NDAS reviewsDisproportionality, unpublished evaluations, replacement tools escaping legacy scrutiny
GermanyPlace forecasting, structured terrorism assessment, police data fusionStrong constitutional informational self-determination, federal/state police statutes, EU data lawStructured documentation of RADAR-iTE and some PRECOBS evaluationsBroad data fusion and opaque security prioritization
NetherlandsNational place forecasting plus multiagency person managementPolice Data Act, administrative and data-protection controls, municipal data-sharing arrangementsGovernment evaluation of Top600; descriptive CAS researchFamily spillovers, intervention attribution, unclear current model performance
ChinaPopulation-scale fusion, biometric tracking, anomaly detection, security watchlistingBroad security statutes with limited independent judicial constraintNo reliable causal public-safety evaluationDiscriminatory mass surveillance, arbitrary detention, protected-behavior monitoring
IndiaRapid state and district pilots integrated with CCTNS and investigation systemsPolice statutes, criminal procedure, constitutional rights, evolving digital-data lawAlmost none publicly availableProcurement opacity, unclear algorithms, insufficient impact assessment
AustraliaPerson-targeting through internal police policy more than advanced MLState police law, ombudsman and conduct commissions, anti-discrimination and privacy regimesBOCSAR and LECC reviews provide unusually concrete operational evidenceCoercive intervention against children and Indigenous communities
CanadaBorder scenario targeting and local police analyticsCharter, Privacy Act, Customs and immigration statutes, federal privacy reviewTransparent retention and masking rules relative to peersNo public targeting precision, recall, demographic impact, or cost-effectiveness
New ZealandYouth and correctional risk instrumentsPrivacy Act, corrections and police law, Treaty and equality obligationsFormal validation work for YORST and long documentation history for RoC*RoICurrent calibration, Māori impact, and downstream intervention effects
FranceExperimental place and municipal demand forecastingPolice and municipal-security law, CNIL and EU data lawLittle independent evidenceUncertain status, vendor-led claims, fragmented municipal procurement
SwitzerlandCanton-level near-repeat forecastingCantonal police statutes and federal/cantonal data protectionSome critical independent reviewWeak evidence of incremental benefit and inconsistent public reporting
Israel/OPTFacial recognition linked to checkpoints and watchlistsDomestic security and military legal regimes; international humanitarian and human-rights law strongly implicatedOperational identification claims onlyMovement restrictions, ethnic discrimination, no meaningful notice or appeal
South KoreaNationwide place-risk analysisPolice and personal-information lawUser-perception study, but little outcome evidenceUndisclosed model, variables, forecast quality, and deployment dosage
JapanSocial-media detection and targeted interventionPolice law, constitutional expression rights, privacy and platform governanceIncreased identification of suspect posts, not demonstrated crime reductionOverbroad language classification and chilling effects
BrazilIntegrated camera networks and biometric watchlist matchingConstitutional privacy and equality principles, data-protection law, municipal procurementCounts of cameras, matches, and arrests; no rigorous causal studyRacial false matches, mass tracking, private-camera integration, scale before validation

Detailed case studies

The U.S. watchlisting ecosystem

The Terrorist Screening Center illustrates why a watchlist should be evaluated as a distributed governance system rather than a database. An originating agency nominates a person; TSC reviews and consolidates the record; screening agencies apply different subsets or operational rules; and frontline officers, airline systems, consular staff, border officials, or foreign partners interpret the resulting match. Biographic and biometric identifiers are intended to reduce identity ambiguity, but name similarity, transliteration, incomplete identifiers, incorrect source intelligence, and stale records can produce false matches or erroneous nominations. The FBI’s public criteria state that nominations cannot be based solely on race, ethnicity, national origin, religion, or First Amendment-protected activity, but the operative evidence and nomination reasoning are generally unavailable to the affected person.

The effects are not uniform. A TSDS record may generate no visible action, repeated secondary screening, questioning, border inspection, visa consequences, denial of boarding, or a law-enforcement encounter. The No Fly List is therefore only the most severe layer. Secure Flight performs passenger matching; ATS-P applies broader border targeting; and other systems combine travel data with law-enforcement and intelligence information. The systems are connected but should not be treated as a single algorithm.

Litigation demonstrates several distinct failure modes. In Ibrahim v. DHS, an FBI agent checked the wrong box on a nomination form, producing years of consequences. In Latif v. Holder, a federal court found the then-existing No Fly redress process constitutionally inadequate. Kashem v. Barr examined revised procedures and upheld important portions, illustrating that procedural sufficiency is context dependent. FBI v. Fikre concerned whether removal from the list mooted a lawsuit; the Supreme Court held in 2024 that the government had not carried its burden merely by removing the plaintiff without assuring that the challenged conduct would not recur. Tanzin v. Tanvir permitted damages claims under the Religious Freedom Restoration Act against federal officials alleged to have used No Fly placement to pressure men to become informants.

The technical evidence remains inadequate. The government publishes neither the number of distinct screening transactions nor the denominators needed to calculate false-positive rates. A redress correction rate cannot substitute for accuracy because many travelers do not know why they were screened, some never file, and the agency may correct identity resolution without acknowledging an erroneous underlying nomination. Likewise, the number of prevented boardings is not a measure of terrorist-risk precision.

Assessment: technical validity D/unknown; causal public-safety benefit D/unknown; legality and procedural fairness C− because courts have forced improvements but secrecy remains; distributional and human-rights impact D because the affected population and error distribution are not disclosed. Pre-crime intensity: 16/20. Recommended disposition: independently audited and legally constrained, with prohibition on coercive consequences based solely on a watchlist match. [Case confidence: 98/100.]

Traveler-risk scoring: ATS-P, Secure Flight, and the terminated Quiet Skies program

Secure Flight answers a comparatively narrow question: whether passenger identity information matches a security-list record and what screening instruction follows. ATS-P asks a broader question: whether a traveler or itinerary matches targeting rules, risk indicators, law-enforcement information, or intelligence. Quiet Skies went further by selecting some travelers for observation based on travel patterns and behavioral indicators even when they were not on a principal terrorist watchlist. These are not equivalent functions, yet they can converge on the same traveler and create cumulative suspicion.

The principal governance problem is that travelers ordinarily see only the downstream friction: an inability to obtain a boarding pass, repeated “SSSS” screening, questioning, device searches, an air-marshal presence, or a border referral. They generally do not receive the rule, source record, confidence level, or agency responsible for the flag. Human review exists, but it can be confirmatory rather than independent when the reviewer lacks access to source reliability or is institutionally rewarded for avoiding false negatives.

Quiet Skies provides a warning about evaluation by institutional assertion. DHS terminated the program on June 5, 2025, stating that it had failed to stop a terrorist attack. That statement may be relevant to cost and governance, but it does not establish that the system produced zero security value: preventing a rare event is intrinsically difficult to measure. Conversely, an agency could not validate the system merely by pointing to investigations or screening discoveries after a flag. A valid evaluation would compare outcomes against a defined counterfactual, publish how many travelers were selected, distinguish true security discoveries from ordinary regulatory violations, and account for the high base-rate advantage enjoyed by broad screening.

Canada’s scenario-based targeting system offers a useful procedural contrast. CBSA publicly describes automatic scenario matching, officer review, progressive masking of PNR information, a 3.5-year ordinary retention ceiling, and simulation of scenarios on depersonalized historical data. Yet Canada also does not publish sufficient precision, recall, false-positive, or demographic-impact data. Transparency about process is therefore better than transparency about performance.

Assessment: ATS-P technical validity D/unknown, Secure Flight C/unknown, Quiet Skies E because its principal rules and evidence were secret. Causal benefit is unavailable for all three at the necessary level of granularity. Pre-crime intensity: ATS-P 15, Secure Flight 14, Quiet Skies 17. Recommended disposition: Secure Flight and ATS-P may continue only with match-quality audits, source-record review, strict purpose limits, publishable aggregate error statistics, and real-time redress; Quiet Skies-style covert surveillance based on broad behavioral rules should remain discontinued. [Case confidence: 96/100.]

Behavioral threat assessment at the FBI and Secret Service

Behavioral threat assessment is frequently mischaracterized as an attempt to calculate who will become violent. The FBI and Secret Service instead describe a process of evaluating concerning conduct in context: communications, planning, fixation, grievance, leakage, weapons acquisition, target research, stressors, deterioration, capability, and protective factors. The FBI emphasizes that no demographic or personality profile reliably identifies a future attacker and that threat assessment is not synonymous with predicting violence.

This distinction matters because the operational objective is often management rather than classification. A multidisciplinary team may interview a person, contact family or employers, encourage treatment, address a grievance, restrict access to a site, evaluate firearm access under applicable law, or open an investigation where there is independent evidence. The most defensible model continuously updates its assessment and seeks to reduce risk rather than label the subject permanently.

Nevertheless, threat assessment has three structural vulnerabilities. First, rare-event prediction remains implicit even where agencies disclaim prediction. A team must decide which cases deserve resources, and broad referral systems can generate many more cases than can be investigated. Second, “concerning behavior” can overlap with protected expression, disability, unpopular political belief, religious practice, adolescent conduct, or mental-health crisis. Third, intervention records may migrate into police intelligence systems, affecting later decisions long after the original concern has passed.

Evaluation is underdeveloped. Secret Service case studies identify behavioral patterns among known attackers, but retrospective prevalence among perpetrators does not establish prospective specificity. Many nonviolent people exhibit grievances, fascination with violence, or disturbing communications. The proper unit of evaluation is therefore not whether assessors can identify common attacker characteristics; it is whether a defined triage and management protocol improves outcomes compared with ordinary practice without generating disproportionate surveillance, coercion, or unnecessary criminalization.

Assessment: technical validity B− for structured case practice but D for rare-event prediction; causal benefit C/unknown; legality and fairness B− when interventions rest on independent lawful grounds and decline to use protected activity as a proxy; human-rights impact C+, highly dependent on local implementation. Intensity: FBI BTAC 9, Secret Service NTAC/BTAU 8. Recommended disposition: permitted with safeguards, including service-first options, documented behavioral nexus, separation from durable criminal-intelligence files, periodic closure, and explicit prohibitions on ideology, religion, disability, or protected speech as sufficient grounds. [Case confidence: 94/100.]

Chicago’s Strategic Subject List

Chicago’s SSL is a canonical example of a technically framed score entering an institution without a stable theory of use. Beginning in 2012, successive models estimated whether an individual would become a victim or offender in a shooting. Later versions used age, arrest and victimization history, weapon and violent incidents, activity trends, and, in one version, narcotics arrests and gang affiliation. The output changed from numerical scores to tiers in the Crime and Victimization Risk Model.

The Chicago OIG found that the department could not reliably reconstruct how scores were used, who accessed them, what training governed interventions, or whether adverse consequences were consistently linked to the intended preventive strategy. Earlier versions had not received adequate evaluation before operational use. A core design ambiguity—predicting victimization and offending in the same label—made the intervention problem especially acute. A person at high risk of being shot might need protection and services, while a person thought likely to shoot someone might receive enforcement attention. Combining both in “party to violence” scoring invited officers to treat vulnerability as dangerousness.

A RAND evaluation did not find that being placed on the list reduced the probability of homicide victimization. It found evidence consistent with increased arrests among listed people, suggesting that the intervention may have changed police behavior more readily than individual behavior. Subsequent research reported increased not-guilty outcomes among some targeted persons, though that study should be interpreted cautiously because it concerned a specific design and comparison.

The SSL also demonstrates why “algorithmic bias” is too narrow a diagnosis. Even a perfectly calibrated model could be harmful if its high-risk category prompted poorly specified enforcement. Conversely, an imperfect model might be less harmful if used solely to offer voluntary services without adverse data sharing. Chicago lacked a controlled intervention protocol, individual notice, correction process, clear expiration rule, and reliable audit trail.

Assessment: technical validity D; causal benefit D−; legality and procedural fairness D; distributional impact D, with foreseeable amplification of historically concentrated police data. Intensity: 14/20. Recommended disposition: do not revive; any successor person-level violence model should be paused until prospectively validated and tied to a lawful, separately evaluated intervention. [Case confidence: 99/100.]

Los Angeles: LASER and PredPol

Los Angeles operated two analytically distinct systems. PredPol forecast locations and times for specified crime types. LASER combined place analysis with “chronic offender” bulletins and intensive attention to named people. Treating them as one predictive-policing system obscures the greater rights intensity of LASER.

PredPol’s design was comparatively parsimonious: recent crime events generated short-term forecast boxes. A randomized Los Angeles trial reported reductions in property crime and received an OJP “Promising” rating. But the LAPD Inspector General found that the department lacked adequate records to isolate PredPol’s effect from patrol decisions and other strategies. The experimental result therefore supports the possibility of benefit under a defined implementation, not a general conclusion that PredPol or proprietary forecasting works in every deployment.

LASER’s chronic-offender component relied on police records, field interviews, criminal history, supervision status, and intelligence. The OIG found inconsistent criteria for selection and removal and substantial gaps in oversight. Such inconsistency undermines both validity and equal treatment: two similarly situated persons could be treated differently depending on division practice, while officers could interpret a bulletin as authorization for proactive contact. The department discontinued LASER in 2019.

The vendor history illustrates market consolidation. PredPol later renamed itself Geolitica, ceased operations at the end of 2023, and transferred customers, employees, and intellectual property to SoundThinking, a company that had previously acquired HunchLab. Agencies can thus continue a forecasting function after the public associates the original brand with controversy. Procurement oversight must follow functionality, data, and model lineage, not merely the corporate name.

Assessment: PredPol technical validity C+, causal benefit C, legality/fairness C, distributional impact C−; LASER technical validity D, causal benefit D, legality/fairness D, distributional impact D. Intensity: PredPol 6, LASER 15. Recommended disposition: transparent place-based methods may be piloted under randomized evaluation; person bulletins of the LASER type should remain discontinued unless rebuilt around stringent evidence and notice requirements. [Case confidence: 97/100.]

Pasco County’s intelligence-led policing program

Pasco County demonstrates how a risk list can become a standing authorization for intervention. The sheriff’s office identified adults as prolific offenders and children as at risk, then used repeated home visits, questioning, citations, code enforcement, school-related information, and contacts with relatives or landlords. Investigative reporting described residents receiving visits even when officers lacked suspicion of a new crime.

The critical issue was not whether an algorithm autonomously ordered each visit. The program institutionalized a future-oriented logic: prior history, associations, family circumstances, or school data justified repeated attention intended to disrupt anticipated offending. Human discretion amplified rather than cured the model’s risks because officers had broad options and weak external limits.

The consequences extended beyond conventional policing. Housing, code compliance, family stability, schooling, and relationships could be affected. Children were especially vulnerable to self-reinforcing records: police attention generated new observations, which could then sustain the judgment that attention remained necessary.

A 2024 settlement required the sheriff’s office not to return to the challenged program and included monetary relief. Settlement is not equivalent to a final judicial determination on every constitutional issue, but it is strong evidence that unbounded proactive targeting could not be defended as ordinary intelligence-led policing.

No credible causal evaluation demonstrated that the program reduced serious crime relative to comparable areas. Counts of visits, arrests, or code violations would not establish benefit because those outputs are partly generated by the intervention itself. A valid evaluation would have measured serious victimization, offending, displacement, family harm, school effects, and complaints against a comparison group.

Assessment: technical validity E/D; causal benefit E/unknown; legality and procedural fairness F-level concern; distributional and human-rights impact E. Intensity: 18/20. Recommended disposition: prohibited in its documented form. [Case confidence: 96/100.]

The United Kingdom: HART, NDAS, and London’s intelligence matrices

Durham’s HART model is one of the more clearly specified police risk tools. It used 34 variables and a random-forest classifier trained on custody events to assign detained persons to risk categories over a two-year period. Its output informed eligibility for Checkpoint, a diversion program. The intended intervention could be beneficial, but denying a diversion opportunity based on predicted risk is still a consequential decision.

Two design features deserve attention. First, risk thresholds reportedly emphasized avoiding false negatives in the high-risk category, which can increase false positives. The normative cost of errors depends on the consequence: falsely denying diversion may be more serious than unnecessarily offering support. Second, historical police data encode differential detection and enforcement. Even if race is excluded, postcode, prior contacts, and offense history may reproduce disparities.

The NDAS Most Serious Violence project was more ambitious. Its stated use case contemplated predicting people likely to commit a first serious gun or knife offense within about 24 months. Official project material indicates that this work was paused because of ethical and predictive-analytics concerns. That decision is an important example of responsible nondeployment: a technically feasible model may still lack an acceptable intervention, legal basis, or positive predictive value.

London’s Gangs Violence Matrix was not primarily a machine-learning model, but it exerted greater pre-crime intensity. The matrix identified alleged gang members and people at risk of violence, facilitated data sharing, and was linked to enforcement and safeguarding decisions. The ICO found serious data-protection failures, and public review showed that Black people were dramatically overrepresented. The Met discontinued it in February 2024 and introduced the Violence Harm Assessment, which uses harm scoring and intelligence to identify people involved or likely to be involved in violence.

The replacement raises a governance test: a successor should not inherit legitimacy merely because it uses academically recognized harm weights, publishes a DPIA, or omits “gang” from its name. It requires fresh validation of population selection, scoring, intervention, proportionality, racial impact, retention, and outcomes.

Assessment: HART intensity 11, NDAS violence model 13, GVM 15, VHA provisionally 12. HART may be independently tested; the NDAS person-prediction model should remain paused; GVM should remain discontinued; VHA should be time-limited and independently audited before expansion. [Case confidence: 96/100.]

Germany and the Netherlands

Germany presents three different governance models. PRECOBS forecasts burglary locations. RADAR-iTE structures professional assessment of police-known persons in the Islamist-extremism context. hessenDATA integrates databases for link and pattern analysis.

PRECOBS has low pre-crime intensity because it does not ordinarily name a person. Its evidence problem is incremental value. Near-repeat burglary is a known criminological phenomenon; police can map recent burglaries without purchasing a proprietary alert system. The necessary comparison is therefore PRECOBS versus transparent hotspot mapping and ordinary analysts, not PRECOBS versus random patrol. Evaluations have not established durable, transferable superiority, and few studies report officer compliance, patrol minutes, displacement, or total cost.

RADAR-iTE is person based but is described as a structured, rule-based instrument rather than a machine-learning prediction. It standardizes assessment and may improve consistency across German police agencies. Yet standardization can also propagate a flawed construct nationally. Because the assessed population is drawn from police-known persons in a defined ideological milieu, the fairness question begins before scoring: who enters the pool, on what evidence, and how are lawful religious or political activities separated from behavior connected to violence?

The Federal Constitutional Court’s 2023 judgment on automated police data analysis is globally significant. It did not prohibit all cross-database analytics. It held that broad provisions in Hesse and Hamburg did not sufficiently limit the data, purposes, methods, and thresholds for highly intrusive automated analysis. The ruling recognizes that combining lawfully held data can create a qualitatively new intrusion.

The Netherlands’ CAS is closer to PRECOBS but reportedly uses a wider set of spatial and demographic variables. Top600 and Top400 are closer to intensive case management: they coordinate police, prosecutors, municipal agencies, care, and youth services around named people and families. Such integration can deliver support, but it can also allow a police risk designation to influence welfare, education, housing, and family interventions without a single accountable decision maker. Government evaluation found mixed outcomes, and civil-society investigations document notice and stigmatization concerns.

Assessment: German PRECOBS 5/20, RADAR-iTE 13, hessenDATA 14; Dutch CAS 5, Top600/Top400 16. Place systems may be piloted with transparent baselines; RADAR-iTE requires independent validation and entry-pool auditing; broad data fusion must follow the German constitutional standard; multiagency lists require notice, purpose separation, and individual correction. [Case confidence: 94/100.]

China’s IJOP and surveillance ecosystem

Xinjiang’s Integrated Joint Operations Platform represents the highest-intensity model in the inventory. Human Rights Watch’s reverse engineering and interviews indicate that IJOP aggregated large volumes of data and generated flags based on behavior that may be lawful and ordinary. Those flags could lead to investigation, questioning, and detention. The system operated within a wider infrastructure of checkpoints, device inspections, biometric collection, cameras, household visits, and administrative control.

The distinction between prediction and detection is largely immaterial at this intensity. IJOP may flag an observed behavior rather than calculate a formal probability of future offending, but authorities treat the behavior as an indicator of future security risk. Broad categories and data fusion produce what is functionally a preemptive suspicion regime.

Police Cloud and Sharp Eyes expand the architecture beyond Xinjiang. Police Cloud platforms seek to connect records across institutions and map relationships; Sharp Eyes and Skynet provide visual identification and tracking infrastructure. Procurement research suggests a fragmented but vast market, making central claims that a single accuracy rate governs the system implausible.

No credible public evaluation establishes precision, recall, calibration, or causal public-safety benefit. A high volume of “findings” cannot validate the system where the state defines ordinary behavior as suspicious and can compel interviews or detention. The denominator—the total number of monitored people and flags—is unavailable.

OHCHR’s Xinjiang assessment found serious human-rights violations and concluded that the scale and discriminatory character of arbitrary detention and restrictions may constitute crimes against humanity. The core governance defects are therefore not remediable through model-card publication or improved facial-recognition accuracy. They arise from discriminatory purpose, coercive consequences, absence of independent courts, lack of contestation, and use of protected conduct.

Assessment: technical validity not meaningfully demonstrable; causal benefit E/unknown; legality and procedural fairness under international standards E; distributional and human-rights impact E. Intensity: IJOP 20, Police Cloud 18, Sharp Eyes 15. Recommended disposition: prohibited, with preservation of records for accountability and remedies. [Case confidence: 96/100.]

Australia’s Suspect Target Management Plan

New South Wales’s STMP shows that predictive policing can be produced by an internal police policy rather than software. Police identified people considered at high risk of offending and subjected them to proactive attention. Children and young people were included, and domestic-violence variants used related targeting principles.

The Law Enforcement Conduct Commission’s Operation Tepito examined the treatment of children over a five-year period and found serious problems including maladministration. The significance of the investigation lies in its focus on intervention, not merely score construction. A formally reasonable risk judgment does not justify stops, visits, searches, or bail checks that lack independent legal grounds or become oppressive through repetition.

STMP also demonstrates asymmetric measurement. Police can readily count target contacts, searches, arrests, and detected breaches. They cannot readily observe crimes that did not occur, harms caused by repeated police contact, school disengagement, family stress, or displacement. An evaluation reporting reduced recorded offending among targets may be confounded by age, regression to the mean, incapacitation, intervention selection, or changes in police recording.

Indigenous overrepresentation and the treatment of children require special analysis. A tool relying on prior police contact enters a feedback system in which communities already subject to intensive enforcement generate more records and more candidates for targeting. Human discretion cannot solve this if officers use the same historical data and institutional assumptions.

NSW Police ended the named STMP in 2023, but successor practices must be evaluated functionally. A renamed high-risk offender policy that preserves selection, proactive contact, and data-sharing logic should inherit the same safeguards and audit obligations.

Assessment: technical validity D, causal benefit C−/uncertain, legality and procedural fairness D, distributional impact D/E for children and heavily policed communities. Intensity: 17/20. Recommended disposition: the documented model should remain ended; any successor must prohibit target status as an independent legal basis for stops, searches, visits, or sanctions. [Case confidence: 98/100.]

Comparative evidence, intensity scorecard, and taxonomy of effects

Evidence-quality grading system

Each program receives four grades:

GradeTechnical validityCausal public-safety benefitLegality and procedural fairnessDistributional and human-rights impact
AProspectively and independently validated for the exact population, target, threshold, and use; calibration and error distributions publishedCredible randomized or strong quasi-experimental evidence of net benefit, including displacement and dosageClear authority, necessity, proportionality, notice, contestation, audit, and enforceable safeguardsIndependently tested subgroup effects, accessible remedies, no material unjustified disparity
BGood validation but limited external replication or incomplete subgroup evidenceCredible outcome evidence with some attribution limitsGenerally lawful and reviewable, with remediable procedural gapsImpact assessed with manageable residual risks
CPartial or internal validation; incomplete calibration or threshold evidenceSuggestive association or limited quasi-experimentAuthority exists but notice, review, documentation, or proportionality is incompleteMaterial risks identified but not comprehensively measured
DWeak, outdated, vendor-led, or nontransferable evidenceNo credible causal demonstration or adverse findingsSerious legal or procedural deficienciesSignificant disparity, stigmatization, or rights concerns
ESecret, unverifiable, conceptually invalid, or grossly mismatched to decisionNo usable evidence or evidence of net harmFundamentally incompatible with due process or rights normsSevere discriminatory, coercive, or population-scale harm
UUnavailableUnavailableUndeterminedUndetermined

No program in this inventory receives an overall A. Some components—for example, identity resolution under controlled laboratory conditions—could have high technical performance, while the operational system remains unevaluated.

Pre-crime intensity scale

Ten factors are scored from zero to two, for a maximum of 20:

FactorZeroOneTwo
Named-person targetingAggregate place/resource onlySmall group or accountNamed individual
Explicit future predictionDescriptiveImplicit risk anticipationExpress forecast of future conduct
Protected or noncriminal behaviorExcludedIndirect or occasionalMaterial input or trigger
OpacityPublic and reproduciblePartial disclosureSecret rules or evidence
Consequence severityPlanning onlyRepeated attention or service eligibilityMovement restriction, detention, search, arrest, custody
Human discretionIndependent and empoweredGuided reviewRubber-stamp or automated consequence
Notice and contestabilityTimely notice and appealLimited/retrospectiveNone
Data breadthNarrow incident dataMultiple police datasetsCross-domain population data
False-positive exposureLow and measurableUnknown/moderateRare target plus broad population
DurationEphemeralMonthsYears or indefinite

Interpretation: 0–4 analytical support; 5–8 anticipatory allocation; 9–12 consequential risk prioritization; 13–16 targeted preemption; 17–20 coercive pre-crime regime.

Program scorecard

These are commission judgments, not agency ratings. “Disposition” is the recommended posture under the governance framework developed below.

ProgramIntensityTechnicalCausal benefitLegality/fairnessDistributional impactRecommended disposition
TSC/TSDS16DUC−DIndependently test and constrain
Secure Flight14CUCC−Permit with safeguards and audit
Quiet Skies/Silent Partner17EE/UDDProhibit recurrence; retain termination
ATS-P15D/UUC−U/DIndependently test and constrain
FBI BTAC9B−C/UB−CPermit with safeguards
Secret Service NTAC/BTAU8B−C/UBCPermit with safeguards
Chicago SSL/CVRM14DD−DDKeep discontinued
LAPD LASER15DDDDKeep discontinued
LAPD PredPol6C+CCC−Pilot only under independent test
Pasco County ILP18E/DE/UEEProhibit
Plainfield Geolitica5DUCUDo not redeploy without validation
Durham HART11C+C/UCCIndependently test
NDAS serious-violence model13D/UUD/UUKeep paused
London Gangs Matrix15DD/UDE/DKeep discontinued
London Violence Harm Assessment12C/UUCU/CTime-limited pilot and audit
German PRECOBS5CC−/UB−C/UPilot with transparent baseline
RADAR-iTE13CUCC/DIndependently test and constrain
hessenDATA14C/UUD before judgmentU/DPause unless narrow statutory test met
Dutch CAS5CU/C−B−/CUPilot with public evaluation
Amsterdam Top600/Top40016C−C−C−DPause high-impact components
Xinjiang IJOP20EE/UEEProhibit
China Police Cloud18E/UUEEProhibit high-impact uses
Sharp Eyes/Skynet15D/UUD/EE/DProhibit indiscriminate biometric tracking
Maharashtra MARVEL7UUU/CUControlled pilot only
Ghazipur AI-SPS8UUUUPause pending disclosure
NSW STMP17DC−D/ED/EKeep ended
CBSA scenario targeting14C/UUC+UPermit with safeguards and audit
Vancouver GeoDASH5UUC/UUVerify status; no expansion without test
NZ YORST12CUCU/CRevalidate before use
NZ RoC*RoI13B−/CU/CCC/UPermit advisory use with recalibration
France PAVED4U/C−UC/UUNo rollout without new test
Marseille M-Pulse5UUC/UUPilot only
Swiss PRECOBS5C−U/C−B−UPilot only with public metrics
Red Wolf/Wolf Pack18UUE/DEProhibit current discriminatory use
South Korean Pre-CAS6U/C−UU/CUIndependently test
Tokyo social-media AI11UUC/UU/CNarrow pilot with speech safeguards
São Paulo Smart Sampa14U/CUC−/DD/UPause biometric expansion pending audit

Metrics that must replace anecdotal “success”

Precision is the share of alerts that correspond to the specified outcome. Recall is the share of actual outcomes that the system successfully flags. A high-recall system can have extremely low precision when the event is rare. If serious violence occurs among one percent of a screened population, a model with 80 percent sensitivity and 90 percent specificity would flag roughly 10.7 percent of the population but only about 7.5 percent of flags would be true positives. The remaining approximately 92.5 percent would be false positives.

Calibration asks whether people assigned, for example, a 20 percent risk actually experience the defined outcome about 20 percent of the time, including within relevant demographic and geographic groups. Calibration alone does not establish usefulness or fairness. A model can be calibrated but too inaccurate for coercive action.

False-positive burden must be measured in people, encounters, and time—not merely percentages. A one-percent error rate across 100 million screenings can produce one million erroneous flags. Burden also includes repeated screening of the same person, downstream database propagation, lost travel, investigative visits, employment or housing consequences, and the effort needed to seek correction.

Base rates must be reported for the exact population and outcome. Programs often improve apparent performance by selecting a population already enriched for police contact. That does not show that screening the general public is effective.

Geographic coverage requires the denominator of all grid cells and time periods, not merely the percentage of crimes that happened somewhere inside a broad forecast area. Forecast-box size, duration, overlap, and total area flagged must be disclosed.

Intervention dosage measures what police actually did after an alert: patrol minutes, stops, visits, surveillance hours, service offers, searches, citations, referrals, or restrictions. A model cannot be evaluated without determining whether officers followed it and whether the intervention differed from ordinary practice.

Displacement includes movement of crime across adjacent locations, times, offense types, victims, or enforcement channels. Diffusion of benefits should also be measured.

Cost effectiveness must include licenses, integration, data cleaning, training, analyst time, patrol diversion, false-positive investigation, litigation, redress, security, and eventual system replacement—not merely vendor fees.

Publicly unavailable metrics dominate this inventory. Precision, recall, calibration, demographic false-positive rates, override rates, intervention dosage, and cost effectiveness are unavailable for the TSC/TSDS, ATS-P, Quiet Skies, RADAR-iTE, IJOP, Police Cloud, Red Wolf, Smart Sampa, Pre-CAS, MARVEL, Ghazipur AI-SPS, Vancouver GeoDASH, M-Pulse, and most Top600/Top400 decision points. They are only partially available for HART, YORST, RoC*RoI, PredPol, PRECOBS, and CAS.

Taxonomy of potential benefits and harms

DomainPlausible benefitNecessary proofCharacteristic harm
Resource allocationMore patrol or services at high-need places and timesComparison with transparent hotspot and analyst baselinesOverpolicing of areas that generate more recorded data
Threat triageFaster review of urgent casesImproved time-to-intervention without excessive false positivesProtected speech or disability reframed as dangerousness
Victim protectionIdentifying people facing elevated victimizationVoluntary, beneficial services and reduced harmVictims treated as suspects or exposed to unwanted police contact
DiversionDirecting eligible persons away from prosecutionHigher completion and lower net system involvementHigh-risk classification used to deny diversion
Border securityPrioritizing limited inspection capacityIncreased detection of serious target harms per inspectionTravel delay, denial, detention, and discriminatory profiling
Identity matchingFinding a wanted or missing personOperational false-match rates and human verificationWrong-person arrest and mass tracking
Network analysisRevealing genuinely relevant associationsEvidentiary validation of link relevanceGuilt by association and surveillance of family, community, or religion
Multiagency coordinationCombining enforcement and supportClear purpose separation and demonstrable service benefitPolice risk labels influencing housing, education, welfare, or immigration
Crime forecastingDirecting patrol to short-term hotspotsProspective benefit beyond simple mappingFeedback loops, displacement, and self-validating arrest data
StandardizationReducing arbitrary variation among officersBetter reliability and outcomes than professional judgmentScaling the same conceptual bias across an entire jurisdiction

The most common documented harms are: erroneous identity matching; false inclusion; stigmatization; denial of opportunity; intensified stops or searches; repeated visits; movement restriction; chilling of speech, religion, association, or protest; family and community spillover; racial, ethnic, religious, Indigenous, age, disability, and class disparities; feedback loops; mission creep; stale-risk persistence; data leakage; vendor lock-in; and inability to correct the source data.

The most credible benefits are narrower: structured documentation, consistent triage, faster identity resolution where match quality is high, more deliberate allocation of scarce resources, and multidisciplinary management of genuinely concerning behavior. These benefits do not require secret person-level criminality prediction.

Comparative law, oversight, and the problem of functional rebranding

United States

No comprehensive U.S. statute governs predictive law enforcement. Applicable constraints are distributed across constitutional doctrine, administrative law, privacy statutes, sector-specific authorities, state laws, consent decrees, collective-bargaining rules, procurement law, and local ordinances.

The First Amendment limits the use of religion, political belief, association, journalism, protest, and speech as grounds for investigation or adverse action, although protected activity may sometimes be considered as contextual evidence when closely connected to an independently lawful inquiry. The critical distinction is between evidence of planning or a true threat and mere ideology, grievance, or association.

The Fourth Amendment generally requires individualized justification for stops, searches, seizures, and some forms of surveillance. A risk score or watchlist label does not itself create reasonable suspicion. Officers must be able to articulate current, particularized facts, and courts should not permit agencies to bootstrap suspicion from an undisclosed model trained on earlier police activity.

The Fifth and Fourteenth Amendments protect due process and equal protection. Severity, duration, error probability, secrecy, and the feasibility of notice influence the required procedures. No Fly List litigation shows that national-security labeling can create a constitutionally significant liberty burden.

The Administrative Procedure Act can provide review of final federal agency action, but doctrines involving standing, finality, state secrets, classified information, and national security can limit access. The Privacy Act offers correction and access rights but contains significant exemptions for law-enforcement and national-security systems. State constitutional privacy provisions and local surveillance ordinances may be more protective.

The core U.S. governance defect is fragmentation without a mandatory public record. A local police department may purchase analytics under an ordinary software contract; a federal agency may characterize a model as intelligence methodology; and a fusion center may share outputs across jurisdictions. No single institution is responsible for measuring cumulative impact.

European Union

The EU AI Act’s prohibition rules have applied since February 2, 2025. Article 5 prohibits certain individual criminal-risk assessment or prediction when based solely on profiling or personality traits, while preserving AI that supports a human assessment already grounded in objective and verifiable facts directly linked to criminal activity. The wording is important: it does not prohibit every police risk system, and “not solely” must not become a loophole in which a trivial additional fact legitimizes an otherwise profiling-based prediction.

Law-enforcement, biometric, migration, border, and criminal-justice systems can fall within high-risk categories. As of August 2, 2026, however, the recently enacted AI Omnibus had extended the application date for many Annex III high-risk requirements to December 2, 2027, while AI Act governance, prohibitions, and specified transparency rules were already in force. The Commission stated that the Omnibus entered into force on July 27, 2026. Agencies therefore cannot treat the delayed high-risk compliance date as permission to ignore existing data-protection, equality, constitutional, or human-rights law.

The Law Enforcement Directive, implemented through national law, separately limits processing by competent authorities. Article 11 restricts decisions based solely on automated processing that produce adverse legal effects or similarly significant effects unless authorized by law with appropriate safeguards, including at least human intervention. “Human intervention” must be substantive: the reviewer needs the data, competence, time, authority, and institutional independence to reject the output.

Data-protection principles—lawfulness, purpose limitation, data minimization, accuracy, storage limitation, security, accountability, necessity, and proportionality—apply even where a system does not qualify as AI. Special-category and biometric data require heightened justification. Equality law prohibits direct and indirect discrimination, and national constitutional law may impose stronger limits.

Council of Europe and international human rights

The Council of Europe Framework Convention on Artificial Intelligence and Human Rights, Democracy and the Rule of Law opened for signature in September 2024. It establishes lifecycle duties concerning risk and impact assessment, accountability, transparency, oversight, procedural safeguards, and remedies, while allowing states latitude in implementation. The EU ratified the Convention in May 2026; the Treaty Office’s chart should be consulted for each country’s signature and ratification status.

The Convention should be read as a floor, not a validation mechanism. A completed impact assessment does not make an unlawful system lawful. National-security exclusions, flexible implementation, and delayed domestic legislation may weaken immediate practical effect.

The ICCPR protects liberty and security, privacy, movement, expression, association, peaceful assembly, equality, and effective remedies. Restrictions must be lawful, necessary, proportionate, and non-discriminatory. Watchlisting, border scoring, mass facial recognition, and behavioral monitoring can implicate several rights simultaneously.

A system used in armed conflict or occupation may also engage international humanitarian law. Biometric checkpoint control of a protected population cannot be assessed solely as a data-protection issue; freedom of movement, discrimination, arbitrary detention, collective surveillance, and occupation law are central.

National approaches

Germany provides the strongest judicial statement in this inventory. The Federal Constitutional Court’s 2023 decision treated automated cross-database analysis as capable of producing new knowledge and deeper intrusion than the component records. Statutes must specify the permissible data, methods, purposes, triggering thresholds, and protection of persons not connected to wrongdoing.

The United Kingdom combines public-law rationality, the Human Rights Act, Equality Act, data-protection law, policing statutes, ICO enforcement, inspectorates, and local mayoral oversight. The Gangs Matrix episode shows both the value and lateness of this oversight: serious defects persisted for years before discontinuation. The VHA’s published operating material is an improvement, but independent evidence remains necessary.

The Netherlands relies on police-data legislation, municipal administrative law, and multiagency agreements. Distributed responsibility is a recurring problem: an individual may experience a combined police, welfare, youth, housing, and prosecution intervention without a single appealable risk decision.

Canada offers relatively detailed PNR retention and masking rules. Its Charter, Privacy Act, customs and immigration statutes, and review institutions can constrain targeting, but effective challenge remains difficult where scenario rules and source intelligence are secret.

Australia and New Zealand rely heavily on police and correctional statutes, privacy regimes, anti-discrimination law, conduct commissions, ombuds institutions, courts, and internal policy. NSW’s LECC shows the importance of oversight with access to case files rather than only model documentation.

China has enacted extensive cybersecurity, data, counterterrorism, and personal-information legislation, but formal legality does not provide meaningful protection where the state defines broad political, religious, or ethnic conduct as security risk and independent review is absent.

How agencies evade or dilute oversight

The most common techniques are functional rather than necessarily deceptive:

TechniqueGovernance effectRequired response
Calling a forecast “intelligence”Invokes secrecy and avoids algorithm rulesRegulate consequential inference regardless of label
Calling a score “decision support”Suggests a human cure without measuring relianceAudit override rates, reviewer time, and actual decisional influence
Replacing “prediction” with “risk” or “harm”Avoids public association with predictive policingApply functional tests: future orientation, target, and consequence
Rebranding a discontinued systemResets public attention while preserving inputs and interventionsRequire lineage disclosure and successor-system review
Embedding analytics in a larger platformMakes the model appear to be ordinary records managementAssess each analytic module and cross-database capability
Classifying rules as operationally sensitivePrevents independent replicationPermit secure auditor access and publish bounded summaries
Procuring a service rather than softwareAvoids asset inventories and capital reviewCover licenses, hosted services, data exchanges, and consulting
Running a “pilot” indefinitelyAvoids permanent-program approvalsImpose end dates, sample limits, and automatic deletion
Separating the score from the interventionObscures causal responsibilityEvaluate the complete sociotechnical decision chain
Framing biometric systems as identification onlyOmits the watchlist and downstream consequenceRegulate enrollment, list quality, match thresholds, and action rules

A governing statute should therefore define covered systems by function and effect: any computational or structured analytic process used to infer, rank, identify, or prioritize a person, group, place, event, or transaction for the purpose of anticipating, preventing, investigating, or managing crime, violence, security threats, border risks, or public disorder.

Procurement market and model governance instruments

Vendor and procurement analysis

The market is not limited to specialized “predictive policing” firms. It includes enterprise data-fusion companies, cloud providers, camera manufacturers, facial-recognition developers, GIS firms, records-management vendors, consulting companies, systems integrators, and government-created corporations.

Geolitica/PredPol and SoundThinking show product-line consolidation. A controversial brand can disappear while its personnel, customers, intellectual property, or functionality moves into a broader public-safety platform. SoundThinking’s earlier acquisition of HunchLab reinforces the need for model-lineage and corporate-successor disclosure.

Palantir’s hessenDATA role shows how an analytics platform can create predictive or preemptive capacity without selling a product called predictive policing. Search, entity resolution, link analysis, and cross-database pattern discovery can materially influence whom police investigate. The German constitutional decision appropriately focused on capability and legal thresholds rather than marketing terminology.

Accenture’s NDAS participation demonstrates the influence of consulting and integration contractors in problem definition, data linkage, and prototype development. Intellectual ownership of the final model is less important than who chooses targets, variables, thresholds, and success measures.

LogObject’s PRECOBS illustrates a narrower proprietary product whose claimed value must be tested against low-cost, transparent alternatives. The procurement question is not simply “Does it predict better than chance?” but “Does it outperform standard near-repeat analysis after all costs and operational differences are included?”

Engie’s M-Pulse role shows the convergence of urban-management and policing analytics. Data gathered for mobility, events, lighting, or municipal services can migrate into security forecasting.

China’s Sharp Eyes ecosystem and São Paulo’s Smart Sampa demonstrate decentralized camera markets. No single vendor controls the end-to-end system; public agencies combine cameras, networks, watchlists, command centers, facial recognition, and private feeds. Responsibility can become diffused across hardware suppliers, integrators, software developers, list owners, and frontline users.

Procurement contracts should require:

  1. government ownership or perpetual access to input, output, version, threshold, and audit logs;
  2. disclosure of model lineage, subcontractors, pretraining data, third-party components, and corporate transfers;
  3. no trade-secret restriction on judicial, regulator, defense, or accredited-auditor access;
  4. reproducible export of the model or decision logic where technically possible;
  5. fixed termination, data-return, deletion, and interoperability provisions;
  6. security testing and incident notification;
  7. limits on secondary vendor use of police or public data;
  8. predetermined performance and rights thresholds tied to payment;
  9. indemnification that does not displace public accountability; and
  10. a prohibition on vendor control of public communications about accuracy.

Model algorithmic impact assessment

An AIA should be completed before procurement, repeated before deployment, and updated after material changes. It should be signed by the operational chief, data-protection or privacy officer, civil-rights officer, technical lead, legal counsel, procurement officer, and independent reviewer.

AIA fieldRequired content
Purpose and necessityPrecisely defined harm, affected population, decision, and evidence that a computational system is necessary
Less intrusive alternativesComparison with staffing, services, hotspot mapping, professional analysis, warrant-based investigation, and nonpolice interventions
Legal authorityStatutory provision, constitutional analysis, data authority, retention authority, information-sharing authority
Prediction targetObservable outcome, time horizon, unit, exclusions, and why it is legally and operationally relevant
InputsEvery variable, source, collection authority, quality assessment, missingness, proxy risk, and update frequency
PopulationEligibility and exclusion rules, geographic coverage, children, protected groups, non-suspects, visitors, and bystanders
ModelArchitecture, training and validation periods, feature transformations, thresholds, version, and uncertainty
Output and interfaceScore, alert, rank, map, confidence interval, explanation, and what the user sees
InterventionEvery permitted and prohibited response, evidentiary threshold, escalation path, and service option
Human reviewQualifications, time, information, independence, override authority, and documentation
MetricsPrecision, recall, calibration, false-positive burden, coverage, subgroup error, dosage, displacement, and cost
Rights impactPrivacy, equality, expression, association, movement, liberty, family, child, disability, and Indigenous rights
SecurityAccess controls, logging, adversarial risk, data poisoning, identity theft, model extraction, and vendor access
Notice and redressNotice timing, explanation, correction, appeal, emergency exception, and remedy
RetentionSource, feature, output, alert, audit, and case-record retention; deletion triggers
GovernanceNamed owner, independent auditor, oversight body, public reporting, complaint mechanism, and sunset
Exit planTermination, data deletion, vendor transition, record preservation for litigation, and successor review

A red-rated AIA factor—lack of legal authority, use of protected conduct as a sufficient predictor, inability to define the outcome, no feasible notice for a severe consequence, no independent validation access, or a coercive automated action—should stop deployment rather than be “balanced” against claimed benefits.

Predeployment field-testing protocol

Laboratory reconstruction. Auditors reproduce data extraction, feature generation, model training, threshold selection, and outputs. Records are checked for duplicates, identity merges, missingness, coding drift, leakage from future information, and label contamination.

Retrospective validation. The model is frozen and tested on a later, untouched period. Results must include all requested metrics, confidence intervals, subgroup distributions, temporal drift, and comparisons against simple baselines.

Shadow mode. For a defined period, the system produces outputs that cannot affect people or deployment. Analysts document what action they would have taken. Shadow mode reveals alert volume, operational feasibility, data latency, and disagreement without creating intervention harm.

Prospective limited field test. Where ethically permissible, randomized or stepped-wedge designs compare the new system with standard practice. Place-based tests randomize comparable locations and measure patrol dosage. Person-based tests should ordinarily evaluate voluntary supportive interventions, not heightened surveillance or coercion.

Intervention fidelity. Every downstream action is logged. The evaluation distinguishes no action, service referral, officer contact, stop, search, surveillance, border examination, arrest, denial, and database dissemination.

Spillover and displacement. Adjacent areas, unflagged groups, family members, neighboring time windows, and alternative offenses are monitored.

Independent stopping rules. The trial must stop for serious unlawful action, subgroup disparity above the predetermined limit, data breach, performance below baseline, excessive false-positive burden, or unanticipated high-impact use.

Posttrial decision. A public report must precede continuation. Failure to demonstrate net benefit results in deletion and termination, not an indefinite pilot.

Independent-audit specification

An accredited auditor must have secure access to source data, sampling frames, model artifacts, rules, documentation, logs, contracts, overrides, case files, and complaint records. The agency and vendor may redact public details that would enable evasion, but may not withhold them from the auditor.

The audit must examine:

Audit domainMinimum test
Data provenanceLawful source, purpose compatibility, completeness, error correction, and representativeness
Identity resolutionFalse matches, merges, splits, transliteration, aliases, and biometric threshold behavior
Construct validityWhether the target corresponds to the claimed harm and decision
Model validityOut-of-time and out-of-place performance, calibration, uncertainty, and baseline comparison
Subgroup performanceRace, ethnicity, national origin, religion where lawful to test, sex, age, disability, geography, socioeconomic proxies, Indigenous status, and intersectional groups
OperationsUser comprehension, alert fatigue, reliance, override rates, workarounds, intervention consistency
LegalityAuthority for each input, inference, disclosure, consequence, and retention period
RightsNecessity, proportionality, chilling effects, family spillovers, and access to remedy
SecurityAccess, exfiltration, insider use, vendor use, poisoning, and incident response
OutcomesCrime or harm effects, services delivered, displacement, false-positive burden, and cost
GovernanceAIA compliance, change control, procurement compliance, complaints, notices, and sunset

Audit reports should publish methods, aggregate results, material limitations, agency responses, remediation deadlines, and the auditor’s conclusion. Classified annexes should be available to courts and designated oversight bodies. Agencies should not select or pay auditors through arrangements that condition compensation on a favorable outcome; a regulator, inspector general, or pooled independent fund is preferable.

Minimum public-disclosure standard

Before deployment, the agency should publish:

  1. official program and successor names;
  2. responsible owner and all participating agencies;
  3. vendor, subcontractors, and contract value;
  4. purpose, legal authority, and affected decisions;
  5. whether named persons, groups, accounts, or places are targeted;
  6. prediction target and time horizon;
  7. categories of inputs and expressly excluded data;
  8. automation level and human-review process;
  9. permitted and prohibited interventions;
  10. retention and sharing rules;
  11. validation design and full aggregate metrics;
  12. demographic and geographic impact results;
  13. independent-audit schedule;
  14. notice, correction, complaint, and appeal procedures;
  15. security-incident history;
  16. model and policy change log;
  17. current operational status;
  18. number of persons, places, or transactions assessed;
  19. annual counts of alerts and downstream actions; and
  20. sunset date and criteria for continuation.

Operationally sensitive details may be withheld only after a written, reviewable finding showing a specific evasion risk. Withholding exact weights does not justify withholding the target, data categories, population, intervention, aggregate performance, demographic effects, retention, or redress.

Individual notice, correction, and appeal model

Notice should be event based. A person must receive notice when an algorithmic or structured risk output materially contributes to denial of travel, enhanced recurring screening, placement in a person-management program, targeted police visits, denial of diversion, custody classification, movement restriction, or dissemination to a nonpolice agency.

The notice should identify the responsible agency, program, decision, date, broad input categories, source agencies, nature of the output, human reviewer, retention period, and method of challenge. It need not disclose operational details that would compromise a live investigation, but delayed notice should follow when the risk expires.

Correction must reach both source and derivative records. Fixing a misspelled name in one screening system is insufficient if the erroneous association, score, or alert remains in partner databases.

Appeals should proceed through three levels: prompt internal review by a person not involved in the original decision; independent administrative review with access to classified or sensitive evidence through secure procedures; and judicial review. Severe consequences require expedited relief. The government should bear the burden of establishing continued inclusion after a prima facie showing of error or staleness.

Where notice is temporarily withheld, an independent body must review the withholding and impose an expiration date. Aggregate secrecy cannot become permanent individualized nonaccountability.

Rules for retention, protected activity, human review, and emergencies

Retention. Unacted-on place forecasts should ordinarily be deleted within 90 days after evaluation. Person-level alerts that do not lead to a substantiated case should be deleted within 30 to 180 days depending on severity. Watchlist and threat records require mandatory periodic re-justification, not automatic renewal. Child and school records should receive shorter periods. Audit logs may be retained longer in segregated form to support accountability, but must not be repurposed for operational intelligence.

Protected activity. Religion, political belief, protest, association, immigration advocacy, journalism, legal representation, disability, mental-health status, and constitutionally protected speech may not constitute sufficient grounds for a risk designation. Consideration is permissible only where a specific act is directly relevant to a lawful inquiry, documented, necessary, and interpreted in context. Proxies and inferential reconstruction are covered by the same rule.

Human review. Reviewers must receive training, source reliability, uncertainty, alternative explanations, and relevant exculpatory information. They must record agreement or override and explain any high-impact decision. An officer’s ability to click “approve” is not meaningful review. No arrest, search, detention, travel denial, custody escalation, or adverse benefits decision should rest solely on a score, watchlist match, facial-recognition result, or algorithmic alert.

Emergency use. Emergency deployment is limited to an imminent and specific threat of death or serious bodily harm; must be authorized by a designated senior official; may last no more than seven days without judicial or independent renewal; may use only necessary data; must log every query and consequence; and must receive retrospective legal, technical, and rights review. Emergency data may not automatically populate ordinary watchlists or intelligence files.

Decision framework, research gaps, and policy recommendations

Decision framework

DispositionGoverning testSystems in this inventory
ProhibitedDiscriminatory or protected-behavior-based person prediction; coercive action solely from profiling or automated output; indiscriminate biometric tracking tied to watchlists; no feasible contestation; purpose intrinsically incompatible with rightsIJOP; documented Pasco model; discriminatory Red Wolf deployment; recreation of Quiet Skies-style covert behavioral surveillance; Police Cloud uses aimed at political or ethnic population control; automated criminality prediction based solely on profiling
PausedMaterial rights consequence with no independent validation, unclear legal basis, missing population or intervention definition, or uncontrolled data fusionNDAS serious-violence model; biometric expansion of Smart Sampa; high-impact components of Top600/Top400; Ghazipur AI-SPS; broad hessenDATA-type analytics absent narrow legislation; unvalidated VHA expansion
Piloted under strict limitsPlausible low-intensity benefit, reversible consequences, narrow data, measurable target, and ethical prospective testPRECOBS, CAS, Pre-CAS, M-Pulse, MARVEL’s noncoercive investigative functions, transparent place-based forecasting, Tokyo social-media triage with speech safeguards
Independently tested before continued useExisting consequential program with incomplete public evidence but potentially legitimate purposeTSDS/TSC procedures, ATS-P, RADAR-iTE, HART, YORST, RoC*RoI, CBSA scenario targeting, successor person-harm tools
Permitted with safeguardsDefined lawful purpose, no solely automated coercion, strong human review, notice and correction, independent audit, demonstrated net benefitNarrow identity matching with human verification; structured behavioral threat management; nonpersonal hotspot analysis; advisory recidivism tools with current calibration; border targeting with published aggregate performance and appeal

Systems should automatically move to a more restrictive category when they add named-person outputs, protected-behavior inputs, cross-domain data, biometric identification, severe consequences, longer retention, or wider sharing.

Research gaps and paired policy responses

Research gapPolicy recommendation
G1. No global registry identifies operational, terminated, renamed, or successor systems.P1. Create mandatory national and international public registries covering algorithms, structured risk tools, watchlists, biometric systems, vendors, pilots, and successors.
G2. Watchlisting denominators, match errors, nomination reversals, and repeat-screening burdens are secret.P2. Require annual watchlisting statistics, secure independent audits, subgroup analysis, and publication of bounded error and redress outcomes.
G3. Agencies rarely publish intervention dosage after a forecast or score.P3. Log every material downstream action and evaluate the model and intervention as a single system.
G4. Place-based studies often compare a product with ordinary patrol rather than transparent analytical baselines.P4. Require comparison against hotspot mapping, analyst judgment, and problem-oriented policing before procurement renewal.
G5. Person models often combine offending and victimization.P5. Prohibit combined labels; design separate protection and enforcement pathways with distinct legal thresholds.
G6. Calibration and error distributions by subgroup are rarely available.P6. Mandate out-of-time, out-of-place, and subgroup calibration reporting with confidence intervals and minimum sample rules.
G7. The people who receive no intervention after a flag are omitted from evaluation.P7. Preserve segregated evaluation records so all alerts, including unacted-on alerts, enter precision and burden calculations.
G8. False positives are counted as records rather than lived consequences.P8. Measure delays, visits, searches, surveillance hours, missed travel, family effects, and correction costs.
G9. Causal evidence for threat-assessment teams is thin.P9. Fund ethical multisite evaluations comparing structured, service-oriented threat management with ordinary referral practice.
G10. Replacement systems evade legacy oversight.P10. Require functional lineage assessments whenever data, personnel, vendors, target populations, or interventions carry into a successor.
G11. Trade-secret claims block defense, judicial, and auditor access.P11. Make unrestricted regulator, court, defense-expert, and accredited-auditor access a nonwaivable procurement condition.
G12. Cross-agency data sharing obscures who made the consequential decision.P12. Assign a legally accountable decision owner and provide a single portal for correction across all recipients.
G13. Children’s risk tools lack longitudinal evidence about education, family, and developmental harm.P13. Presumptively prohibit police person-risk scoring of children; permit only independently approved, voluntary service tools with short retention.
G14. Indigenous, racial, ethnic, religious, disability, and intersectional impacts are incompletely measured.P14. Require community-governed impact studies and empower equality or human-rights bodies to suspend systems.
G15. Biometric accuracy is tested in laboratories rather than operational watchlist conditions.P15. Test camera quality, crowd conditions, demographic performance, list quality, operator behavior, and wrong-person consequences in the deployment environment.
G16. Cost studies omit integration, analyst time, litigation, redress, and decommissioning.P16. Use total lifecycle cost and cost per independently verified beneficial outcome.
G17. National-security systems lack counterfactual evaluation.P17. Establish cleared independent evaluation units able to conduct controlled tests and publish nonclassified findings.
G18. Data drift and policy drift are rarely monitored after launch.P18. Mandate continuous monitoring and automatic suspension after material performance, population, purpose, or intervention changes.
G19. Emergency powers can seed permanent databases.P19. Impose seven-day default limits, independent renewal, post-use notice, segregated storage, and deletion absent a substantiated case.
G20. There is little evidence about community-level legitimacy, reporting behavior, and displacement from services.P20. Include trust, willingness to report crime, service access, complaints, and community survey measures in every public-safety evaluation.

Commission recommendations

The paired responses above produce a coherent twenty-point policy program. The commission should give priority to the following implementation sequence.

Establish a legal presumption against named-person prediction. A public authority seeking to predict that a named person will offend, become violent, or pose a security threat should bear a heightened burden of necessity and validity. Systems based solely or predominantly on profiling, personality, protected conduct, family association, neighborhood, or historical police contact should be prohibited where they can contribute to coercive action.

Separate prediction from legal authority. A forecast never supplies the legal grounds for a stop, search, detention, arrest, border denial, home visit, code action, or custody escalation. Each action requires its own current statutory and constitutional justification.

Regulate watchlists as adjudicative infrastructure. Nomination, matching, dissemination, consequence, retention, and redress must be governed together. A watchlist should not be treated as an intelligence file beyond procedural review when it predictably affects travel or liberty.

Prefer services over enforcement for vulnerability predictions. Where a model identifies potential victimization, the default response should be voluntary, confidential support. Declining services must not increase a person’s risk category or be treated as evidence of dangerousness.

Require independent validation before procurement, not after controversy. Agencies should not purchase a system on the basis of vendor demonstrations, retrospective fit, or experience in another jurisdiction. Validation must concern the local population, target, data, threshold, interface, and intervention.

Create enforceable sunset clauses. Every covered program should expire after no more than three years, and every pilot after no more than one year, unless renewed through public findings of legality, necessity, effectiveness, proportionality, and cost effectiveness.

Build defense and judicial access into the system. When an output contributes to prosecution, detention, sentencing, parole, or another adjudicative decision, the affected person must be able to challenge the model, data, match, source reliability, and human use. Trade secrets cannot override fair-trial rights.

Ban solely automated high-impact law-enforcement decisions. The prohibition should include arrest referrals, search authorizations, travel denials, custody changes, diversion denial, inclusion in intensive person-management programs, and biometric stop decisions.

Treat human review as an auditable control. Agencies must disclose review time, override rates, disagreement patterns, reviewer training, and whether users can see uncertainty and exculpatory data.

Fund independent public alternatives. Where place-based forecasting may be useful, governments should develop transparent, reproducible hotspot and workload tools rather than become dependent on opaque proprietary systems.

Final determination

The global record does not support either of two extreme claims: that all predictive analysis is inherently illegitimate, or that better data automatically produces objective policing.

Low-intensity geographic analysis can sometimes improve planning, especially when it uses narrow event data, transparent methods, short retention, and noncoercive outputs. Structured behavioral threat assessment can improve documentation and coordination when it avoids demographic profiling, distinguishes protected expression from conduct, and emphasizes voluntary risk reduction. Correctional tools may support consistent planning when independently calibrated and never treated as determinative.

The case for person-level preventive enforcement is much weaker. Chicago, Los Angeles, Pasco County, London’s Gangs Matrix, and NSW’s STMP show that risk designation can transform vulnerability, association, or historical contact into recurring police attention without proving public-safety benefit. National-security and traveler systems magnify this problem through secrecy and severe consequences. IJOP, Police Cloud applications, and discriminatory biometric checkpoint systems demonstrate that pre-crime infrastructure can become a mechanism of population control rather than crime prevention.

The decisive governance question is therefore not “Does the system use AI?” It is:

What claim about a person, place, or future event is being made; on what data and evidence; who acts on it; what consequence follows; how often is it wrong; who bears the error; and can the affected person obtain an effective remedy?

Any agency unable to answer those questions should not deploy the system. Any program whose answer depends on protected behavior, secret evidence, indiscriminate population data, unreviewable automation, or severe action against people not reasonably suspected of wrongdoing should be prohibited. Systems with plausible but unproven benefits should be paused or confined to genuine, time-limited trials. Only systems demonstrating lawful purpose, incremental public-safety value, controlled false-positive burden, equitable effects, meaningful human judgment, transparent oversight, and individual contestability should be permitted to continue.