Security / Resilience / Autonomous Systems

H-R02: Autonomous Incident Detection, Classification, and Machine-Speed Triage Architecture

Report summary

The convergence of Information Technology (IT) and Operational Technology (OT) has transformed the built environment into a highly contested cyber-physical domain, rendering legacy compliance-centric security models fundamentally inadequate [institutional analysis]1. Standard security frameworks tra

Status
Research archive item
Category
Security / Resilience / Autonomous Systems
Length
5,916 words
Reading time
27 minutes
Report type
evaluation

Key topics

  • Security / Resilience / Autonomous Systems
  • Security
  • Resilience
  • Autonomous Systems
  • AI
  • Agentic Web
  • .NET
  • Runtime
  • Privacy

Research provenance

Archive status
Research archive item
Content identity
sha256:c50f106d7684077e0abfc4a614ce27f65f424b9e6ac2c2bc40a24653a996f35d

For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.

This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.

Full report

On this page

1. Executive Decision Brief

The convergence of Information Technology (IT) and Operational Technology (OT) has transformed the built environment into a highly contested cyber-physical domain, rendering legacy compliance-centric security models fundamentally inadequate \[institutional analysis\]1. Standard security frameworks traditionally interpreted the legal standard of care as administrative adherence to accepted best practices, heavily relying on static checklists and validation-tuned heuristics \[observed deployment or practice\]1. However, these methodologies fail to provide finite-sample control over operational error rates in safety-critical environments, frequently resulting in automated defense mechanisms inducing self-inflicted physical outages \[peer-reviewed research finding\]1. To achieve machine-speed triage without introducing catastrophic fragility, organizational defenses must transition to systems security engineering (SSE) frameworks rooted in verifiable mathematical risk control and engineered navigation \[policy proposal\]1. This report defines the comprehensive architecture required for the autonomous detection, correlation, classification, and triage of cyber-physical incidents. By strictly integrating Mondrian (class-conditional) conformal prediction for threshold calibration, Dempster-Shafer theory for sensor fusion, W3C PROV semantic graphs for causal correlation, and OASIS OpenC2 for standardized actuation, systems can execute machine-speed responses while rigorously bounding both false positive and false negative rates \[technical proposal\]3. High-confidence anomalies trigger deterministic containment protocols via Collaborative Automated Course of Action Operations (CACAO) playbooks, whereas low-confidence observations seamlessly degrade via exponential decay functions to prevent stale data from poisoning active state models \[reasoned inference\]6. Ultimately, this architectural paradigm shifts the standard of care from passive notification to deterministic, physics-aware resilience, ensuring that critical infrastructure maintains a safe state even during advanced persistent compromises \[institutional analysis\]1.

2. Definitions and Scope

To achieve autonomous triage at machine speed, telemetry must be deterministically classified into distinct operational states. The architecture strictly delineates these states to prevent automated logic from applying malicious containment protocols to benign hardware failures \[technical proposal\]. An anomaly is defined as a statistically significant deviation from established conformal prediction bounds or physical invariant models that currently lacks verified attribution or causal linking \[reasoned inference\]. Conversely, a fault represents a determinable hardware or software malfunction originating from internal entropy, physical wear, or logic errors, presenting without malicious provenance or external exploitation vectors \[institutional analysis\]. An attack is characterized as a sequence of actions demonstrating malicious intent, validated through causal provenance graphs and matching known threat behaviors or unauthorized state mutations \[peer-reviewed research finding\]8. Furthermore, the system must recognize maintenance, which is a pre-authorized, scheduled mutation of system state or configuration, verified by valid cryptographic tokens and corresponding change-management telemetry \[observed deployment or practice\]. Benign drift constitutes a gradual, non-malicious divergence in telemetry, such as environmental temperature shifts affecting analog sensor resistance, that maintains core system invariants and safely degrades in priority over time \[technical proposal\]. Finally, a sensor failure denotes the complete loss, erratic fluctuation, or physical blinding of a telemetry source, characterized by a sudden drop in signal-to-noise ratio or a complete loss of the data feed, independent of the underlying physical process changes \[reasoned inference\].

3. Historical and Technical Context

Historically, cybersecurity incident response relied on linear, human-driven lifecycles that fundamentally lacked the capacity to operate at machine speed. For example, the National Institute of Standards and Technology (NIST) Special Publication 800-61 Revision 2 provided a stage-by-stage playbook—Detection, Analysis, Containment, Eradication, Recovery, and Post-Incident Activity—designed specifically for human responders \[stale or superseded\]9. However, as the velocity of automated threats increased, these manual procedural models resulted in severe analyst burnout, unmanageable alert queues, and missed intrusions \[observed deployment or practice\]2. The reliance on validation-tuned heuristic thresholds meant that deployed systems carried no finite-sample control over their operational error rates, degrading unpredictably under distribution shifts and causing high false positive rates \[peer-reviewed research finding\]2. In April 2025, NIST SP 800-61 Revision 3 superseded Revision 2, fundamentally restructuring incident response to align with the NIST Cybersecurity Framework (CSF) 2.0 \[established standard or law\]10. This revision shifted the focus from reactive, stage-by-stage containment to continuous monitoring (DE.CM) and broader organizational risk management \[current official policy\]10. Concurrently, the publication of NIST SP 800-160 Volumes 1 and 2 introduced systems security engineering (SSE) as a critical subdiscipline, arguing that cyber-physical systems require physics-based hazard mitigation, hazard-specific traceability, and non-digital fallbacks \[current official policy\]1. This paradigm shift from viewing security as an administrative compliance exercise to treating resilience as a fundamental, engineered physical property demands an entirely new architecture capable of withstanding degraded states without triggering catastrophic, self-inflicted failures \[institutional analysis\]1.

4. Current Standards, Law, Policy, and Deployed Practice

The operationalization of autonomous incident handling relies on the strict integration of several authoritative standards and protocols designed to replace proprietary, non-interoperable frameworks \[observed deployment or practice\]. NIST SP 800-160 Volume 2 defines the cyber resiliency engineering lifecycle, legally and practically requiring systems to anticipate, withstand, recover from, and adapt to adverse conditions \[current official policy\]12. This standard mandates that systems maintain essential functions despite an adversary establishing a foothold, directly rejecting the premise that compliance controls focused solely on boundary defense equate to actual resilience \[established standard or law\]12. To facilitate interoperable, machine-speed orchestration, the architecture relies heavily on standards defined by the Organization for the Advancement of Structured Information Standards (OASIS). The Collaborative Automated Course of Action Operations (CACAO) version 2.0 standard specifies machine-readable JSON playbooks that orchestrate sequential, parallel, and conditional security logic blocks \[established standard or law\]6. These playbooks actuate defenses via Open Command and Control (OpenC2), a standardized machine-to-machine language that decouples high-level security intents from specific proprietary implementation technologies using precise actuator profiles \[established standard or law\]4. Furthermore, for forensic and causal correlation, the W3C PROV standard is deployed to model data provenance as a directed acyclic graph, mapping subjects, objects, and events to enable robust semantic causality tracking \[established standard or law\]8.

5. Architecture and Data Models

Telemetry and Data Correlation

To accurately distinguish between anomalies, faults, attacks, and benign drift, an autonomous system requires exhaustive, multi-dimensional telemetry spanning identity tokens, network flow metrics, host-level execution traces, physical-process sensor data, and hardware state \[technical proposal\]4. Correlation cannot rely on simple, linear timestamp matching, as adversaries routinely manipulate system clocks or exploit temporal race conditions \[reasoned inference\]. Instead, the architecture utilizes the W3C PROV data model to construct systemic provenance graphs \[peer-reviewed research finding\]8. By defining processes and threads as active subjects, files and network sockets as passive objects, and system calls as the connecting events, the system establishes a mathematically verifiable causal dependency \[peer-reviewed research finding\]8. This semantic-aware linkage natively resists adversarial timing manipulation because the topology of the directed acyclic graph remains invariant regardless of the localized timestamp \[reasoned inference\]8. Furthermore, to handle discrepancies between varying sensor inputs—such as a network intrusion detection system reporting an attack while a programmable logic controller (PLC) reports normal physical function—the architecture employs Dempster-Shafer theory \[peer-reviewed research finding\]17. This mathematical framework fuses evidence by calculating belief and plausibility bounds across contradictory sensors, preserving the uncertainty without forcing a premature or brittle probabilistic collapse \[peer-reviewed research finding\]5.

Architectural Minimization of False Positives and Negatives

The deployment of singular detection mechanisms is insufficient for safety-critical environments. The architecture requires a multi-model consensus approach, leveraging the strengths of diverse analytical paradigms while mathematically bounding their respective failure modes \[technical proposal\].

Analytical MethodArchitectural Function and FocusFalse Positive ProfileFalse Negative ProfileMachine-Speed Viability
Rule EnginesDeterministic signature matching and strict boundary enforcement.Low (if signatures are strict).High (fails to detect novel/zero-day attacks).High (O(1) lookup latency).
Anomaly DetectionStatistical deviation identification from baseline behavioral telemetry.High (frequently flags benign operational drift).Low (capable of catching novel adversarial methods).High (optimized for stream processing).
Supervised LearningClassification via labeled historical datasets.Moderate (highly susceptible to distribution shifts).Moderate (vulnerable to adversarial evasion).High (fast inference post-training).
Self-Supervised LearningDeep representation learning on massive unlabeled data sets.Moderate.Low (robust to topological and distribution shifts).Moderate (computationally heavy).
Causal ModelsPhysics and logic-based invariant checking (e.g., thermodynamic laws).Very Low (physics laws are immutable).Low (assuming the physical environment is fully mapped).High (highly deterministic).
Bayesian MethodsProbabilistic updating of beliefs based on sequential evidence.Moderate (highly dependent on accurate prior probabilities).Moderate.Moderate (requires continuous parameter tuning).
Graph MethodsW3C PROV semantic linkages and dependency mapping.Low (highly context-aware and relationally driven).Low (difficult for adversaries to spoof full graphs).Low (graph traversal introduces high latency).
Multi-Agent CouncilsAgent consensus (e.g., LLMs wrapped in CACAO orchestration).Low (cross-validation filters errors).Low.Low (unacceptable latency for machine-speed).

Table 1: Comparison of Analytical Methods for Autonomous Classification \[reasoned inference\]. To synthesize these methods effectively, the system integrates Digital Twins and Invariant Checking \[technical proposal\]19. Cyber-physical systems operate under immutable kinematic and thermodynamic laws. By running a mathematical digital twin in parallel to the physical plant, the system continually evaluates telemetry against these invariants \[peer-reviewed research finding\]19. If IT telemetry indicates normal operation, but the OT telemetry demonstrates a physical state that violates the thermodynamic invariant model, the system registers a high-confidence anomaly overriding the IT-based rule engines \[reasoned inference\].

6. Failure Modes and Adversarial Cases

The deployment of autonomous triage systems introduces novel attack surfaces and architectural fragility points. If the system is not engineered to withstand active subversion, adversaries will weaponize the automation to induce self-inflicted outages \[institutional analysis\].

25 Implementation Failure Modes

\#Failure Mode NameDescription and Mechanism of Failure
1Compliance TheaterReplacing engineered physical navigation with static, administrative notification checklists.
2Dependency ExplosionGraph correlation algorithms failing to prune benign OS background noise, exhausting memory.
3Threshold DegenerationRelying on validation-tuned heuristic thresholds instead of mathematically proven conformal bounds.
4Minority Class CollapseMarginal conformal prediction severely under-covering rare attacks, dropping below 1% coverage.
5Missing Contextual FallbackThe absence of analog or mechanical fallbacks when primary digital sensors completely fail.
6Stale Evidence PersistenceFailing to decay historical indicators, causing cascading, non-relevant false positives.
7Semantic GapDisconnect between high-level OpenC2 JSON commands and low-level proprietary device APIs.
8Actuator SaturationFlooding firewalls with thousands of OpenC2 'deny' commands, exhausting the control plane.
9Playbook DriftCACAO playbooks diverging from actual network topologies over time, rendering actions useless.
10Orphaned AlertsAlerts falling exactly into the dual-threshold deferral band during periods when no human is available.
11Identity Spoofing TrustBlindly trusting cryptographic tokens without validating the physical provenance of the request.
12Clock DesynchronizationNTP poisoning destroying the causal ordering required for W3C PROV graph construction.
13Sensor BlindingFailing to detect a sudden drop in the signal-to-noise ratio, interpreting silence as safety.
14Cascading FailoverAutomated service migration triggering rapid resource exhaustion on secondary failover nodes.
15Dempster-Shafer ConflictThe mathematical inability of the algorithm to resolve highly conflicting evidence masses.
16Infinite Triage LoopsAutomated CACAO playbooks triggering each other in an unbreakable cyclic dependency.
17Unbounded AutomationAutomating physical safety shutdowns without requiring a mandatory human operator override token.
18Poisoned BaselinesTraining anomaly detection algorithms during an active, completely undetected network compromise.
19Alert Fatigue ReintroductionOver-notifying SOC analysts of 'deferred' uncertified alerts, recreating the original problem.
20Metadata LeakagePublic evidence models inadvertently exposing internal network topologies to external observers.
21Format ObsolescenceRelying on deprecated incident formats like RFC 5070 instead of modern OpenC2/CACAO.
22Lack of Graceful ExtensibilityThe system failing completely under stress rather than stretching and degrading cleanly.
23Improper Class ConditioningFailing to use Mondrian subsets during calibration, mathematically masking minority failure rates.
24Over-Reliance on LLMsUtilizing generative AI for deterministic routing without rigorous, API-level wrapper limits.
25Insufficient Calibration DataFabricating conformal bounds when the calibration pool size fails the [Figure omitted from source export] equation.

Table 2: 25 Implementation Failure Modes for Autonomous Systems \[institutional analysis\]1.

50-Case Adversarial Test Set

To assure operational resilience, the architecture must maintain finite-sample error control when subjected to the following 50 adversarial manipulation scenarios \[technical proposal\].

\#CategoryAdversarial Scenario / Inject Description
1Poisoned TelemetryInjecting slow-drift thermal data to stay below standard anomaly detection thresholds.
2Poisoned TelemetryFalsifying network flow metadata via a deeply compromised hypervisor layer.
3Poisoned TelemetrySending zero-variance sensor readings to perfectly simulate a normal, idle machine state.
4Poisoned TelemetryInjecting synthetic W3C PROV events to force the creation of false causal branches.
5Poisoned TelemetryAltering PLC register readbacks via Man-in-the-Middle while modifying physical states.
6Poisoned TelemetryFlooding the central SIEM with valid but completely irrelevant diagnostic application logs.
7Poisoned TelemetryReplaying historic, cryptographically signed normal traffic to mask a parallel data exfiltration.
8Timing ManipulationDesynchronizing network NTP by precisely 5 seconds to break PROV causality graph generation.
9Timing ManipulationDelaying critical physical sensor packets to trigger false timeout and failover faults.
10Timing ManipulationReordering TCP segments specifically to confuse and bypass deep packet inspection engines.
11Timing ManipulationSending attack events at the exact mathematical boundary of the rolling detection time window.
12Timing ManipulationExploiting execution race conditions between the detection phase and OpenC2 actuation phase.
13Timing ManipulationImplementing gradual temporal drift over months to evade rapid-spike anomaly detection.
14Timing ManipulationStretching attack chains over 6 months to force memory models to decay critical evidence.
15Identity SpoofingForging OAuth tokens with valid scopes but physically impossible geolocations.
16Identity SpoofingHijacking a legitimate machine identity (mTLS) immediately post-authentication.
17Identity SpoofingDuplicating an RFID physical badge while the user is actively logged in at a different facility.
18Identity SpoofingUtilizing a revoked identity token milliseconds before the central CRL synchronizes.
19Identity SpoofingAchieving privilege escalation inside a highly trusted OpenC2 actuator Docker container.
20Identity SpoofingSpoofing a CACAO playbook digital signature to inject malicious mitigation logic.
21Identity SpoofingExecuting automated attacks utilizing a heavily monitored, compromised domain controller account.
22Correlated Sensor ErrorSimultaneously spoofing voltage and temperature readings to perfectly match digital twin physics models.
23Correlated Sensor ErrorPhysically blinding optical sensors with lasers while injecting fake, normal telemetry.
24Correlated Sensor ErrorCompromising the shared serial data bus serving multiple redundant physical sensors.
25Correlated Sensor ErrorExploiting a zero-day shared firmware flaw across diverse, supposedly independent sensor brands.
26Correlated Sensor ErrorInducing severe electromagnetic interference (EMI) to degrade analog signal integrity.
27Correlated Sensor ErrorCovertly manipulating the physical environment (e.g., bypassing HVAC to heat the server room).
28Correlated Sensor ErrorCutting primary optical communication lines to force a fallback to unencrypted out-of-band radio.
29Model EvasionCrafting adversarial network payloads meticulously designed to sit just below Mondrian confidence thresholds.
30Model EvasionExploiting minority-class under-coverage vulnerabilities in poorly calibrated marginal conformal predictors.
31Model EvasionUtilizing Living-off-the-Land (LotL) native binaries to perfectly mimic standard administrative behavior.
32Model EvasionAltering malware execution flow to break graph-based behavioral signatures and PROV mappings.
33Model EvasionCreating synthetic anomalies purely to force an automated, highly disruptive, and unsafe failover.
34Model EvasionModifying executable file extensions to successfully bypass superficial rule engine classification.
35Model EvasionTriggering physical actions that explicitly and dangerously contradict the digital twin's predicted state.
36Alert FloodsGenerating 100,000 low-priority alerts in 60 seconds to completely exhaust the SOC processing queue.
37Alert FloodsTriggering widespread, benign threshold crossings as a smokescreen to mask a targeted data extraction.
38Alert FloodsForcing the dual-threshold conformal system to degrade to 100% manual deferral by starving the calibration pool.
39Alert FloodsExhausting the CACAO orchestration engine memory limit using deeply nested infinite while-loops.
40Alert FloodsSaturating the Dempster-Shafer evidence fusion engine with massive volumes of conflicting data points.
41Alert FloodsTriggering multiple legitimate, resource-intensive playbooks simultaneously to cause CPU starvation.
42Alert FloodsCreating thousands of orphan nodes in the W3C PROV graph to break backward causal tracking.
43Missing DataIntentionally dropping 30% of network telemetry via intermediate routing blackholes.
44Missing DataCovertly disabling host-level audit logging (e.g., ETW/auditd) mid-execution of an exploit chain.
45Missing DataSelectively erasing specific causal link events in the PROV graph immediately before transmission.
46Missing DataSimulating a rapid ransomware attack that exclusively encrypts the local log buffer.
47Missing DataDeliberately starving the calibration dataset to successfully prevent conformal threshold calculation.
48Missing DataWithholding physical sensor data entirely while IT systems fraudulently report normal operations.
49Missing DataBlocking DNS resolution specifically for external threat intelligence enrichment feeds.
50Missing DataSevering the connection to the OpenC2 actuator precisely during the execution of mitigation logic.

Table 3: 50-Case Adversarial Test Set for Autonomous Resilience Validation \[technical proposal\]21.

7. Evidence and Currentness Requirements

Thresholds and Confidence Models Without Inventing Precision

To avoid the catastrophic danger of validation-tuned heuristics, the architecture strictly employs Mondrian (class-conditional) Conformal Prediction \[peer-reviewed research finding\]3. Standard, marginal conformal prediction fails fundamentally in cybersecurity due to extreme class imbalances (e.g., a 1:345 threat-to-benign ratio). This imbalance allows a model set to 90% global coverage to hit its target by covering 96% of the majority class while effectively covering less than 1% of the critical minority threat class, leaving rare attacks completely unflagged \[peer-reviewed research finding\]3. Mondrian Conformal Prediction resolves this by calibrating quantiles independently for each specific class, mathematically guaranteeing finite-sample coverage within every category. The architecture leverages a Dual-Threshold Conformal Deferral model \[peer-reviewed research finding\]2. This model operates independently of alert prevalence by utilizing two operator-chosen error budgets: [Figure omitted from source export] (maximum acceptable false positive rate for benign escalation) and [Figure omitted from source export] (maximum acceptable false negative rate for threat misses) \[peer-reviewed research finding\]2.

1. Auto-Escalate Zone: Alerts with a nonconformity score greater than [Figure omitted from source export] are escalated automatically.

2. Auto-Close Zone: Alerts with a score less than [Figure omitted from source export] are closed automatically.

3. Deferral Band: Alerts scoring between [Figure omitted from source export] and [Figure omitted from source export] lack mathematical certainty and are deferred to a human analyst.

Crucially, this system refuses to invent precision. A threshold physically exists only if its calibration pool meets the mathematical requirement. For the upper threshold, the benign calibration pool ([Figure omitted from source export]) must satisfy [Figure omitted from source export] \[peer-reviewed research finding\]2. For a strict budget of [Figure omitted from source export], the system requires at least 99 clean calibration alerts. If the data is insufficient, the system declares the zone infeasible, gracefully closing the automated pathway and deferring to manual triage, transforming uncertified risks into measurable workload rather than silent failures \[peer-reviewed research finding\]2.

Evidence Freshness and State Degradation

Cybersecurity state is not static; telemetry degrades in relevance over time. To handle evidence freshness, the architecture applies the concept of a Loss with Decaying Factor (LDF) to its belief models \[peer-reviewed research finding\]7. The system utilizes an exponential decay function to continuously degrade confidence in stale or contradictory observations: [Figure omitted from source export] where [Figure omitted from source export] is the initial confidence state, [Figure omitted from source export] is the decay constant (the hazard rate), and [Figure omitted from source export] represents the elapsed time \[peer-reviewed research finding\]24. If a telemetry source ceases to report, the system's trust in that subsystem's security state decays exponentially. Once the confidence score crosses a predefined lower bound, the system gracefully defaults to a fail-safe, un-trusted state, preventing old, out-of-date telemetry from poisoning the active classification models \[reasoned inference\]7.

Evidence Required for Autonomous Actions

Automated actions are strictly categorized based on their reversibility, operational friction, and safety-critical impact. The required evidence threshold scales proportionately with the severity of the action \[technical proposal\].

  • Safe to Automate Immediately (Low Friction): Actions such as querying endpoints for running processes, blocking inbound perimeter ports, generating SOC tickets, or executing honeypot routing.
  • Evidence Needed: A single nonconformity score crossing the [Figure omitted from source export] threshold in the IT telemetry layer, governed by budget [Figure omitted from source export].
  • Requires Additional Evidence (Moderate Friction): Actions such as isolating internal endpoints, revoking user identity tokens, or terminating active database sessions.
  • Evidence Needed: Multi-modal consensus (e.g., both Network and Host PROV graph causal links) AND a Mondrian class-conditional confidence exceeding 95%.
  • Outside Automatic Path (High Friction / Safety Critical): Actions such as the service migration of critical OT workloads, shutting down industrial PLCs, or triggering the physical failover of electrical power systems.
  • Evidence Needed: Mandatory human-in-the-loop authorization via a CACAO playbook assignment step. Automation is only permitted if explicitly hardcoded into the system's mechanical, non-digital physical fallbacks.

8. Operational and Institutional Implications

The adoption of class-conditional conformal prediction and automated OpenC2 actuation fundamentally alters the economics and operation of the Security Operations Center (SOC) \[institutional analysis\]. By mathematically budgeting acceptable false positive ([Figure omitted from source export]) and false negative ([Figure omitted from source export]) rates prior to deployment, the organization can align cyber risk directly with financial and operational tolerances, transforming uncertified, unpredictable models into stable engineering constructs \[reasoned inference\]2. However, this transition demands highly structured assurance cases and significant capital investment. Organizations must heavily incentivize the cost of care, subsidizing the lifecycle costs of resilient architectures—such as hardware-enforced unidirectional gateways and analog, out-of-band recovery networks—that traditional IT compliance checklists entirely ignore \[policy proposal\]1. Ultimately, a higher standard of care requires making resilient, "Left-of-Boom" architectures a viable business decision rather than an unfunded regulatory expectation \[institutional analysis\]1.

9. Public-Versus-Protected Information Boundary

Organizations require a mechanism to definitively prove operational resilience and regulatory compliance to the public without exposing sensitive internal network topologies or live monitoring feeds \[technical proposal\]. The public operational-evidence model achieves this by relying on zero-knowledge proofs and aggregated cryptographic telemetry digests. Rather than exposing live SIEM dashboards or verbose incident reports, the system automatically publishes daily, mathematically verifiable assertions on a public, tamper-evident ledger \[peer-reviewed research finding\]27. These public assertions include:

1. The total volume of events processed and triaged.

2. The guaranteed error bounds ([Figure omitted from source export], [Figure omitted from source export]) successfully enforced by the conformal predictors.

3. Hash-linked cryptographic proofs of CACAO playbook executions, structurally verifying that standardized response actions were taken \[reasoned inference\]6. This model confirms the presence of an active, mathematically rigorous defense posture without ever implying or requiring live public monitoring when no verified, authorized source is connected.

10. Implementation Roadmap

Complete Incident Lifecycle Design

The architecture defines a complete, closed-loop incident lifecycle, orchestrated entirely via CACAO playbooks \[technical proposal\]:

1. First Observation: Multi-modal telemetry is ingested, correlated via W3C PROV, and scored for nonconformity.

2. Validation: The Dual-Threshold Conformal Deferral model strictly routes the observation (Auto-close, Auto-escalate, or Defer to human).

3. Correlation & Fusion: Dempster-Shafer theory calculates belief bounds across conflicting IT and OT sensors to prevent logic collapse.

4. Classification: The digital twin invariant engine checks physical constraints, mapping the confirmed anomaly to MITRE ATT\&CK taxonomy.

5. Mitigation: An OpenC2 actuator executes the deterministic containment command (e.g., deny network flow).

6. Closure: The CACAO playbook finalizes the evidence logs, applying cryptographic hashes, and outputs the zero-knowledge proof.

7. Reopening: Automatically triggered if new causal dependencies are discovered that mathematically link to historical PROV nodes.

8. Correction: Re-evaluating false positives, utilizing human analyst feedback to return the data to the conformal calibration pool.

9. Supersession: The localized incident is merged into a larger, strategic APT campaign tracker, superseding the original alert.

25 Implementation Patterns

\#Pattern NameDescription
1Dual-Threshold WrapperWrapping existing ML scoring models with conformal prediction boundary layers.
2W3C PROV Native LoggingModifying internal software to output logs natively in the structured PROV-DM format.
3OpenC2 Actuator ProxiesDeploying translation middleware for legacy firewalls to seamlessly accept OpenC2 JSON.
4Decaying Trust StoreImplementing continuous exponential [Figure omitted from source export] decay on all active identity session tokens.
5Physics-Informed Digital TwinRunning kinematic and thermodynamic models in parallel to raw OT sensor streams.
6Out-of-Band FallbackHardwiring a secure, analog serial connection specifically for emergency actuator commands.
7CACAO Modular PlaybooksBuilding small, highly reusable playbook snippets called dynamically by a master workflow.
8Dempster-Shafer Fusion NodeA dedicated microservice responsible for resolving conflicts between distinct telemetry types.
9Graceful Extensibility LoopsDesigning systems to intentionally shed non-essential computational load under cyber duress.
10Ledger-Backed AttestationStoring response actions and W3C PROV graphs on a cryptographic ledger for auditability.
11Calibration Pool RotationContinuously sliding the window of benign data to gracefully adapt to environmental baseline drift.
12Semantic Graph PruningAutomatically deleting orphaned nodes in the PROV graph after 7 days to conserve RAM.
13Hardware-Enforced SegmentationUsing physical unidirectional data diodes to protect the core OpenC2 orchestrator from compromise.
14Class-Conditional QuantilesSplitting calibration data strictly by label to mathematically protect minority threat classes.
15Token-Based Rate LimitingPreventing alert floods by strictly tokenizing the SOC triage processing queues.
16Fallback to ManualDefaulting to 100% manual deferral if calibration pools drop below mathematical limits.
17Automated Hypothesis TestingUsing Bonferroni corrections to test multiple threat hypotheses simultaneously without compounding errors.
18Zero-Knowledge Evidence ExportsGenerating public resilience reports without exposing intellectual property or schema designs.
19Mondrian Calibration StratificationStratifying telemetry contexts (e.g., VIP users vs standard) for highly accurate calibration.
20Actuator Status CallbacksRequiring cryptographically signed receipts from OpenC2 actuators immediately post-action.
21Threat-Agnostic RecoveryBuilding automated rebuilds exclusively from immutable infrastructure images rather than cleaning malware.
22Causal Forward TrackingAutomatically tracing the potential blast radius from an entry point using graph edges.
23Causal Backward TrackingTracing a detected anomaly back to its precise root origin via PROV edge reversal.
24Hybrid Human-Machine TeamingOrchestrating specific CACAO steps to pause and wait for a human supervisor token.
25Dynamic Risk BudgetingAdjusting error budgets [Figure omitted from source export] and [Figure omitted from source export] dynamically based on DEFCON-style operational states.

Table 4: 25 Implementation Patterns for Autonomous Triage Architecture \[technical proposal\].

11. Test and Assurance Plan

To assure system integrity and compliance with the modernized standard of care, organizations must undergo rigorous mathematical and operational assurance testing \[institutional analysis\]1.

30 Direct-Answer Items for Assurance Audits

\#QuestionDirect Answer
1Is the False Positive Rate mathematically bounded?Yes, strictly via Mondrian conformal prediction targeting the operator budget [Figure omitted from source export].
2How is data provenance and causality stored?As a directed acyclic graph following the W3C PROV-DM semantic standard.
3What stops playbook orchestration infinite loops?CACAO workflows contain hard, immutable limits on while execution cycles.
4How are conflicting IT and OT sensors handled?Dempster-Shafer theory calculates distinct belief and plausibility intervals.
5Does the system invent thresholds without data?No, if [Figure omitted from source export], the zone is explicitly closed and deferred to humans.
6How is physical safety ensured during automation?Digital twin invariant checking acts as a definitive veto against unsafe actions.
7Are the orchestration playbooks proprietary?No, they utilize the open, interoperable OASIS CACAO v2.0 JSON standard.
8How are physical and logical actions executed?Via OASIS OpenC2 standardized, machine-to-machine control commands.
9What happens when evidence ages or goes silent?It degrades mathematically via an exponential decay function (LDF).
10How is minority class coverage handled?Class-conditional calibration (Mondrian CP) guarantees coverage for rare threats.
11Is live monitoring data exposed publicly?No, aggregated zero-knowledge cryptographic proofs are utilized.
12What strictly defines a benign operational drift?A gradual telemetry change that perfectly maintains system physical invariants.
13How is identity token spoofing countered?By correlating logical identity tokens with physical process PROV graphs.
14Can the system operate entirely offline?Yes, local OpenC2 actuator profiles execute pre-cached, highly deterministic fallback logic.
15Does the system use generic LLMs for routing?Only within strict API wrappers; never for deterministic, machine-speed routing.
16How are legacy physical systems supported?Through deployable OpenC2 actuator proxy translation layers.
17What triggers an incident reopening?New causal W3C PROV links mathematically mapping to historical event nodes.
18How is analyst alert fatigue minimized?By auto-closing low-threat events strictly within the certified [Figure omitted from source export] budget bound.
19What is the legal standard of care?Engineered navigation and mathematical risk control, replacing administrative compliance.
20How are false negatives directly controlled?The lower deferral threshold ([Figure omitted from source export]) is calibrated to the strict budget [Figure omitted from source export].
21Are non-digital fallbacks required?Yes, mechanical interlocks and analog governors are mandatory for safety-critical OT.
22How is timing manipulation stopped?Causal dependency graphs (W3C PROV) natively resist simple timestamp spoofing.
23Can the orchestrator be flooded by alerts?Rate-limiting and tokenized queuing structurally prevent resource exhaustion.
24How are False Positives systematically corrected?They are fed directly back into the continuous Mondrian calibration pool.
25What if IT logs and OT physics logs disagree?Consensus algorithms structurally defer to the physics-based OT invariants.
26Is there a manual override for automation?Yes, cryptographic human-in-the-loop tokens can preempt any active CACAO playbook.
27How is network actuator saturation prevented?State-aware proxies consolidate and prune redundant OpenC2 commands.
28What strictly defines an incident?A statistically significant deviation lacking verified benign physical attribution.
29Are system backups trusted blindly for recovery?No, recovery playbooks mandate stringent cryptographic integrity checks prior to restoration.
30How is regulatory compliance documented?Through immutable cryptographic ledgers recording all CACAO playbook actions.

Table 5: 30 Direct-Answer Items for Assurance Audits \[technical proposal\]2.

30 Page Concepts for Organizational Playbooks

\#Page Concept\#Page Concept
1W3C PROV Architecture Schema16Conformal Calibration Pool Maintenance
2OpenC2 Actuator Network Map17False Positive Feedback Loop
3Mondrian Conformal Math Primer18Threat Intelligence Ingestion API
4Dempster-Shafer Fusion Logic19Zero-Knowledge Reporting Standards
5CACAO Playbook JSON Templates20Out-of-Band Communications Protocol
6Exponential Decay Tuning Guide21Mechanical Interlock Override Procedures
7Digital Twin Invariant Definitions22Identity Token Revocation Workflows
8Telemetry Ingestion Routing23Hardware-Enforced Segmentation Map
9Dual-Threshold Configuration24Sensor Blinding Detection Heuristics
10Incident Lifecycle Flowchart25NTP Synchronization and Auditing
11Fallback and Degradation Policies26Causal Graph Traversal Algorithms
12Safety-Critical Action Matrix27Ledger-Backed Attestation Queries
13Human-in-the-Loop Token Guide28Adversarial Testing Emulation Scripts
14API Key and Cert Management29SOC Analyst Triage Dashboard UI
15OT/IT Boundary Firewall Rules30Executive Cyber-Physical Risk Dashboard

Table 6: 30 Page Concepts for Organizational Documentation \[technical proposal\].

12. Open Research Questions

1. Non-Digital Integration: How can mechanical interlocks and analog governors natively communicate state changes to a digital OpenC2 orchestrator without introducing exploitable digital attack surfaces? \[hypothesis\]

2. Quantum-Resistant Attestation: How will the computational overhead of post-quantum cryptographic standards impact the execution latency of ledger-backed W3C PROV causality graphs at machine speed? \[unknown\]

3. Cross-Organizational Orchestration: How can automated CACAO playbooks safely execute collaborative response actions across distinct legal and sovereign boundaries without violating data privacy and data sovereignty laws? \[policy proposal\]

13. Contradiction Register

  • Validation-Tuned Heuristics vs. Conformal Prediction: Traditional IT security tools rely heavily on manually tuned thresholds, leading to unpredictable failure under distribution shifts. This directly contradicts the SSE mandate for finite-sample guarantees, which is exclusively provided by conformal prediction models \[disputed claim\]2.
  • Linear Lifecycle vs. Continuous Monitoring: NIST SP 800-61 Revision 2 utilized a strictly phased, linear approach to incident handling (Detection, Analysis, Containment). Revision 3 completely contradicts this by aligning with CSF 2.0, treating incident response as a fluid, continuous monitoring loop integrated deeply into enterprise risk management \[current official policy\]10.

14. Claim-Status Table

ClaimLabelDate
Cyber-physical systems require engineered resilience, not administrative compliance.\[institutional analysis\]Aug 2026
NIST SP 800-61 Rev 3 supersedes Rev 2, aligning with CSF 2.0 and DE.CM.\[established standard or law\]Aug 2026
W3C PROV prevents dependency explosion via causal tracking in threat modeling.\[peer-reviewed research finding\]Aug 2026
Mondrian CP maintains minority-class coverage under extreme data imbalance.\[peer-reviewed research finding\]Aug 2026
OpenC2 standardizes technology-agnostic actuator commands across platforms.\[established standard or law\]Aug 2026
Trust and Evidence states decay exponentially over time (LDF).\[peer-reviewed research finding\]Aug 2026
Dual-threshold systems mathematically degrade to manual triage if [Figure omitted from source export] is low.\[reasoned inference\]Aug 2026

Table 7: Status of Major Architectural Claims \[reasoned inference\].

15. Source-Quality Table

Source TopicAuthority LevelBias / Limitation
NIST SP 800-160 (Vol 1 & 2\)High (Primary Government Standard)Highly theoretical in nature; requires extensive translation to implement practically.
NIST SP 800-61 (Rev 2 & 3\)High (Primary Government Standard)Rev 2 is deprecated; Rev 3 shifts focus heavily to management over technical playbooks.
Conformal Prediction (CADES)High (Peer-Reviewed ML Research)Evaluated primarily on static datasets; real-time latency in graph structures requires optimization.
W3C PROV & Provenance GraphsHigh (International Web Standard)Implementation overhead can significantly impact CPU/Memory on edge IoT devices.
OASIS OpenC2 & CACAOHigh (International Security Standard)Vendor adoption is ongoing; currently requires middleware proxies for legacy systems.
Exponential Decay (LDF)Moderate (Peer-Reviewed Research)The decay rate ([Figure omitted from source export]) requires careful, empirical tuning per environment to avoid premature forgetting.

(Research cutoff date: August 2026\)

Works cited

1. Beyond the IT Checklist: Engineering a Reasonable Standard of Care for Cyber Safety, https://arxiv.org/html/2606.13612

2. Dual-Threshold Conformal Deferral for Trustworthy Security Alert Triage \- Preprints.org, https://www.preprints.org/manuscript/202607.1357

3. Class-Conditional Conformal Prediction for Reliable Anomaly Detection Under Extreme Class Imbalance \- MDPI, https://www.mdpi.com/2504-4990/8/7/190

4. Bill of Materials Integration in OpenC2 for Enhanced Context Discovery \- WebThesis, https://webthesis.biblio.polito.it/39723/1/tesi.pdf

5. Real-time threat, impact analysis and response automation for SOC/CSIRT operations \- PvIB, https://www.pvib.nl/kenniscentrum/documenten/soccrates-real-time-threat-impact-analysis-and-response-automation-for-soc-csirt-operations

6. OASIS CACAO v2.0 standard is released \- JCOP, https://jcop.eu/blogposts/blog16.html

7. Temporal Decay Loss for Adaptive Log Anomaly Detection in Cloud Environments \- PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC12073674/

8. Threat detection and investigation with system-level provenance graphs: A survey \- Northwestern Computer Science, https://users.cs.northwestern.edu/\~ychen/Papers/ProvGraph\_survey\_2021.pdf

9. NIST Computer Security Incident Handling Guide: A Comprehensive Overview \- GitHub, https://github.com/tomwechsler/Ethical\_Hacking\_and\_Penetration\_Testing/blob/main/Documentation/NIST\_Computer\_Security\_Incident\_Handling\_Guide.md

10. Updated NIST Incident Response Guidance: SP 800-61 Rev. 3 \- Tandem, https://tandem.app/blog/updated-nist-incident-response-guidance-sp-800-61-rev-3

11. Towards intelligence driven automated incident response \- WebThesis, https://webthesis.biblio.polito.it/22865/1/tesi.pdf

12. Draft SP 800-160 Vol. 2, Systems Security Engineering, https://csrc.nist.gov/files/pubs/sp/800/160/v2/ipd/docs/sp800-160-vol2-draft.pdf

13. NIST SP 800-160 focuses on plugging security into systems engineering to develop defensible, survivable systems \- Industrial Cyber, https://industrialcyber.co/nist/nist-sp-800-160-focuses-on-plugging-security-into-systems-engineering-to-develop-defensible-survivable-systems/

14. A comparative Study on Cyber Threat Intelligence: The Security Incident Response Perspective, https://epub.uni-regensburg.de/52721/1/Accepted\_CTI%20Survey\_IEEE\_COMST\_public.pdf

15. Flurry: a Fast Framework for Reproducible Multi-layered Provenance Graph Representation Learning \- arXiv, https://arxiv.org/pdf/2203.02744

16. Security Approaches for Data Provenance in the Internet of Things: A Systematic Literature Review \- arXiv, https://arxiv.org/html/2407.03466v1

17. (PDF) Distributed attack prevention using Dempster-Shafer theory of, https://www.academia.edu/36967755/Distributed\_attack\_prevention\_using\_Dempster\_Shafer\_theory\_of\_evidence

18. Multisensor Data Fusion in IoT Environments in Dempster–Shafer Theory Setting: An Improved Evidence Distance-Based Approach \- MDPI, https://www.mdpi.com/1424-8220/23/11/5141

19. Cyber-Physical Systems Security: A Comprehensive Review of Anomaly Detection Techniques \- arXiv, https://arxiv.org/html/2502.13256v2

20. Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark \- arXiv, https://arxiv.org/html/2607.27143

21. Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT\&CK-Aligned Triage as a Worked Instance \- arXiv, https://arxiv.org/html/2608.12444v1

22. Automatic incident response solutions: a review of proposed solutions' input and output, https://www.researchgate.net/publication/373483648\_Automatic\_incident\_response\_solutions\_a\_review\_of\_proposed\_solutions'\_input\_and\_output

23. Class-Conditional Conformal Prediction for Reliable Anomaly Detection Under Extreme Class Imbalance \- ResearchGate, https://www.researchgate.net/publication/408397599\_Class-Conditional\_Conformal\_Prediction\_for\_Reliable\_Anomaly\_Detection\_Under\_Extreme\_Class\_Imbalance

24. Exponential Decay Law | Law | Research Starters \- EBSCO, https://www.ebsco.com/research-starters/law/exponential-decay-law

25. (PDF) Temporal Effects of Contributing Factors in Insider Risk Assessment: Insider Threat Indicator Decay Characteristics \- ResearchGate, https://www.researchgate.net/publication/369529292\_Temporal\_Effects\_of\_Contributing\_Factors\_in\_Insider\_Risk\_Assessment\_Insider\_Threat\_Indicator\_Decay\_Characteristics

26. Is there a Half-Life for the Success Rates of AI Agents? \- Toby Ord, https://www.tobyord.com/writing/half-life

27. HermBuild: Hermeticity-Based Reproducible Builds with Verifiable Provenance and Software Supply Chain Attestation \- JETIR.org, https://www.jetir.org/papers/JETIR2511644.pdf

28. Achieving Cyber Resilience through standards-based, machine, https://phoeni2x.eu/2024/08/29/achieving-cyber-resilience-through-standards-based-machine-processable-executable-incident-response-business-continuity-playbooks/