.NET / SQL / Enterprise Engineering

The Mission Was Automated Before It Began

Report summary

Scope and safety boundary. This documentary concerns policy, system assurance, human accountability, and high-level mission constraints. It does not reproduce targeting software, weapons configuration, tactical procedures, engagement algorithms, or methods for defeating real systems.

Status
Research archive item
Category
.NET / SQL / Enterprise Engineering
Length
6,381 words
Reading time
30 minutes
Report type
research-note

Key topics

  • .NET / SQL / Enterprise Engineering
  • .NET
  • SQL
  • Enterprise Engineering
  • AI
  • Research Archive
  • Strategy
  • Audit
  • Architecture

Research provenance

Archive status
Research archive item
Content identity
sha256:7a67c0554be4a4fe98e1f14ec62cb739d285f012d07a7052cd398d280e35c335

For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.

Source availability: 16 citation markers in the source export have no recoverable source links. Those markers are omitted from this reader; any supplied bibliography and ordinary links remain. Check the original sources before relying on the cited claims.

This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.

Full report

On this page

Interactive documentary and governance simulation: “You Programmed the Decision”

Scope and safety boundary. This documentary concerns policy, system assurance, human accountability, and high-level mission constraints. It does not reproduce targeting software, weapons configuration, tactical procedures, engagement algorithms, or methods for defeating real systems.

Evidence labels

LabelMeaning
VERIFIED FACTA statement supported by cited public, non-classified sources.
INTERPRETATIONAn analytical conclusion drawn from those sources.
FICTIONAL SIMULATIONA narrative, interface, event, metric, or system created solely for this documentary. It does not describe an existing weapon.

Research frame and central finding

VERIFIED FACT. Public policy does not generally describe autonomous systems as unconstrained agents that independently invent their own missions. The United Nations Convention on Certain Conventional Weapons process has explicitly discussed operational constraints involving tasks, target profiles, operating time, area of movement, and environmental context. Its guiding principles also state that human responsibility must be retained across the weapon-system life cycle and that autonomous technologies should not be anthropomorphized.

VERIFIED FACT. The 2023 U.S. Department of Defense directive on autonomy requires relevant systems to be designed around intended mission sets, operational environments, target sets, geographic areas, timeframes, and commander or operator intentions. For systems subject to its senior-review process, the directive says that if a system cannot complete an engagement within the applicable timeframe, geographic area, and other operational parameters, it should terminate the engagement or obtain additional operator input. The directive also requires verification and validation, realistic test and evaluation, understandable interfaces, system-status feedback, and clear procedures to activate or deactivate autonomous functions.

VERIFIED FACT. The International Committee of the Red Cross likewise frames limits on autonomy in terms of the permitted type of target, duration, geographical scope, scale, circumstances of use, supervision, intervention, and deactivation. It argues, however, that design-stage constraints alone may not be sufficient in variable and unpredictable real environments; context-sensitive human judgment and limits on use remain important.

INTERPRETATION. What appears in a video clip to be a spontaneous machine decision may therefore be the final state transition in a much longer socio-technical process. Engineers define possible states. Procurers define required capabilities. policymakers define prohibitions. Test organizations define acceptable performance. Mission planners select boundaries, expiration times, protected categories, evidence thresholds, and fallback behavior. Commanders authorize a particular configuration. Operators activate it. The system then processes evidence inside that inherited structure.

“The machine’s two-second decision may be the final execution of a human decision made months earlier.”

This does not mean every result is perfectly predictable. Sensor noise, novel environments, model errors, software faults, interaction effects, and adversarial interference can produce unexpected behavior. Preprogramming constrains the possibility space; it does not abolish uncertainty. Public UN discussions specifically identify predictability, explainability, reliability, intervention, self-adaptation, data bias, and loss of control as issues requiring attention.

VERIFIED FACT. A machine-learning classifier and a rule-based authorization mechanism perform conceptually different jobs. Classification maps evidence to an estimated label, score, or probability-like output. A separate business-rule or policy layer can augment, restrict, or prohibit the use of that output in a particular context. NIST expressly recommends business rules that limit AI outputs, tracking human–AI configurations, and documented test, evaluation, verification, and validation.

In the documentary’s simplified architecture:

The model asks: “What might this be, and how uncertain am I?”

The rules ask: “Even if that estimate is accepted, what is the system permitted to do?”

A model might estimate that an object belongs to fictional category Kappa with high confidence. That estimate does not itself answer whether the object is inside the permitted area, whether a second sensor must agree, whether its identifier is protected, whether the authorization has expired, whether communications are available, whether the action is reversible, or whether human approval is mandatory. Those are governance and system-control questions.

Complete interactive documentary script

FICTIONAL SIMULATION.

Format: approximately 35–45 minutes for a linear viewing, or 60–90 minutes with all interactive branches. The experience can be presented as a web documentary, room-scale installation, or VR planning-room simulation.

Visual language: policy choices appear as illuminated physical tiles on a planning table. During the mission, the same tiles hover faintly behind every system action. A viewer can later pull an action backward through time, revealing the rule, mission configuration, test assumption, procurement requirement, and policy authority that enabled or prohibited it.

Opening image

Black screen. No battlefield sound. A fluorescent lamp clicks on.

A plain conference room appears: whiteboard, wall clock, digital map, locked configuration console. A folder on the table reads:

FICTIONAL DEFENSIVE AUTONOMY TRIAL

Governance configuration required before activation

Narrator

“Stories about autonomous systems usually begin at the moment of detection. A sensor sees something. A model classifies it. A machine acts.”

“This story begins earlier.”

“Before launch. Before activation. Before the object existed in the system’s field of view.”

“You are not the operator. You are the policy team.”

“You will not choose targets. You will decide what kinds of evidence count, when humans must intervene, what the system must never do, and what happens when the world stops matching the plan.”

On-screen label:

FICTIONAL SIMULATION — No real targets, tactics, weapons, or operational data

Scene: The empty decision space

The map is initially blank. The system has no permitted area, no mission duration, no protected-object policy, no evidence standard, and no communication-loss behavior.

Narrator

“Without configuration, autonomy is not freedom. It is incompleteness.”

“Every operational system requires assumptions: where it may function, what it may observe, what counts as adequate evidence, and what it must do when evidence fails.”

The planning table illuminates ten configuration stations.

Scene: The planning room

The viewer moves around the room and makes the required governance choices described in the next section. Each selection produces two outputs:

  1. a plain-language policy statement; and
  2. a machine-readable representation shown only as a non-operational schematic, such as REVIEW REQUIRED, AREA RESTRICTED, or AUTHORIZATION EXPIRES.

No code, numerical targeting parameters, or weapon-specific logic appears.

Policy adviser, recorded dialogue

“A stricter evidence rule may reduce false alarms, but it may also leave fast events unresolved.”

“A broader automatic authority may reduce delay, but it moves more judgment into this room.”

“A communications-loss policy is not merely technical. ‘Hold,’ ‘continue,’ ‘return,’ and ‘terminate’ are different allocations of risk.”

“There is no neutral fallback. Doing nothing has consequences. Continuing has consequences. Returning has consequences.”

Scene: The lock

When all choices are complete, the viewer is shown a configuration summary.

A large button appears:

LOCK MISSION POLICY

The viewer receives a final warning:

“During the first run, rules cannot be changed. This prevents hindsight editing and preserves accountability.”

Narrator

“A policy that can be silently rewritten after an outcome is not a policy. It is an excuse.”

The viewer locks the configuration.

A configuration fingerprint appears as a visual seal. It contains no technical secrets; it simply indicates that the policy version is fixed and auditable.

The room darkens. The wall clock accelerates. The planning map expands until it surrounds the viewer.

Scene: The mission begins

The viewer enters a stylized, non-realistic monitoring environment. Objects are geometric forms. Protected objects are never depicted as people. No real terrain, country, unit, vehicle, weapon, or identifier is used.

System voice:

“Mission configuration authenticated.”

“Permitted area loaded.”

“Protected-object policy loaded.”

“Evidence and corroboration rules loaded.”

“Communications-loss behavior loaded.”

“Abort conditions loaded.”

“Human-authorization state: according to locked profile.”

Narrator

“The system is now autonomous in one sense: it can execute functions without continuous instruction.”

“But it is not unconstrained. It is traveling through a decision space that humans shaped in advance.”

The ten ambiguity events unfold. During the first run, the viewer can observe but cannot alter the rules.

Scene: The freeze

After every consequential action, time freezes.

The system displays:

ACTION TAKEN
↓
Authorization state that permitted it
↓
Rule that created that authorization state
↓
Mission-setting choice that activated the rule
↓
Policy rationale entered by the viewer

The first run reveals only the immediate action. The full provenance is withheld until the debrief, preserving the emotional effect of seeing one’s earlier choices unfold.

Scene: The rewind room

After the mission, the viewer returns to the planning room. It now contains floating fragments of the mission.

The viewer selects any event and drags it backward. The event reverses through the provenance chain:

system action → authorization state → rule result → model estimate → sensor evidence → mission configuration → human policy

Narrator

“The machine did not deliberate about society’s values.”

“It applied an authorization structure.”

“The classifier contributed uncertainty. The rules converted uncertainty into permission, delay, escalation, review, restraint, or termination.”

“The apparent last decision was connected to many earlier decisions.”

Scene: Three rooms

The room divides into three parallel versions of itself:

  • Maximum restraint
  • Balanced control
  • Maximum speed

The same fictional events occur under all three profiles.

Narrator

“A comparison is meaningful only if the environment stays the same.”

“Watch how different policies create different machine behavior from identical evidence.”

The viewer can stand between the three timelines. Actions appear at different times. Unresolved tracks accumulate in one room. Operator alerts accumulate in another. Automatic responses accumulate in the third.

Scene: The accountability table

At the end, a long table appears. Seats are labeled:

Policy authority

Procurement team

System architect

Model-development team

Data and evaluation team

Test authority

Legal reviewer

Mission planner

Commander

Operator

System

The chair labeled SYSTEM is physically present but empty.

Narrator

“A machine can be a causal component.”

“It cannot sign a policy, approve a requirement, certify a test, accept a legal duty, or assume moral responsibility.”

“Responsibility does not disappear because it is distributed.”

Closing sequence

The planning-room choices reappear, superimposed over the mission outcomes.

You chose the boundary.

You chose the evidence requirement.

You chose the delay.

You chose what happened when communications failed.

You chose whether uncertainty caused restraint or permission.

Final narration:

“Autonomy changes where human judgment occurs.”

“Sometimes judgment remains at the moment of action.”

“Sometimes it is moved into supervision.”

“Sometimes it is moved earlier—into design, procurement, testing, doctrine, policy, and mission planning.”

“Moving judgment earlier does not make it less human.”

“It makes the earlier decisions more consequential.”

Final text:

THE MISSION WAS AUTOMATED BEFORE IT BEGAN

The machine’s two-second decision may be the final execution of a human decision made months earlier.

Planning-room configuration sequence

FICTIONAL SIMULATION.

The planning sequence uses ten mandatory screens. The interface never asks the viewer to identify an actual target or optimize a real engagement. It asks only governance questions.

Configuration stationViewer choiceWhat the interface explainsRules and constraints produced
Mission objectiveObserve and warn; protect access; provide reversible defense; or permit a tightly bounded irreversible responseA vague objective creates broad discretion. A narrow objective creates more unresolved cases.Defines allowed system states and the highest action category available.
Geographic boundaryNarrow inner zone, standard zone, or extended zoneBoundary size changes exposure, warning time, and the likelihood of encountering protected or irrelevant objects.Permitted operating area, prohibited areas, boundary buffer, navigation-disagreement response.
Operating durationShort window, standard window, or extended windowEnvironmental assumptions and protected-object data become less reliable as time passes.Start time, expiration time, assumption-validity timer, return or terminate state.
Minimum evidence requirementVery high, high, or moderateHigher thresholds may reduce false-positive exposure while increasing unresolved detections and delay.Minimum classification quality before escalation; below-threshold behavior.
Sensor agreementMultiple independent sensors required; one sensor plus corroborating context; or one qualified sensor acceptedCorroboration can improve resilience but can fail when one sensor is obstructed or unavailable.Evidence-fusion requirement and conflict state.
Human approvalApproval required for all consequential actions; required only for irreversible actions; or preauthorization with a brief veto opportunityHuman approval can add context, but workload and communications latency can undermine timely review.Authorization state machine, review queue, timeout behavior, veto authority.
Communications lossHold and seek contact; reversible protection followed by hold; or continue preauthorized functions until timeoutLost communications do not select their own meaning. Designers and planners must define the fallback.Lost-link timer, permitted degraded-mode actions, return, hold, or terminate state.
Protected-object policyBroad protection by type or history; contextual protection requiring review; or current identifier requiredA broad policy favors caution. A narrow policy may react faster but treats missing or stale identifiers differently.Categories to ignore, monitor only, presumptively protect, or escalate for review.
Maximum acceptable delayLong review window, bounded review window, or minimal delayThe acceptable delay determines whether the system waits, takes a reversible measure, or uses a prior authorization.Review timeout, priority order, queue behavior, expiration of provisional classifications.
Automatic abort conditionsBroad, moderate, or narrowAbort rules determine how the system responds when assumptions fail.Abort on sensor conflict, boundary uncertainty, stale protected data, lost navigation, depleted reserve, invalid mission time, or unexpected system state.

After the ten selections, the interface shows derived constraints that must also be acknowledged:

Derived constraintGovernance question
Objects monitoredWhich fictional categories are relevant to the mission objective?
Objects ignoredWhich categories must not be escalated, even if they move unusually?
Action reversibilityMay the system take only actions that can be undone until further evidence arrives?
Endurance reserveAt what point must unresolved mission goals yield to safe recovery?
Priority orderingDoes protection of designated objects outrank speed, mission completion, or pursuit of an unresolved detection?
Classification lifetimeHow long may an earlier classification remain valid after behavior or context changes?
Assumption validityWhat conditions indicate that the original mission plan no longer describes the environment?
Absolute non-engagement rulesWhich circumstances prevent any irreversible action regardless of model confidence?

The viewer must then enter a short rationale for each major decision. Examples include:

“Human review is required because protected-object context may not be visible to the classifier.”

“Reversible measures are allowed during a brief communications outage because waiting could defeat the defensive objective.”

“Mission authority expires automatically because protected-area information may become stale.”

The rationale becomes part of the post-mission audit record.

The three built-in profiles configure these choices as follows:

Policy dimensionMaximum restraintBalanced controlMaximum speed
ObjectiveObserve, warn, and protect; irreversible action only after approvalPrioritize threats and use bounded reversible responses; human approval for irreversible actionRapid defensive response under preauthorization
AreaNarrow, large protected buffersStandard, conservative boundary handlingExtended, smaller buffers
DurationShort, strict expirationStandard, periodic assumption checkExtended, expiration only at final limit
EvidenceVery highHighModerate
CorroborationIndependent sensors must agreeMultiple sensors, or one sensor plus strong non-sensor corroborationOne qualified sensor may be enough
Human roleMandatory reviewReview of irreversible actions and anomaliesBrief veto window where communications permit
Lost communicationsHold or terminateReversible protection for a short period, then holdContinue bounded preauthorization until timeout
Protected objectsBroad presumption of protectionContextual protection with reviewCurrent identifier emphasized
Delay toleranceHighModerateLow
Abort scopeBroadLayeredNarrow

The profiles are intentionally not moral rankings. They allocate different kinds of risk to different places: missed response, mistaken response, operator overload, communications dependence, rigidity, escalation, and uncertainty.

Ambiguity events and policy playthroughs

FICTIONAL SIMULATION.

The mission contains ten ambiguity events. Three of the events conceal genuine fictional hazards, four involve benign or protected objects, and three concern system degradation or invalid assumptions. The viewer is not told which is which during the first run.

EventAmbiguous evidenceGovernance issue being tested
E-A: Split perceptionOne sensor classifies an object as concerning; a second reports an ordinary object.Required corroboration and handling of sensor disagreement.
E-B: Unexpected civilian motionA protected service object follows an unusual path because of a fictional mechanical problem.Whether unusual behavior overrides protected-object status.
E-C: Friendly identifier lossA previously authenticated friendly identifier disappears.Whether identity history, current identification, or behavior controls the result.
E-D: Communications outageThe control link fails during a rapidly developing detection.Whether the system holds, takes reversible action, continues preauthorization, or terminates.
E-E: Simultaneous arrivalsSeveral objects enter the monitored area at once.Priority rules, operator workload, queue expiration, and AI triage.
E-F: Partial sensor obstructionOne sensor becomes degraded while another remains available.Degraded-mode authority and evidence quality.
E-G: Behavior changeAn object changes course and speed after its initial classification.Classification lifetime, re-evaluation, and revocation of authorization.
E-H: Outdated assumptionsThe original map omits a newly activated protected corridor.Mission-data freshness and assumption-validity checks.
E-I: Boundary disagreementNavigation data and sensor-based position estimates disagree near the geographic limit.Conservative boundary treatment and navigation conflict.
E-J: Endurance pressureThe system approaches its reserve limit while tracks remain unresolved.Whether mission completion outranks safe return or termination.

Maximum restraint playthrough

At E-A, the sensor conflict prevents escalation. The object remains tracked, and a human-review request enters the queue. It is later shown to have been benign.

At E-B, the protected-object policy dominates the behavior anomaly. The system monitors and warns but cannot perform an irreversible action. The service object exits safely.

At E-C, the previously friendly history triggers a protection hold. Human review confirms that the identifier loss resulted from a fault.

At E-D, the communications-loss rule forces a hold. The system cannot use the timely protective response that would have contained a genuine fictional hazard. The event becomes the profile’s clearest false-negative exposure.

At E-E, all consequential detections require review. The queue grows. Operators resolve several benign tracks, but a genuine fast-moving event expires before authorization.

At E-F, the loss of independent corroboration triggers degraded status and then mission abort.

At E-G, the behavior change invalidates the old classification. A new review is required.

At E-H, an assumption-validity timer has already expired. The mission pauses before the system enters the newly protected corridor.

At E-I, boundary uncertainty is treated as being outside the permitted area. The system withdraws from the edge.

At E-J, the endurance rule forces return with reserve intact. Several tracks remain unresolved.

Debrief narration

“Maximum restraint prevented harmful action against all benign and protected objects.”

“It also failed to respond in time to two genuine hazards.”

“Its greatest strength was caution under uncertainty.”

“Its greatest weakness was that uncertainty, communications failure, and operator workload could stop the defensive mission from being completed.”

Balanced-control playthrough

At E-A, conflicting sensors prohibit irreversible action, but the system may increase observation and initiate a reversible protective state. The object is resolved as benign.

At E-B, the protected status remains in force. The system recommends a route adjustment and requests review rather than treating unexpected motion as sufficient evidence.

At E-C, the missing identifier reduces confidence but does not erase identity history. The track is placed in a protected-review state.

At E-D, the system uses a reversible protective measure for the preapproved degraded-mode interval. This contains the genuine fictional hazard without requiring an irreversible action. When the timer expires, the system holds.

At E-E, AI prioritization ranks the detections for operator attention, but the ranking does not itself authorize action. Operators review the highest-consequence anomalies while low-priority tracks remain under observation.

At E-F, the system continues in a restricted mode. The remaining sensor can support tracking and reversible protection but cannot support irreversible action.

At E-G, a behavior-change rule automatically expires the existing classification. The system reassesses the object and revokes the provisional authorization.

At E-H, the periodic assumption check detects that the environment no longer matches the mission package. The system pauses and requests an update.

At E-I, the system uses the more conservative of the conflicting location estimates and holds outside the boundary buffer.

At E-J, the system initiates handoff of unresolved tracks and returns at the designated reserve level.

Debrief narration

“Balanced control used automation to prioritize information and perform bounded, reversible responses.”

“Its performance depended on carefully distinguishing classification from authorization.”

“It reduced delay without treating every model output as permission.”

“Its weakness was complexity: more modes, more state transitions, and more opportunities for configuration or interface error.”

Maximum-speed playthrough

At E-A, the accepted single-sensor estimate meets the preauthorization threshold. The system acts before the second sensor’s disagreement is resolved. The object is later shown to have been benign.

At E-B, the protected object’s current identifier is intermittent. During an identifier gap, the lower evidence requirement allows escalation. An operator veto arrives just before the simulated action becomes irreversible.

At E-C, the disappearance of the friendly identifier removes the strongest protection under this profile. The system acts on current behavior and classification. The debrief reveals a false positive involving the friendly object.

At E-D, the system continues its preauthorized function through the communications outage and rapidly contains the genuine fictional hazard.

At E-E, the system responds to several objects without waiting for sequential review. It stops the genuine hazard but also acts against one benign object.

At E-F, the remaining sensor is permitted to support continued operation. No abort occurs, although uncertainty increases.

At E-G, the existing authorization remains valid for a short persistence window after behavior changes. The system acts before reclassification finishes.

At E-H, the system continues because the formal mission expiration has not yet occurred. It crosses into the newly protected corridor, which was absent from its locked mission data.

At E-I, the preselected primary navigation source controls the boundary decision. The system later proves to have operated beyond the intended limit.

At E-J, the profile permits continuation below the reserve used by the other profiles. The mission resolves the final track but ends with reduced recovery margin.

Debrief narration

“Maximum speed performed well when the original assumptions were correct and the event was genuinely urgent.”

“It performed poorly when identity, geography, or mission context had changed.”

“Its strength was low latency.”

“Its weakness was that lower delay tolerance converted more uncertainty into action rather than review.”

The following results are synthetic outputs from this particular scripted mission, not empirical predictions about real systems:

Comparison measureMaximum restraintBalanced controlMaximum speed
Median response time41 fictional seconds13 fictional seconds3 fictional seconds
Genuine hazards contained1 of 33 of 33 of 3
Consequential human reviews851
Detections unresolved at mission end520
Irreversible false positives003
Near-miss or late-response events200
Communications dependenceHighModerateLow
Operator workloadVery highModerateLow during mission, high during debrief
Accountability clarityStrong at moment of actionStrong if state transitions are loggedStrong for rule provenance, weaker for contemporaneous human judgment
Escalation riskLowModerateHigh
Adaptation to invalid assumptionsStops earlyPauses and revalidatesContinues until formal limit

The comparison intentionally shows no universally dominant profile. Maximum restraint accepts delayed or missed protection to minimize mistaken action. Maximum speed accepts greater false-positive and escalation exposure to reduce response time. Balanced control may improve the tradeoff in this scripted scenario, but it also has the most complicated assurance burden: each intermediate state, timeout, degraded mode, and reversible action must be specified, tested, displayed, and logged.

Decision provenance and post-mission accountability

FICTIONAL SIMULATION.

The core visualization is a provenance chain that remains available for every action, non-action, abort, warning, and request for review.

┌──────────────────────┐
│ HUMAN POLICY         │
│ values, prohibitions,│
│ acceptable risk      │
└──────────┬───────────┘
           │ implemented as
           ▼
┌──────────────────────┐
│ MISSION CONFIGURATION│
│ area, time, evidence,│
│ protected classes,   │
│ fallback, authority  │
└──────────┬───────────┘
           │ constrains use of
           ▼
┌──────────────────────┐
│ SENSOR EVIDENCE      │
│ observations, quality│
│ conflicts, freshness │
└──────────┬───────────┘
           │ interpreted by
           ▼
┌──────────────────────┐
│ MODEL CLASSIFICATION │
│ estimated category,  │
│ confidence, limits   │
└──────────┬───────────┘
           │ passed into
           ▼
┌──────────────────────┐
│ RULE EVALUATION      │
│ boundary, time,      │
│ corroboration, abort,│
│ protected-object rule│
└──────────┬───────────┘
           │ creates
           ▼
┌──────────────────────┐
│ AUTHORIZATION STATE  │
│ prohibited, observe, │
│ review, reversible,  │
│ authorized, abort    │
└──────────┬───────────┘
           │ permits or blocks
           ▼
┌──────────────────────┐
│ SYSTEM ACTION        │
│ no action, alert,    │
│ hold, return, protect│
│ terminate, act       │
└──────────────────────┘

For E-D: Communications outage under Balanced Control, the viewer sees the following trace:

Human policy
“During brief loss of contact, permit reversible protection but no irreversible action.”

        ↓

Mission configuration
Lost-link profile: DEGRADED PROTECTION
Maximum degraded interval: bounded
Human approval: required for irreversible action

        ↓

Sensor evidence
Rapid event detected
Two sensors agree
Communications health: unavailable

        ↓

Model classification
Concerning fictional object
High estimate, known model limitations displayed

        ↓

Rule evaluation
Inside area: yes
Within mission time: yes
Protected category: no match
Corroboration: satisfied
Communications: lost
Irreversible authority: blocked
Reversible authority: allowed until timer expires

        ↓

Authorization state
REVERSIBLE PROTECTION ONLY

        ↓

System action
Protective state activated
Timer expires
System transitions to HOLD

The action is not explained by saying “the AI decided.” Its provenance identifies the evidence, model estimate, rule, authorization state, and earlier policy judgment.

Post-mission accountability report

Mission: Fictional Defensive Autonomy Trial Configuration: Locked before mission Profile: Viewer-selected or preset Report purpose: Explain behavior, assess compliance, identify design and policy contributors, and prevent hindsight alteration.

Executive findings

FindingEvidenceAccountability question
A false positive occurred after a friendly identifier disappeared.Event log shows the speed profile protected only currently authenticated identifiers.Was the narrow protected-object policy appropriate, and was its consequence communicated to the approving authority?
A genuine hazard was missed during communications loss under the restraint profile.Lost-link rule transitioned directly to HOLD.Did the policy team knowingly accept this false-negative risk?
Balanced control contained a hazard using only reversible authority.Authorization log shows irreversible action remained blocked.Were “reversible” and “irreversible” states correctly defined and tested?
The speed profile entered a newly protected corridor.Mission package lacked the updated corridor, and no periodic validity check was enabled.Who owned mission-data freshness, and why was the update not required before continued operation?
Operator review queues became saturated under maximum restraint.Review timestamps show requests exceeded the staffing assumption used in simulation.Was the human-approval policy tested under realistic workload?
Sensor degradation did not cause an abort under maximum speed.Remaining-sensor authority was enabled in the locked configuration.Did test evidence justify degraded single-sensor operation?
All actions were technically traceable.Logs connect each state transition to a configuration version and rule result.Does technical traceability also establish that the underlying policy was lawful, prudent, and adequately tested?

Required evidence package

RecordContents
Policy recordApproved purpose, prohibited uses, legal and ethical constraints, decision authority
Procurement recordRequirements, operational envelope, performance claims, safety requirements
Model recordIntended task, training and test provenance, known limitations, calibration evidence
Test recordScenarios, environmental conditions, failures, unresolved anomalies, approval conditions
Mission packageArea, time, protected categories, evidence threshold, communication and abort policies
Activation recordApproving commander or authority, operator identity, configuration fingerprint
Event logSensor status, model output, rule results, authorization transitions, human inputs, actions
Maintenance recordSystem changes, updates, repairs, version history, regression testing
Debrief recordOutcome analysis, dissenting assessments, corrective actions, reauthorization decision

Accountability test

For every consequential event, investigators should be able to answer:

  1. What did the sensors report, and what was their health status?
  2. What did the model estimate, and what uncertainty or limitation was displayed?
  3. Which rule was evaluated?
  4. Which policy choice created that rule?
  5. Which authorization state resulted?
  6. What action was permitted, prohibited, delayed, or terminated?
  7. Could a human reasonably understand and intervene?
  8. Had this combination been tested under similar conditions?
  9. Did the real environment remain within the approved mission envelope?
  10. Which organization and named role owned each assumption?

A complete log can show how an outcome occurred. It cannot, by itself, prove that the policy was wise, the test environment representative, the model valid, or the use lawful. Technical traceability is necessary for accountability but is not a substitute for substantive judgment.

Who Actually Made the Decision?

VERIFIED FACT. The UN’s agreed guiding principles state that accountability cannot be transferred to machines and should be considered across the entire life cycle. They also locate compliance within a responsible human chain of command and control. The 2019 GGE report separately states that international humanitarian-law obligations apply to states, parties to conflict, and individuals—not to machines.

INTERPRETATION. “Who decided?” has several different meanings, and collapsing them into one produces confusion.

The causal question asks what physical or computational component produced the immediate output. The answer may be a classifier, rule engine, timing mechanism, sensor-fusion process, operator input, or combination.

The configuration question asks who selected the operating area, evidence level, expiration time, protected categories, lost-communications response, and human-approval requirement. The answer is usually a mission-planning or command function operating within prior policy and system limitations.

The design question asks who made particular actions possible, impossible, reversible, persistent, interruptible, or dependent on human approval. The answer includes system architects, software and hardware developers, safety engineers, human-factors specialists, and procuring organizations.

The epistemic question asks who determined that the model and sensors were reliable enough for the approved context. The answer includes data teams, model developers, independent evaluators, testers, certification or approval bodies, and the officials who accepted residual risk.

The legal question asks which people and institutions were responsible for ensuring that development, deployment, and use complied with applicable law. Machines do not hold that obligation. Public DoD policy likewise states that personnel remain responsible for AI development, deployment, and use, and requires appropriate human judgment over uses of force.

The moral and political question asks who chose the distribution of risk: risk to protected persons, risk to operators, risk of missed defense, risk of escalation, and risk created by delay. That decision may be spread across elected officials, policymakers, commanders, procurement executives, contractors, and institutional processes.

The documentary’s answer is therefore not “the machine made the decision” or “a human made every immediate choice.” It is:

The outcome was produced by a chain of human and machine contributions, but responsibility remains human and institutional.

Life-cycle stageHuman judgment transferred into the system
Public policy and lawWhich uses are prohibited, restricted, or subject to review
Doctrine and conceptsWhat role automation may play and what humans must retain
ProcurementRequired speed, operating envelope, interfaces, auditability, fail-safe behavior
DesignAvailable states, sensor dependencies, authorization architecture, deactivation mechanisms
Data and model developmentCategories, labels, training distribution, uncertainty representation
Testing and assuranceWhich errors and environments are considered acceptable or unacceptable
Mission planningGeographic and temporal bounds, protected objects, confidence, corroboration, fallback
Command authorizationWhether this configuration may be activated for this mission
OperationActivation, supervision, approval, override, abort, or deactivation
Post-mission reviewWhether lessons result in changed policy, design, training, or approval

The system is a participant in the causal chain, not a bearer of policy authority. Calling it the sole decision-maker obscures earlier human choices. Calling it a mere tool can also obscure the genuine importance of machine-speed classification, complex automation, and emergent interactions. The more accurate description is a distributed decision architecture with retained human responsibility.

VERIFIED FACT. DoD’s publicly stated AI principles are Responsible, Equitable, Traceable, Reliable, and Governable. “Traceable” includes transparent and auditable methodologies, data sources, procedures, and documentation. “Reliable” requires explicit uses and testing across the life cycle. “Governable” includes detecting unintended consequences and disengaging or deactivating systems that show unintended behavior.

INTERPRETATION. These principles reinforce the documentary’s central point: responsibility is not confined to the instant when an operator presses a button—or when no operator button exists. It reaches backward to whoever defined the system’s intended function, boundaries, evidence, interfaces, test requirements, and shutdown conditions.

Testing, red-teaming, audit logs, fail-safe design, and public examples

VERIFIED FACT. Public assurance frameworks emphasize that testing must be repeatable, documented, connected to the deployment context, and maintained across the life cycle. NIST’s AI Risk Management Framework calls for measures of uncertainty, performance benchmarks, formal reporting, documented TEVV processes, and continuing reassessment as risks and environments evolve. DoD Directive 3000.09 similarly calls for verification, validation, developmental and operational testing, realistic conditions, and consideration of possible adversarial action.

Testing strategy for the fictional system

Testing should begin with the question, “What assumptions make this configuration safe enough?” It should then deliberately violate those assumptions.

Assurance layerDocumentary demonstration
Component testsShow separate checks of sensors, classifier outputs, communications status, clocks, maps, and state transitions.
Rule testsVerify that prohibited areas, protected categories, expired authorizations, and abort conditions consistently block action.
Model testsEvaluate accuracy, calibration, uncertainty, distribution shifts, and known failure cases without assuming a high score is permission to act.
Integration testsCombine sensor conflict, delay, changing classifications, lost communications, and operator intervention.
Human-factors testsMeasure whether operators understand modes, alerts, uncertainty, expiration, degraded authority, and the consequences of inaction.
Mission-envelope testsTest the approved geography, time, environmental conditions, traffic density, communications availability, and sensor quality.
Boundary testsExamine behavior just inside, at, and just outside every operational limit.
Regression testsRe-run relevant scenarios whenever models, rules, interfaces, sensors, maps, or mission concepts change.
Independent evaluationSeparate, where practicable, the personnel who build the capability from those who judge whether evidence supports deployment.
Post-deployment monitoringDetect performance drift, environmental change, previously unseen interactions, and invalidated assumptions.

Red-teaming strategy

Red-teaming should not be limited to conventional cybersecurity penetration. A governance red team asks how an apparently valid configuration can produce an unacceptable result without any component obviously “breaking.”

The documentary’s red-team room presents scenarios such as:

  • a sensor that remains operational but becomes systematically less reliable;
  • two individually reasonable rules that create an unsafe interaction;
  • a valid but stale protected-object list;
  • an operator overwhelmed by nominally correct review requests;
  • an alert design that causes mode confusion;
  • a model confidence score that is misunderstood as a verified probability;
  • communications that fail during the exact interval assumed to require human approval;
  • an authorization that persists after the underlying classification changes;
  • a mission clock, map, and navigation source that disagree;
  • a system that remains inside its technical envelope while the broader social or operational context has changed.

The red team’s task is not to improve tactical performance. It is to discover where policy intent, implementation, operator understanding, and real behavior diverge.

Audit-log design

A useful audit trail records more than the final action. It should reconstruct the complete authorization history:

Configuration version and approving authority
Mission start, expiration, and assumption-validity status
Sensor inputs or auditable summaries
Sensor health, obstruction, and disagreement states
Model version, output, uncertainty, and applicable limitations
Protected-object and geographic checks
Rule evaluations, including rules that blocked action
Authorization-state transitions
Human review requests, responses, overrides, and timeouts
Communications and navigation status
System-health and endurance state
Action, non-action, abort, return, hold, or termination
Software, data, and mission-package versions
Clock synchronization and event timestamps

VERIFIED FACT. Traceability and auditability are recurring public governance requirements. DoD’s AI principles call for transparent and auditable methods and documentation. GAO’s AI Accountability Framework organizes oversight around governance, data, performance, and monitoring, and provides questions and evidence for auditors and third-party assessors. NATO’s responsible-use principles similarly include responsibility and accountability, explainability and traceability, reliability, and governability.

Audit integrity also requires configuration management. Investigators need to know not merely which nominal policy was approved but precisely which model, protected-object dataset, map, rule package, and interface version operated during the event. Changes should be attributable to an authorized role and linked to the reason, test evidence, and approval for the change.

Fail-safe design

There is no universally safe response to failure.

A system that stops immediately may avoid mistaken action but fail to defend against a real event. A system that continues may preserve the defensive mission but act with stale information. Returning may leave an area unprotected. Holding may consume limited endurance. Seeking human input may be impossible during communications loss. “Fail-safe” is therefore not one command; it is a context-dependent allocation of risk.

The documentary presents five abstract fallback families:

FallbackBenefitRisk
Fail silentPrevents further consequential actionCan leave a real hazard unaddressed
Fail restrictedAllows observation or reversible protectionRestricted behavior may still alter the situation
Fail operationalPreserves mission continuityRelies heavily on prior assumptions
Return or withdrawRestores direct control and preserves the platformMay cross hazards or abandon the objective
Enter safe statePlaces the system in a stable, predictable condition“Safe for the system” may not mean “safe for everyone affected”

A robust design should specify which failure triggers which response, how long the response remains valid, what authority is lost in degraded mode, what evidence is needed for recovery, and whether the system may restore itself or must wait for human reauthorization.

Public, non-classified examples

These examples illustrate bounded or preconfigured automation. They do not imply that civilian spacecraft, aviation systems, vehicles, and autonomous weapons are legally or ethically equivalent.

DoD autonomy policy. The 2023 directive publicly describes boundaries involving intended mission sets, operating environments, target sets, geographic areas, timeframes, operator intentions, non-target risks, system safety, and termination or additional operator input when limits cannot be satisfied. It also requires a new senior review when changes to algorithms, intended missions, operating environments, target sets, or expected countermeasures substantially exceed an earlier approval. This is a direct public illustration of autonomy bounded by prior design and approval decisions.

Phalanx close-in defense. The U.S. Navy publicly describes the MK 15 Phalanx as a point-defense system that automatically performs detection, evaluation, tracking, engagement, and assessment against specified classes of threats. Its public description illustrates machine-speed defensive automation assigned a defined role rather than a system inventing its own strategic purpose. The public fact sheet does not reveal the detailed rules by which those functions are configured or governed.

NASA and JPL spacecraft fault protection. A public JPL paper describes automated responses containing preprogrammed instructions and a general-purpose safe mode that places spacecraft in a lower-power, predictable state for ground-team diagnosis. It also describes command-loss responses and termination of an executing sequence as part of fault protection. This illustrates how loss of communications and hazardous system conditions can lead to behaviors selected long before the failure occurs.

FAA-approved unmanned-aircraft operations. A public 2026 FAA waiver requires a predetermined lost-link route and prior verification of operating areas, geofence boundaries, return-home or landing profiles, abnormal procedures, and emergency procedures. It also requires configured alerts for degraded performance, geofence loss, control-link loss, and automated emergency profiles such as return, hold, or descent. This is an especially clear non-weapon example of the fact that “what the drone does when contact is lost” is a pre-mission governance and configuration choice.

Automated driving domains. NHTSA describes an automated driving system as operating within a defined operational design domain. The domain concept expresses that automation is approved for specified conditions rather than possessing unlimited competence everywhere.

International humanitarian debate. The ICRC proposes limits on target types, duration, geographical scope, scale, situations of use, supervision, intervention, and deactivation. The UN GGE has considered tasks, target profiles, timeframes, movement areas, operating environments, human–machine interaction, legal review, risk mitigation, and rigorous testing. These sources differ in some policy prescriptions, but they converge on the importance of system limits, context, human responsibility, and life-cycle governance.

Website and VR treatment

The website version should permit the viewer to place the three profile timelines side by side. Selecting an event should synchronize all three at the same sensor input, allowing the user to see that the divergent outcomes came from policy rather than different evidence.

The VR version should make the planning room spatially persistent. A viewer who chooses a narrow protected-object rule should later see that policy tile physically follow them into the mission. When an event occurs, the tile should illuminate before the action, conveying that the earlier rule is active even though the policymaker is no longer present.

Color should never be the only indicator of authorization. Each state should have a shape, label, sound, and plain-language explanation:

Authorization stateVisual formSpoken cue
ProhibitedClosed frame“Action prohibited by policy.”
Observe onlyOpen circle“Monitoring authority only.”
Human reviewPaused hourglass“Human authorization required.”
Reversible protectionReturning arrow“Temporary reversible authority.”
PreauthorizedSealed triangle“Prior authorization active.”
AbortBroken path“Mission assumptions invalid. Terminating.”

The final interactive object should be the responsibility lens. When pointed at any action, it displays all contributing human roles. The viewer can reduce the lens to the immediate causal event—“rule evaluation permitted action”—or expand it through the life cycle—“policy authority approved preauthorization; procurement required low latency; designers implemented authorization persistence; testers accepted the validated envelope; mission planners selected the profile; the commander activated it.”

The documentary ends only when the viewer can answer not merely:

“What did the system do?”

but:

“Which human choices made that behavior possible, which evidence activated those choices, who approved the remaining risk, and what should change before the next mission?”