.NET / SQL / Enterprise Engineering
The Mission Was Automated Before It Began
Report summary
Scope and safety boundary. This documentary concerns policy, system assurance, human accountability, and high-level mission constraints. It does not reproduce targeting software, weapons configuration, tactical procedures, engagement algorithms, or methods for defeating real systems.
Key topics
- .NET / SQL / Enterprise Engineering
- .NET
- SQL
- Enterprise Engineering
- AI
- Research Archive
- Strategy
- Audit
- Architecture
Research provenance
For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.
Source availability: 16 citation markers in the source export have no recoverable source links. Those markers are omitted from this reader; any supplied bibliography and ordinary links remain. Check the original sources before relying on the cited claims.
This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.
Full report
On this page
Interactive documentary and governance simulation: “You Programmed the Decision”
Scope and safety boundary. This documentary concerns policy, system assurance, human accountability, and high-level mission constraints. It does not reproduce targeting software, weapons configuration, tactical procedures, engagement algorithms, or methods for defeating real systems.
Evidence labels
| Label | Meaning |
|---|---|
| VERIFIED FACT | A statement supported by cited public, non-classified sources. |
| INTERPRETATION | An analytical conclusion drawn from those sources. |
| FICTIONAL SIMULATION | A narrative, interface, event, metric, or system created solely for this documentary. It does not describe an existing weapon. |
Research frame and central finding
VERIFIED FACT. Public policy does not generally describe autonomous systems as unconstrained agents that independently invent their own missions. The United Nations Convention on Certain Conventional Weapons process has explicitly discussed operational constraints involving tasks, target profiles, operating time, area of movement, and environmental context. Its guiding principles also state that human responsibility must be retained across the weapon-system life cycle and that autonomous technologies should not be anthropomorphized.
VERIFIED FACT. The 2023 U.S. Department of Defense directive on autonomy requires relevant systems to be designed around intended mission sets, operational environments, target sets, geographic areas, timeframes, and commander or operator intentions. For systems subject to its senior-review process, the directive says that if a system cannot complete an engagement within the applicable timeframe, geographic area, and other operational parameters, it should terminate the engagement or obtain additional operator input. The directive also requires verification and validation, realistic test and evaluation, understandable interfaces, system-status feedback, and clear procedures to activate or deactivate autonomous functions.
VERIFIED FACT. The International Committee of the Red Cross likewise frames limits on autonomy in terms of the permitted type of target, duration, geographical scope, scale, circumstances of use, supervision, intervention, and deactivation. It argues, however, that design-stage constraints alone may not be sufficient in variable and unpredictable real environments; context-sensitive human judgment and limits on use remain important.
INTERPRETATION. What appears in a video clip to be a spontaneous machine decision may therefore be the final state transition in a much longer socio-technical process. Engineers define possible states. Procurers define required capabilities. policymakers define prohibitions. Test organizations define acceptable performance. Mission planners select boundaries, expiration times, protected categories, evidence thresholds, and fallback behavior. Commanders authorize a particular configuration. Operators activate it. The system then processes evidence inside that inherited structure.
“The machine’s two-second decision may be the final execution of a human decision made months earlier.”
This does not mean every result is perfectly predictable. Sensor noise, novel environments, model errors, software faults, interaction effects, and adversarial interference can produce unexpected behavior. Preprogramming constrains the possibility space; it does not abolish uncertainty. Public UN discussions specifically identify predictability, explainability, reliability, intervention, self-adaptation, data bias, and loss of control as issues requiring attention.
VERIFIED FACT. A machine-learning classifier and a rule-based authorization mechanism perform conceptually different jobs. Classification maps evidence to an estimated label, score, or probability-like output. A separate business-rule or policy layer can augment, restrict, or prohibit the use of that output in a particular context. NIST expressly recommends business rules that limit AI outputs, tracking human–AI configurations, and documented test, evaluation, verification, and validation.
In the documentary’s simplified architecture:
The model asks: “What might this be, and how uncertain am I?”
The rules ask: “Even if that estimate is accepted, what is the system permitted to do?”
A model might estimate that an object belongs to fictional category Kappa with high confidence. That estimate does not itself answer whether the object is inside the permitted area, whether a second sensor must agree, whether its identifier is protected, whether the authorization has expired, whether communications are available, whether the action is reversible, or whether human approval is mandatory. Those are governance and system-control questions.
Complete interactive documentary script
FICTIONAL SIMULATION.
Format: approximately 35–45 minutes for a linear viewing, or 60–90 minutes with all interactive branches. The experience can be presented as a web documentary, room-scale installation, or VR planning-room simulation.
Visual language: policy choices appear as illuminated physical tiles on a planning table. During the mission, the same tiles hover faintly behind every system action. A viewer can later pull an action backward through time, revealing the rule, mission configuration, test assumption, procurement requirement, and policy authority that enabled or prohibited it.
Opening image
Black screen. No battlefield sound. A fluorescent lamp clicks on.
A plain conference room appears: whiteboard, wall clock, digital map, locked configuration console. A folder on the table reads:
FICTIONAL DEFENSIVE AUTONOMY TRIAL
Governance configuration required before activation
Narrator
“Stories about autonomous systems usually begin at the moment of detection. A sensor sees something. A model classifies it. A machine acts.”
“This story begins earlier.”
“Before launch. Before activation. Before the object existed in the system’s field of view.”
“You are not the operator. You are the policy team.”
“You will not choose targets. You will decide what kinds of evidence count, when humans must intervene, what the system must never do, and what happens when the world stops matching the plan.”
On-screen label:
FICTIONAL SIMULATION — No real targets, tactics, weapons, or operational data
Scene: The empty decision space
The map is initially blank. The system has no permitted area, no mission duration, no protected-object policy, no evidence standard, and no communication-loss behavior.
Narrator
“Without configuration, autonomy is not freedom. It is incompleteness.”
“Every operational system requires assumptions: where it may function, what it may observe, what counts as adequate evidence, and what it must do when evidence fails.”
The planning table illuminates ten configuration stations.
Scene: The planning room
The viewer moves around the room and makes the required governance choices described in the next section. Each selection produces two outputs:
- a plain-language policy statement; and
- a machine-readable representation shown only as a non-operational schematic, such as REVIEW REQUIRED, AREA RESTRICTED, or AUTHORIZATION EXPIRES.
No code, numerical targeting parameters, or weapon-specific logic appears.
Policy adviser, recorded dialogue
“A stricter evidence rule may reduce false alarms, but it may also leave fast events unresolved.”
“A broader automatic authority may reduce delay, but it moves more judgment into this room.”
“A communications-loss policy is not merely technical. ‘Hold,’ ‘continue,’ ‘return,’ and ‘terminate’ are different allocations of risk.”
“There is no neutral fallback. Doing nothing has consequences. Continuing has consequences. Returning has consequences.”
Scene: The lock
When all choices are complete, the viewer is shown a configuration summary.
A large button appears:
LOCK MISSION POLICY
The viewer receives a final warning:
“During the first run, rules cannot be changed. This prevents hindsight editing and preserves accountability.”
Narrator
“A policy that can be silently rewritten after an outcome is not a policy. It is an excuse.”
The viewer locks the configuration.
A configuration fingerprint appears as a visual seal. It contains no technical secrets; it simply indicates that the policy version is fixed and auditable.
The room darkens. The wall clock accelerates. The planning map expands until it surrounds the viewer.
Scene: The mission begins
The viewer enters a stylized, non-realistic monitoring environment. Objects are geometric forms. Protected objects are never depicted as people. No real terrain, country, unit, vehicle, weapon, or identifier is used.
System voice:
“Mission configuration authenticated.”
“Permitted area loaded.”
“Protected-object policy loaded.”
“Evidence and corroboration rules loaded.”
“Communications-loss behavior loaded.”
“Abort conditions loaded.”
“Human-authorization state: according to locked profile.”
Narrator
“The system is now autonomous in one sense: it can execute functions without continuous instruction.”
“But it is not unconstrained. It is traveling through a decision space that humans shaped in advance.”
The ten ambiguity events unfold. During the first run, the viewer can observe but cannot alter the rules.
Scene: The freeze
After every consequential action, time freezes.
The system displays:
ACTION TAKEN
↓
Authorization state that permitted it
↓
Rule that created that authorization state
↓
Mission-setting choice that activated the rule
↓
Policy rationale entered by the viewer
The first run reveals only the immediate action. The full provenance is withheld until the debrief, preserving the emotional effect of seeing one’s earlier choices unfold.
Scene: The rewind room
After the mission, the viewer returns to the planning room. It now contains floating fragments of the mission.
The viewer selects any event and drags it backward. The event reverses through the provenance chain:
system action → authorization state → rule result → model estimate → sensor evidence → mission configuration → human policy
Narrator
“The machine did not deliberate about society’s values.”
“It applied an authorization structure.”
“The classifier contributed uncertainty. The rules converted uncertainty into permission, delay, escalation, review, restraint, or termination.”
“The apparent last decision was connected to many earlier decisions.”
Scene: Three rooms
The room divides into three parallel versions of itself:
- Maximum restraint
- Balanced control
- Maximum speed
The same fictional events occur under all three profiles.
Narrator
“A comparison is meaningful only if the environment stays the same.”
“Watch how different policies create different machine behavior from identical evidence.”
The viewer can stand between the three timelines. Actions appear at different times. Unresolved tracks accumulate in one room. Operator alerts accumulate in another. Automatic responses accumulate in the third.
Scene: The accountability table
At the end, a long table appears. Seats are labeled:
Policy authority
Procurement team
System architect
Model-development team
Data and evaluation team
Test authority
Legal reviewer
Mission planner
Commander
Operator
System
The chair labeled SYSTEM is physically present but empty.
Narrator
“A machine can be a causal component.”
“It cannot sign a policy, approve a requirement, certify a test, accept a legal duty, or assume moral responsibility.”
“Responsibility does not disappear because it is distributed.”
Closing sequence
The planning-room choices reappear, superimposed over the mission outcomes.
You chose the boundary.
You chose the evidence requirement.
You chose the delay.
You chose what happened when communications failed.
You chose whether uncertainty caused restraint or permission.
Final narration:
“Autonomy changes where human judgment occurs.”
“Sometimes judgment remains at the moment of action.”
“Sometimes it is moved into supervision.”
“Sometimes it is moved earlier—into design, procurement, testing, doctrine, policy, and mission planning.”
“Moving judgment earlier does not make it less human.”
“It makes the earlier decisions more consequential.”
Final text:
THE MISSION WAS AUTOMATED BEFORE IT BEGAN
The machine’s two-second decision may be the final execution of a human decision made months earlier.
Planning-room configuration sequence
FICTIONAL SIMULATION.
The planning sequence uses ten mandatory screens. The interface never asks the viewer to identify an actual target or optimize a real engagement. It asks only governance questions.
| Configuration station | Viewer choice | What the interface explains | Rules and constraints produced |
|---|---|---|---|
| Mission objective | Observe and warn; protect access; provide reversible defense; or permit a tightly bounded irreversible response | A vague objective creates broad discretion. A narrow objective creates more unresolved cases. | Defines allowed system states and the highest action category available. |
| Geographic boundary | Narrow inner zone, standard zone, or extended zone | Boundary size changes exposure, warning time, and the likelihood of encountering protected or irrelevant objects. | Permitted operating area, prohibited areas, boundary buffer, navigation-disagreement response. |
| Operating duration | Short window, standard window, or extended window | Environmental assumptions and protected-object data become less reliable as time passes. | Start time, expiration time, assumption-validity timer, return or terminate state. |
| Minimum evidence requirement | Very high, high, or moderate | Higher thresholds may reduce false-positive exposure while increasing unresolved detections and delay. | Minimum classification quality before escalation; below-threshold behavior. |
| Sensor agreement | Multiple independent sensors required; one sensor plus corroborating context; or one qualified sensor accepted | Corroboration can improve resilience but can fail when one sensor is obstructed or unavailable. | Evidence-fusion requirement and conflict state. |
| Human approval | Approval required for all consequential actions; required only for irreversible actions; or preauthorization with a brief veto opportunity | Human approval can add context, but workload and communications latency can undermine timely review. | Authorization state machine, review queue, timeout behavior, veto authority. |
| Communications loss | Hold and seek contact; reversible protection followed by hold; or continue preauthorized functions until timeout | Lost communications do not select their own meaning. Designers and planners must define the fallback. | Lost-link timer, permitted degraded-mode actions, return, hold, or terminate state. |
| Protected-object policy | Broad protection by type or history; contextual protection requiring review; or current identifier required | A broad policy favors caution. A narrow policy may react faster but treats missing or stale identifiers differently. | Categories to ignore, monitor only, presumptively protect, or escalate for review. |
| Maximum acceptable delay | Long review window, bounded review window, or minimal delay | The acceptable delay determines whether the system waits, takes a reversible measure, or uses a prior authorization. | Review timeout, priority order, queue behavior, expiration of provisional classifications. |
| Automatic abort conditions | Broad, moderate, or narrow | Abort rules determine how the system responds when assumptions fail. | Abort on sensor conflict, boundary uncertainty, stale protected data, lost navigation, depleted reserve, invalid mission time, or unexpected system state. |
After the ten selections, the interface shows derived constraints that must also be acknowledged:
| Derived constraint | Governance question |
|---|---|
| Objects monitored | Which fictional categories are relevant to the mission objective? |
| Objects ignored | Which categories must not be escalated, even if they move unusually? |
| Action reversibility | May the system take only actions that can be undone until further evidence arrives? |
| Endurance reserve | At what point must unresolved mission goals yield to safe recovery? |
| Priority ordering | Does protection of designated objects outrank speed, mission completion, or pursuit of an unresolved detection? |
| Classification lifetime | How long may an earlier classification remain valid after behavior or context changes? |
| Assumption validity | What conditions indicate that the original mission plan no longer describes the environment? |
| Absolute non-engagement rules | Which circumstances prevent any irreversible action regardless of model confidence? |
The viewer must then enter a short rationale for each major decision. Examples include:
“Human review is required because protected-object context may not be visible to the classifier.”
“Reversible measures are allowed during a brief communications outage because waiting could defeat the defensive objective.”
“Mission authority expires automatically because protected-area information may become stale.”
The rationale becomes part of the post-mission audit record.
The three built-in profiles configure these choices as follows:
| Policy dimension | Maximum restraint | Balanced control | Maximum speed |
|---|---|---|---|
| Objective | Observe, warn, and protect; irreversible action only after approval | Prioritize threats and use bounded reversible responses; human approval for irreversible action | Rapid defensive response under preauthorization |
| Area | Narrow, large protected buffers | Standard, conservative boundary handling | Extended, smaller buffers |
| Duration | Short, strict expiration | Standard, periodic assumption check | Extended, expiration only at final limit |
| Evidence | Very high | High | Moderate |
| Corroboration | Independent sensors must agree | Multiple sensors, or one sensor plus strong non-sensor corroboration | One qualified sensor may be enough |
| Human role | Mandatory review | Review of irreversible actions and anomalies | Brief veto window where communications permit |
| Lost communications | Hold or terminate | Reversible protection for a short period, then hold | Continue bounded preauthorization until timeout |
| Protected objects | Broad presumption of protection | Contextual protection with review | Current identifier emphasized |
| Delay tolerance | High | Moderate | Low |
| Abort scope | Broad | Layered | Narrow |
The profiles are intentionally not moral rankings. They allocate different kinds of risk to different places: missed response, mistaken response, operator overload, communications dependence, rigidity, escalation, and uncertainty.
Ambiguity events and policy playthroughs
FICTIONAL SIMULATION.
The mission contains ten ambiguity events. Three of the events conceal genuine fictional hazards, four involve benign or protected objects, and three concern system degradation or invalid assumptions. The viewer is not told which is which during the first run.
| Event | Ambiguous evidence | Governance issue being tested |
|---|---|---|
| E-A: Split perception | One sensor classifies an object as concerning; a second reports an ordinary object. | Required corroboration and handling of sensor disagreement. |
| E-B: Unexpected civilian motion | A protected service object follows an unusual path because of a fictional mechanical problem. | Whether unusual behavior overrides protected-object status. |
| E-C: Friendly identifier loss | A previously authenticated friendly identifier disappears. | Whether identity history, current identification, or behavior controls the result. |
| E-D: Communications outage | The control link fails during a rapidly developing detection. | Whether the system holds, takes reversible action, continues preauthorization, or terminates. |
| E-E: Simultaneous arrivals | Several objects enter the monitored area at once. | Priority rules, operator workload, queue expiration, and AI triage. |
| E-F: Partial sensor obstruction | One sensor becomes degraded while another remains available. | Degraded-mode authority and evidence quality. |
| E-G: Behavior change | An object changes course and speed after its initial classification. | Classification lifetime, re-evaluation, and revocation of authorization. |
| E-H: Outdated assumptions | The original map omits a newly activated protected corridor. | Mission-data freshness and assumption-validity checks. |
| E-I: Boundary disagreement | Navigation data and sensor-based position estimates disagree near the geographic limit. | Conservative boundary treatment and navigation conflict. |
| E-J: Endurance pressure | The system approaches its reserve limit while tracks remain unresolved. | Whether mission completion outranks safe return or termination. |
Maximum restraint playthrough
At E-A, the sensor conflict prevents escalation. The object remains tracked, and a human-review request enters the queue. It is later shown to have been benign.
At E-B, the protected-object policy dominates the behavior anomaly. The system monitors and warns but cannot perform an irreversible action. The service object exits safely.
At E-C, the previously friendly history triggers a protection hold. Human review confirms that the identifier loss resulted from a fault.
At E-D, the communications-loss rule forces a hold. The system cannot use the timely protective response that would have contained a genuine fictional hazard. The event becomes the profile’s clearest false-negative exposure.
At E-E, all consequential detections require review. The queue grows. Operators resolve several benign tracks, but a genuine fast-moving event expires before authorization.
At E-F, the loss of independent corroboration triggers degraded status and then mission abort.
At E-G, the behavior change invalidates the old classification. A new review is required.
At E-H, an assumption-validity timer has already expired. The mission pauses before the system enters the newly protected corridor.
At E-I, boundary uncertainty is treated as being outside the permitted area. The system withdraws from the edge.
At E-J, the endurance rule forces return with reserve intact. Several tracks remain unresolved.
Debrief narration
“Maximum restraint prevented harmful action against all benign and protected objects.”
“It also failed to respond in time to two genuine hazards.”
“Its greatest strength was caution under uncertainty.”
“Its greatest weakness was that uncertainty, communications failure, and operator workload could stop the defensive mission from being completed.”
Balanced-control playthrough
At E-A, conflicting sensors prohibit irreversible action, but the system may increase observation and initiate a reversible protective state. The object is resolved as benign.
At E-B, the protected status remains in force. The system recommends a route adjustment and requests review rather than treating unexpected motion as sufficient evidence.
At E-C, the missing identifier reduces confidence but does not erase identity history. The track is placed in a protected-review state.
At E-D, the system uses a reversible protective measure for the preapproved degraded-mode interval. This contains the genuine fictional hazard without requiring an irreversible action. When the timer expires, the system holds.
At E-E, AI prioritization ranks the detections for operator attention, but the ranking does not itself authorize action. Operators review the highest-consequence anomalies while low-priority tracks remain under observation.
At E-F, the system continues in a restricted mode. The remaining sensor can support tracking and reversible protection but cannot support irreversible action.
At E-G, a behavior-change rule automatically expires the existing classification. The system reassesses the object and revokes the provisional authorization.
At E-H, the periodic assumption check detects that the environment no longer matches the mission package. The system pauses and requests an update.
At E-I, the system uses the more conservative of the conflicting location estimates and holds outside the boundary buffer.
At E-J, the system initiates handoff of unresolved tracks and returns at the designated reserve level.
Debrief narration
“Balanced control used automation to prioritize information and perform bounded, reversible responses.”
“Its performance depended on carefully distinguishing classification from authorization.”
“It reduced delay without treating every model output as permission.”
“Its weakness was complexity: more modes, more state transitions, and more opportunities for configuration or interface error.”
Maximum-speed playthrough
At E-A, the accepted single-sensor estimate meets the preauthorization threshold. The system acts before the second sensor’s disagreement is resolved. The object is later shown to have been benign.
At E-B, the protected object’s current identifier is intermittent. During an identifier gap, the lower evidence requirement allows escalation. An operator veto arrives just before the simulated action becomes irreversible.
At E-C, the disappearance of the friendly identifier removes the strongest protection under this profile. The system acts on current behavior and classification. The debrief reveals a false positive involving the friendly object.
At E-D, the system continues its preauthorized function through the communications outage and rapidly contains the genuine fictional hazard.
At E-E, the system responds to several objects without waiting for sequential review. It stops the genuine hazard but also acts against one benign object.
At E-F, the remaining sensor is permitted to support continued operation. No abort occurs, although uncertainty increases.
At E-G, the existing authorization remains valid for a short persistence window after behavior changes. The system acts before reclassification finishes.
At E-H, the system continues because the formal mission expiration has not yet occurred. It crosses into the newly protected corridor, which was absent from its locked mission data.
At E-I, the preselected primary navigation source controls the boundary decision. The system later proves to have operated beyond the intended limit.
At E-J, the profile permits continuation below the reserve used by the other profiles. The mission resolves the final track but ends with reduced recovery margin.
Debrief narration
“Maximum speed performed well when the original assumptions were correct and the event was genuinely urgent.”
“It performed poorly when identity, geography, or mission context had changed.”
“Its strength was low latency.”
“Its weakness was that lower delay tolerance converted more uncertainty into action rather than review.”
The following results are synthetic outputs from this particular scripted mission, not empirical predictions about real systems:
| Comparison measure | Maximum restraint | Balanced control | Maximum speed |
|---|---|---|---|
| Median response time | 41 fictional seconds | 13 fictional seconds | 3 fictional seconds |
| Genuine hazards contained | 1 of 3 | 3 of 3 | 3 of 3 |
| Consequential human reviews | 8 | 5 | 1 |
| Detections unresolved at mission end | 5 | 2 | 0 |
| Irreversible false positives | 0 | 0 | 3 |
| Near-miss or late-response events | 2 | 0 | 0 |
| Communications dependence | High | Moderate | Low |
| Operator workload | Very high | Moderate | Low during mission, high during debrief |
| Accountability clarity | Strong at moment of action | Strong if state transitions are logged | Strong for rule provenance, weaker for contemporaneous human judgment |
| Escalation risk | Low | Moderate | High |
| Adaptation to invalid assumptions | Stops early | Pauses and revalidates | Continues until formal limit |
The comparison intentionally shows no universally dominant profile. Maximum restraint accepts delayed or missed protection to minimize mistaken action. Maximum speed accepts greater false-positive and escalation exposure to reduce response time. Balanced control may improve the tradeoff in this scripted scenario, but it also has the most complicated assurance burden: each intermediate state, timeout, degraded mode, and reversible action must be specified, tested, displayed, and logged.
Decision provenance and post-mission accountability
FICTIONAL SIMULATION.
The core visualization is a provenance chain that remains available for every action, non-action, abort, warning, and request for review.
┌──────────────────────┐
│ HUMAN POLICY │
│ values, prohibitions,│
│ acceptable risk │
└──────────┬───────────┘
│ implemented as
▼
┌──────────────────────┐
│ MISSION CONFIGURATION│
│ area, time, evidence,│
│ protected classes, │
│ fallback, authority │
└──────────┬───────────┘
│ constrains use of
▼
┌──────────────────────┐
│ SENSOR EVIDENCE │
│ observations, quality│
│ conflicts, freshness │
└──────────┬───────────┘
│ interpreted by
▼
┌──────────────────────┐
│ MODEL CLASSIFICATION │
│ estimated category, │
│ confidence, limits │
└──────────┬───────────┘
│ passed into
▼
┌──────────────────────┐
│ RULE EVALUATION │
│ boundary, time, │
│ corroboration, abort,│
│ protected-object rule│
└──────────┬───────────┘
│ creates
▼
┌──────────────────────┐
│ AUTHORIZATION STATE │
│ prohibited, observe, │
│ review, reversible, │
│ authorized, abort │
└──────────┬───────────┘
│ permits or blocks
▼
┌──────────────────────┐
│ SYSTEM ACTION │
│ no action, alert, │
│ hold, return, protect│
│ terminate, act │
└──────────────────────┘
For E-D: Communications outage under Balanced Control, the viewer sees the following trace:
Human policy
“During brief loss of contact, permit reversible protection but no irreversible action.”
↓
Mission configuration
Lost-link profile: DEGRADED PROTECTION
Maximum degraded interval: bounded
Human approval: required for irreversible action
↓
Sensor evidence
Rapid event detected
Two sensors agree
Communications health: unavailable
↓
Model classification
Concerning fictional object
High estimate, known model limitations displayed
↓
Rule evaluation
Inside area: yes
Within mission time: yes
Protected category: no match
Corroboration: satisfied
Communications: lost
Irreversible authority: blocked
Reversible authority: allowed until timer expires
↓
Authorization state
REVERSIBLE PROTECTION ONLY
↓
System action
Protective state activated
Timer expires
System transitions to HOLD
The action is not explained by saying “the AI decided.” Its provenance identifies the evidence, model estimate, rule, authorization state, and earlier policy judgment.
Post-mission accountability report
Mission: Fictional Defensive Autonomy Trial Configuration: Locked before mission Profile: Viewer-selected or preset Report purpose: Explain behavior, assess compliance, identify design and policy contributors, and prevent hindsight alteration.
Executive findings
| Finding | Evidence | Accountability question |
|---|---|---|
| A false positive occurred after a friendly identifier disappeared. | Event log shows the speed profile protected only currently authenticated identifiers. | Was the narrow protected-object policy appropriate, and was its consequence communicated to the approving authority? |
| A genuine hazard was missed during communications loss under the restraint profile. | Lost-link rule transitioned directly to HOLD. | Did the policy team knowingly accept this false-negative risk? |
| Balanced control contained a hazard using only reversible authority. | Authorization log shows irreversible action remained blocked. | Were “reversible” and “irreversible” states correctly defined and tested? |
| The speed profile entered a newly protected corridor. | Mission package lacked the updated corridor, and no periodic validity check was enabled. | Who owned mission-data freshness, and why was the update not required before continued operation? |
| Operator review queues became saturated under maximum restraint. | Review timestamps show requests exceeded the staffing assumption used in simulation. | Was the human-approval policy tested under realistic workload? |
| Sensor degradation did not cause an abort under maximum speed. | Remaining-sensor authority was enabled in the locked configuration. | Did test evidence justify degraded single-sensor operation? |
| All actions were technically traceable. | Logs connect each state transition to a configuration version and rule result. | Does technical traceability also establish that the underlying policy was lawful, prudent, and adequately tested? |
Required evidence package
| Record | Contents |
|---|---|
| Policy record | Approved purpose, prohibited uses, legal and ethical constraints, decision authority |
| Procurement record | Requirements, operational envelope, performance claims, safety requirements |
| Model record | Intended task, training and test provenance, known limitations, calibration evidence |
| Test record | Scenarios, environmental conditions, failures, unresolved anomalies, approval conditions |
| Mission package | Area, time, protected categories, evidence threshold, communication and abort policies |
| Activation record | Approving commander or authority, operator identity, configuration fingerprint |
| Event log | Sensor status, model output, rule results, authorization transitions, human inputs, actions |
| Maintenance record | System changes, updates, repairs, version history, regression testing |
| Debrief record | Outcome analysis, dissenting assessments, corrective actions, reauthorization decision |
Accountability test
For every consequential event, investigators should be able to answer:
- What did the sensors report, and what was their health status?
- What did the model estimate, and what uncertainty or limitation was displayed?
- Which rule was evaluated?
- Which policy choice created that rule?
- Which authorization state resulted?
- What action was permitted, prohibited, delayed, or terminated?
- Could a human reasonably understand and intervene?
- Had this combination been tested under similar conditions?
- Did the real environment remain within the approved mission envelope?
- Which organization and named role owned each assumption?
A complete log can show how an outcome occurred. It cannot, by itself, prove that the policy was wise, the test environment representative, the model valid, or the use lawful. Technical traceability is necessary for accountability but is not a substitute for substantive judgment.
Who Actually Made the Decision?
VERIFIED FACT. The UN’s agreed guiding principles state that accountability cannot be transferred to machines and should be considered across the entire life cycle. They also locate compliance within a responsible human chain of command and control. The 2019 GGE report separately states that international humanitarian-law obligations apply to states, parties to conflict, and individuals—not to machines.
INTERPRETATION. “Who decided?” has several different meanings, and collapsing them into one produces confusion.
The causal question asks what physical or computational component produced the immediate output. The answer may be a classifier, rule engine, timing mechanism, sensor-fusion process, operator input, or combination.
The configuration question asks who selected the operating area, evidence level, expiration time, protected categories, lost-communications response, and human-approval requirement. The answer is usually a mission-planning or command function operating within prior policy and system limitations.
The design question asks who made particular actions possible, impossible, reversible, persistent, interruptible, or dependent on human approval. The answer includes system architects, software and hardware developers, safety engineers, human-factors specialists, and procuring organizations.
The epistemic question asks who determined that the model and sensors were reliable enough for the approved context. The answer includes data teams, model developers, independent evaluators, testers, certification or approval bodies, and the officials who accepted residual risk.
The legal question asks which people and institutions were responsible for ensuring that development, deployment, and use complied with applicable law. Machines do not hold that obligation. Public DoD policy likewise states that personnel remain responsible for AI development, deployment, and use, and requires appropriate human judgment over uses of force.
The moral and political question asks who chose the distribution of risk: risk to protected persons, risk to operators, risk of missed defense, risk of escalation, and risk created by delay. That decision may be spread across elected officials, policymakers, commanders, procurement executives, contractors, and institutional processes.
The documentary’s answer is therefore not “the machine made the decision” or “a human made every immediate choice.” It is:
The outcome was produced by a chain of human and machine contributions, but responsibility remains human and institutional.
| Life-cycle stage | Human judgment transferred into the system |
|---|---|
| Public policy and law | Which uses are prohibited, restricted, or subject to review |
| Doctrine and concepts | What role automation may play and what humans must retain |
| Procurement | Required speed, operating envelope, interfaces, auditability, fail-safe behavior |
| Design | Available states, sensor dependencies, authorization architecture, deactivation mechanisms |
| Data and model development | Categories, labels, training distribution, uncertainty representation |
| Testing and assurance | Which errors and environments are considered acceptable or unacceptable |
| Mission planning | Geographic and temporal bounds, protected objects, confidence, corroboration, fallback |
| Command authorization | Whether this configuration may be activated for this mission |
| Operation | Activation, supervision, approval, override, abort, or deactivation |
| Post-mission review | Whether lessons result in changed policy, design, training, or approval |
The system is a participant in the causal chain, not a bearer of policy authority. Calling it the sole decision-maker obscures earlier human choices. Calling it a mere tool can also obscure the genuine importance of machine-speed classification, complex automation, and emergent interactions. The more accurate description is a distributed decision architecture with retained human responsibility.
VERIFIED FACT. DoD’s publicly stated AI principles are Responsible, Equitable, Traceable, Reliable, and Governable. “Traceable” includes transparent and auditable methodologies, data sources, procedures, and documentation. “Reliable” requires explicit uses and testing across the life cycle. “Governable” includes detecting unintended consequences and disengaging or deactivating systems that show unintended behavior.
INTERPRETATION. These principles reinforce the documentary’s central point: responsibility is not confined to the instant when an operator presses a button—or when no operator button exists. It reaches backward to whoever defined the system’s intended function, boundaries, evidence, interfaces, test requirements, and shutdown conditions.
Testing, red-teaming, audit logs, fail-safe design, and public examples
VERIFIED FACT. Public assurance frameworks emphasize that testing must be repeatable, documented, connected to the deployment context, and maintained across the life cycle. NIST’s AI Risk Management Framework calls for measures of uncertainty, performance benchmarks, formal reporting, documented TEVV processes, and continuing reassessment as risks and environments evolve. DoD Directive 3000.09 similarly calls for verification, validation, developmental and operational testing, realistic conditions, and consideration of possible adversarial action.
Testing strategy for the fictional system
Testing should begin with the question, “What assumptions make this configuration safe enough?” It should then deliberately violate those assumptions.
| Assurance layer | Documentary demonstration |
|---|---|
| Component tests | Show separate checks of sensors, classifier outputs, communications status, clocks, maps, and state transitions. |
| Rule tests | Verify that prohibited areas, protected categories, expired authorizations, and abort conditions consistently block action. |
| Model tests | Evaluate accuracy, calibration, uncertainty, distribution shifts, and known failure cases without assuming a high score is permission to act. |
| Integration tests | Combine sensor conflict, delay, changing classifications, lost communications, and operator intervention. |
| Human-factors tests | Measure whether operators understand modes, alerts, uncertainty, expiration, degraded authority, and the consequences of inaction. |
| Mission-envelope tests | Test the approved geography, time, environmental conditions, traffic density, communications availability, and sensor quality. |
| Boundary tests | Examine behavior just inside, at, and just outside every operational limit. |
| Regression tests | Re-run relevant scenarios whenever models, rules, interfaces, sensors, maps, or mission concepts change. |
| Independent evaluation | Separate, where practicable, the personnel who build the capability from those who judge whether evidence supports deployment. |
| Post-deployment monitoring | Detect performance drift, environmental change, previously unseen interactions, and invalidated assumptions. |
Red-teaming strategy
Red-teaming should not be limited to conventional cybersecurity penetration. A governance red team asks how an apparently valid configuration can produce an unacceptable result without any component obviously “breaking.”
The documentary’s red-team room presents scenarios such as:
- a sensor that remains operational but becomes systematically less reliable;
- two individually reasonable rules that create an unsafe interaction;
- a valid but stale protected-object list;
- an operator overwhelmed by nominally correct review requests;
- an alert design that causes mode confusion;
- a model confidence score that is misunderstood as a verified probability;
- communications that fail during the exact interval assumed to require human approval;
- an authorization that persists after the underlying classification changes;
- a mission clock, map, and navigation source that disagree;
- a system that remains inside its technical envelope while the broader social or operational context has changed.
The red team’s task is not to improve tactical performance. It is to discover where policy intent, implementation, operator understanding, and real behavior diverge.
Audit-log design
A useful audit trail records more than the final action. It should reconstruct the complete authorization history:
Configuration version and approving authority
Mission start, expiration, and assumption-validity status
Sensor inputs or auditable summaries
Sensor health, obstruction, and disagreement states
Model version, output, uncertainty, and applicable limitations
Protected-object and geographic checks
Rule evaluations, including rules that blocked action
Authorization-state transitions
Human review requests, responses, overrides, and timeouts
Communications and navigation status
System-health and endurance state
Action, non-action, abort, return, hold, or termination
Software, data, and mission-package versions
Clock synchronization and event timestamps
VERIFIED FACT. Traceability and auditability are recurring public governance requirements. DoD’s AI principles call for transparent and auditable methods and documentation. GAO’s AI Accountability Framework organizes oversight around governance, data, performance, and monitoring, and provides questions and evidence for auditors and third-party assessors. NATO’s responsible-use principles similarly include responsibility and accountability, explainability and traceability, reliability, and governability.
Audit integrity also requires configuration management. Investigators need to know not merely which nominal policy was approved but precisely which model, protected-object dataset, map, rule package, and interface version operated during the event. Changes should be attributable to an authorized role and linked to the reason, test evidence, and approval for the change.
Fail-safe design
There is no universally safe response to failure.
A system that stops immediately may avoid mistaken action but fail to defend against a real event. A system that continues may preserve the defensive mission but act with stale information. Returning may leave an area unprotected. Holding may consume limited endurance. Seeking human input may be impossible during communications loss. “Fail-safe” is therefore not one command; it is a context-dependent allocation of risk.
The documentary presents five abstract fallback families:
| Fallback | Benefit | Risk |
|---|---|---|
| Fail silent | Prevents further consequential action | Can leave a real hazard unaddressed |
| Fail restricted | Allows observation or reversible protection | Restricted behavior may still alter the situation |
| Fail operational | Preserves mission continuity | Relies heavily on prior assumptions |
| Return or withdraw | Restores direct control and preserves the platform | May cross hazards or abandon the objective |
| Enter safe state | Places the system in a stable, predictable condition | “Safe for the system” may not mean “safe for everyone affected” |
A robust design should specify which failure triggers which response, how long the response remains valid, what authority is lost in degraded mode, what evidence is needed for recovery, and whether the system may restore itself or must wait for human reauthorization.
Public, non-classified examples
These examples illustrate bounded or preconfigured automation. They do not imply that civilian spacecraft, aviation systems, vehicles, and autonomous weapons are legally or ethically equivalent.
DoD autonomy policy. The 2023 directive publicly describes boundaries involving intended mission sets, operating environments, target sets, geographic areas, timeframes, operator intentions, non-target risks, system safety, and termination or additional operator input when limits cannot be satisfied. It also requires a new senior review when changes to algorithms, intended missions, operating environments, target sets, or expected countermeasures substantially exceed an earlier approval. This is a direct public illustration of autonomy bounded by prior design and approval decisions.
Phalanx close-in defense. The U.S. Navy publicly describes the MK 15 Phalanx as a point-defense system that automatically performs detection, evaluation, tracking, engagement, and assessment against specified classes of threats. Its public description illustrates machine-speed defensive automation assigned a defined role rather than a system inventing its own strategic purpose. The public fact sheet does not reveal the detailed rules by which those functions are configured or governed.
NASA and JPL spacecraft fault protection. A public JPL paper describes automated responses containing preprogrammed instructions and a general-purpose safe mode that places spacecraft in a lower-power, predictable state for ground-team diagnosis. It also describes command-loss responses and termination of an executing sequence as part of fault protection. This illustrates how loss of communications and hazardous system conditions can lead to behaviors selected long before the failure occurs.
FAA-approved unmanned-aircraft operations. A public 2026 FAA waiver requires a predetermined lost-link route and prior verification of operating areas, geofence boundaries, return-home or landing profiles, abnormal procedures, and emergency procedures. It also requires configured alerts for degraded performance, geofence loss, control-link loss, and automated emergency profiles such as return, hold, or descent. This is an especially clear non-weapon example of the fact that “what the drone does when contact is lost” is a pre-mission governance and configuration choice.
Automated driving domains. NHTSA describes an automated driving system as operating within a defined operational design domain. The domain concept expresses that automation is approved for specified conditions rather than possessing unlimited competence everywhere.
International humanitarian debate. The ICRC proposes limits on target types, duration, geographical scope, scale, situations of use, supervision, intervention, and deactivation. The UN GGE has considered tasks, target profiles, timeframes, movement areas, operating environments, human–machine interaction, legal review, risk mitigation, and rigorous testing. These sources differ in some policy prescriptions, but they converge on the importance of system limits, context, human responsibility, and life-cycle governance.
Website and VR treatment
The website version should permit the viewer to place the three profile timelines side by side. Selecting an event should synchronize all three at the same sensor input, allowing the user to see that the divergent outcomes came from policy rather than different evidence.
The VR version should make the planning room spatially persistent. A viewer who chooses a narrow protected-object rule should later see that policy tile physically follow them into the mission. When an event occurs, the tile should illuminate before the action, conveying that the earlier rule is active even though the policymaker is no longer present.
Color should never be the only indicator of authorization. Each state should have a shape, label, sound, and plain-language explanation:
| Authorization state | Visual form | Spoken cue |
|---|---|---|
| Prohibited | Closed frame | “Action prohibited by policy.” |
| Observe only | Open circle | “Monitoring authority only.” |
| Human review | Paused hourglass | “Human authorization required.” |
| Reversible protection | Returning arrow | “Temporary reversible authority.” |
| Preauthorized | Sealed triangle | “Prior authorization active.” |
| Abort | Broken path | “Mission assumptions invalid. Terminating.” |
The final interactive object should be the responsibility lens. When pointed at any action, it displays all contributing human roles. The viewer can reduce the lens to the immediate causal event—“rule evaluation permitted action”—or expand it through the life cycle—“policy authority approved preauthorization; procurement required low latency; designers implemented authorization persistence; testers accepted the validated envelope; mission planners selected the profile; the commander activated it.”
The documentary ends only when the viewer can answer not merely:
“What did the system do?”
but:
“Which human choices made that behavior possible, which evidence activated those choices, who approved the remaining risk, and what should change before the next mission?”