Autonomy 2040

Autonomy 2040

Who Decides to Strike?

Three regulatory choices. Three futures for war.

A scenario study of military AI, law and accountability · 2022–2040

AI is entering the systems through which militaries see, interpret and act upon the world. It can help identify objects, rank possible targets, navigate without communication and coordinate machines at speeds no human operator can match.

The central question is no longer whether AI will be used in warfare. It is what decisions it will be allowed to make, under what conditions, and who will remain responsible when those decisions cause harm.

This is not a prediction of one inevitable future. It is a structured examination of how three different regulatory choices could change warfare between 2026 and 2040.

Why 2040? The immediate risks in this scenario arrive during the late 2020s. The year 2040 shows what happens after early choices become infrastructure, doctrine and precedent.

The argument in 90 seconds

  1. Autonomy is already moving onboard, because communications are jammed.
  2. AI can shape a strike without firing the weapon itself.
  3. A nominal human approval is not necessarily a real decision.
  4. Too little regulation and badly designed regulation create different failure modes.
  5. Testable, risk-based accountability is part of usable defence capability — and lawfulness and civilian protection are requirements in their own right.
What is this?

A public-interest scenario exercise, not a prediction. One shared, evidence-based timeline (2022–2026) divides in 2027 into three regulatory futures, each followed to 2040. Every major entry is labelled by its evidential status, and the assumptions are set out on the Method page so that lawyers, policymakers and researchers can inspect — and contest — them.

Why use scenarios?

Individual policy proposals can sound attractive in isolation. Following them through procurement, conflict, institutional incentives and crisis conditions can reveal consequences that static analysis misses. A scenario is a way of asking a policy the question a crisis will eventually ask it — on paper, first.

What counts as autonomy?

Autonomy is not one thing. It helps to distinguish:

  • autonomous navigation;
  • automatic tracking of a human-selected target;
  • AI-generated target recommendations;
  • autonomous target selection;
  • autonomous application of force;
  • strategic planning across multiple systems.

A drone that navigates without a signal is not the same as a system that chooses whom to attack. The ICRC's definition of autonomous weapons turns on the latter: selecting and engaging targets without further human intervention.

The question is not simply whether a human pressed the button. It is whether that human had the information, time and authority required to make a real decision.

How certain is any of this?

Every major entry states its basis— one of four labels — plus, where a claim rests on reporting, its corroboration status, and the scenario's causal confidence:

documented
An event or policy supported by reliable evidence.
reported
A credible claim that has not been fully or independently verified.
projected
A future development extrapolated from identified incentives and capabilities.
speculative
A lower-confidence possibility included because its consequences could be extreme.

Reported material additionally carries a corroboration status: corroborated (multiple independent sources), partly corroborated, or uncorroborated (a single source, like the 2024 Ukraine claim). A projection is a causal proposition, not evidence. All numerical indicators in the monitor are illustrative scenario indices, not forecast probabilities. Factual claims have a July 2026 evidence cut-off.

The scenario, 2022–2040

Phase I · 2022–2026The transition begins

Shared history · 2022–2024

The battlefield becomes a technology laboratory

documented · corroborated · high confidence
Basis: An event or policy supported by reliable evidence. Corroboration: Supported by more than one independent source. Confidence: the scenario treats this causal step as high-confidence.

Jamming breaks remote control in Ukraine. Militaries answer with fibre-optic tethers, frequency-hopping — and onboard autonomy. One of those answers creates a pathway.

Electronic warfare defines the drone war. When jamming severs the link between operator and aircraft, both sides respond several ways at once: fibre-optic cables that cannot be jammed, frequency-hopping radios, and software that lets the drone finish its task alone.

Autonomy is not the only answer to jamming — but it is the answer that scales. A tether reaches ten kilometres; onboard navigation, tracking and image recognition reach anywhere. Ukrainian drones learn to lock onto a pre-identified target in the final phase of flight and fly the last, jammed kilometre themselves.

That is terminal guidance, not target selection. But the same onboard components — camera, classifier, guidance loop — are a reusable pathway toward classification and engagement. The technical groundwork travels further than the doctrine that ordered it.

The military AI kill chainFIND
sensor fusion, object recognition
FIX
geolocation, data links
TRACK
autonomous tracking
TARGET
ranking, recommendation
ENGAGE
terminal guidance, navigation
ASSESS
battle-damage analysis
── AI functions entering each stage
The kill chain: AI now assists every stage, not only the weapon.Illustration / scenario index — not measured data.Six connected stages of the military kill chain — find, fix, track, target, engage, assess — each annotated with the AI functions entering that stage.

Concepts

Automation vs autonomy

Automation follows fixed rules; autonomy interprets a situation and chooses among actions.

An automated system does the same thing every time: a fuse detonates at a set altitude, an autopilot holds a heading. An autonomous function decides — it classifies what its sensors see and selects a response within some envelope.

The distinction matters because law and safety cases attach to decisions. A drone that flies itself home is automating navigation. A drone that picks which object to hit is making the decision the law cares most about. Most real systems mix both, which is why regulating 'military AI' as one category fails.

SOURCES [10] Defining Autonomy: Why Software, Not Drones, Will Decide the Next War, Center for Strategic and International Studies (CSIS)[5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross

The kill chain

The sequence from finding a target to assessing a strike — AI now assists every link, not just the weapon.

Militaries describe targeting as find, fix, track, target, engage, assess. AI enters each stage separately: sensor fusion and object recognition at find; geolocation at fix; autonomous tracking; ranking and recommendation at target; terminal guidance at engage; battle-damage analysis at assess.

A system can therefore shape a lethal outcome without firing anything — by deciding what a human sees, in what order, with what confidence attached. Regulation aimed only at the trigger misses most of the chain.

SOURCES [10] Defining Autonomy: Why Software, Not Drones, Will Decide the Next War, Center for Strategic and International Studies (CSIS)

Why might this follow?
  1. Jamming breaks continuous remote control
  2. Tethers, new radios and autonomy all compete as fixes
  3. Onboard guidance scales where tethers cannot
  4. The same components enable classification and engagement later
What could prevent or alter this?

Hard software limits on link-loss behaviour — return, loiter, abort — preserve the human decision at the cost of some missions. Some forces accept that cost; others quietly don't.

SOURCES [9] Technological Evolution on the Battlefield, Center for Strategic and International Studies (CSIS)[10] Defining Autonomy: Why Software, Not Drones, Will Decide the Next War, Center for Strategic and International Studies (CSIS)

Shared history · 2024, reported 2026

A threshold is claimed, not established

reported · uncorroborated · low confidence
Basis: A credible claim that has not been fully or independently verified. Corroboration: Rests on a single source with no independent public corroboration. Confidence: the scenario treats this causal step as low-confidence.

A single-source report says ten fully autonomous drones killed people in a 2024 test. What it proves is uncertainty — and how thresholds will actually arrive.

Single-source report; no independent public corroboration.

In June 2026, New Scientist reported a claim by a senior Ukrainian defence-industry figure: a one-off 2024 test in which drones were set to destroy anything in a given area, with confirmed casualties. 'We tried it. It's a test. We never implemented it.'

Treat this as an exercise in evidentiary uncertainty. The public record contains no operational logs, no official confirmation, no geolocated footage, no independent investigation. It should not be called the first autonomous killing, or fully autonomous, except as an attributed claim. Independent analysis in the same period finds no fielded end-to-end autonomy anywhere.

The claim still matters. It shows how thresholds will arrive if they arrive: not announced, but disclosed later, by one voice, unverifiable — and instantly assumed possible by every planner who hears it.

Concepts

Navigation vs target selection

Flying without a signal is not the same as choosing whom to attack.

Autonomous navigation, station-keeping and terminal guidance onto a target a human already chose are established, widely fielded functions. Autonomous target selection — the machine deciding what to engage — is the threshold the ICRC's definition turns on.

The technical pathway is continuous (the same onboard vision that tracks a chosen target can be retasked to find one), but the legal and moral line is not. Keeping the two separate in language is a precondition for regulating either.

SOURCES [5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross[9] Technological Evolution on the Battlefield, Center for Strategic and International Studies (CSIS)

Why might this follow?
  1. Jamming rewards systems that finish without a link
  2. A one-off test is claimed, then disclosed
  3. Verification is impossible from public evidence
  4. Planners must now assume the capability exists
What could prevent or alter this?

Incident-reporting norms and independent investigation capacity would determine whether such claims get examined or simply absorbed. Neither currently exists for this class of event.

SOURCES [12] First deaths from 'killer drones', New Scientist (International Edition, issue of 20 June 2026)[10] Defining Autonomy: Why Software, Not Drones, Will Decide the Next War, Center for Strategic and International Studies (CSIS)

Shared history · 2025

Britain commits to autonomy and machine-speed integration

documented · corroborated · high confidence
Basis: An event or policy supported by reliable evidence. Corroboration: Supported by more than one independent source. Confidence: the scenario treats this causal step as high-confidence.

The Strategic Defence Review orders uncrewed systems at scale and a £1bn Digital Targeting Web. What kind of 'autonomous' matters more than the word.

The SDR 2025 commits Britain to drone swarms, uncrewed vessels and aircraft, faster procurement, and a Digital Targeting Web — due 2027 — linking sensors, deciders and weapons at what the review calls machine speed.

Precision matters here. The review's 'autonomous' spans logistics, navigation, surveillance, mine countermeasures and defensive interception. It does not announce autonomous engagement, and stated UK policy rules out weapons that identify, select and attack targets without context-appropriate human involvement.

What the SDR does commit to is the infrastructure AI-assisted targeting flows through: the sensors, the links, the recommendation engines, the tempo. Policy now sets the direction the battlefield had already taken. Unset, as yet: who tests the claims, what records survive, who answers when it goes wrong.

Concepts

AI decision support

Software that recommends targets or actions while a human formally decides.

Decision-support systems rank candidate targets, propose weapon pairings and estimate collateral risk. They are usually procured outside weapon-review regimes because they do not fire.

Their influence is real: they set the options, the order, the framing and the tempo of human choices. Whether that influence is scrutinised anywhere is one of the scenario's central questions.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Why might this follow?
  1. The threat environment rewards speed and mass
  2. Government commits to uncrewed systems and integration
  3. Targeting infrastructure precedes its assurance regime
What could prevent or alter this?

The Targeting Web's assurance regime is still designable in 2026 — logging, human-control testing and review triggers could be procurement conditions rather than retrofits.

SOURCES [1] The Strategic Defence Review 2025 — Making Britain Safer: secure at home, strong abroad, UK Government[2] Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence, UK Ministry of Defence

Shared history · 2025–2026

Decision cycles compress now

documented · partly corroborated · high confidence
Basis: An event or policy supported by reliable evidence. Corroboration: Core claims are supported by independent sources; at least one supporting citation is still pending verification. Confidence: the scenario treats this causal step as high-confidence.

Target recognition and navigation assistance are already routine. Detection-to-destruction can take thirty seconds. This is the present, not a forecast.

By 2026 the acceleration is documented, not hypothetical. Automatic target recognition and navigation assistance are routine in Ukraine. AI-assisted units report detection-to-destruction times of just over thirty seconds. Large orders — reportedly including a German-funded purchase of 50,000 strike drones with terminal autonomous tracking — are putting terminal autonomy into thousands of systems at once.

Two facts must be held together. Adoption is fast and at scale; and fully end-to-end autonomous engagement remains unfielded — independent analysis finds no true edge autonomy in operation, and a typical FPV strike still involves three to six people.

The compression is the danger that arrives first. Tens-of-seconds cycles leave no room for deliberation to be added later; whatever review exists must already be inside the loop.

Shared history · Late 2026

The accountability fork

documented · corroborated · high confidence
Basis: An event or policy supported by reliable evidence. Corroboration: Supported by more than one independent source. Confidence: the scenario treats this causal step as high-confidence.

Law, review and policy exist. What is missing is a public, standardised, independently verifiable way to test control claims, software changes and incident evidence.

Britain is not lawless here. International humanitarian law applies to every new weapon. Article 36 reviews are conducted. Stated policy requires context-appropriate human involvement, and NATO principles cover accountability, traceability and governability.

What does not exist is an architecture that makes those commitments checkable from outside a programme: standardised tests of what an operator can actually do at operational tempo; rules for when a software change voids an approval; evidence that must survive an incident; anyone independent and technically equipped to verify any of it.

With systems arriving in thousands and decision cycles in seconds, the gap between commitment and verification is now the decision. In 2027, Britain chooses how — or whether — to close it.

The autonomy ladderMACHINEHUMAN1 · INFORMATION ASSISTANCEorganises informationinterprets and decides2 · RECOMMENDATIONproposes targets or actionsinvestigates and decides3 · HUMAN-CONFIRMED ENGAGEMENTtracks a human-selected targetselects and authorises4 · SUPERVISED AUTONOMYselects and attacks within set parametersmonitors, may intervene5 · UNSUPERVISED AUTONOMYselects and attacks without further interventiondefines the mission beforehand6 · STRATEGIC AUTONOMYplans and coordinates across systemssupervises goals, not actions
The autonomy ladder: what the machine does, and what the human still does, at each level. Autonomy is not a single binary condition.Six levels of autonomy from information assistance to strategic autonomy. At each level the machine's role grows and the human's role narrows from deciding each action to supervising goals.

Concepts

Meaningful human control

Not whether a human pressed the button, but whether they had the information, time and authority to decide.

UK policy requires 'context-appropriate human involvement' in weapons that identify, select and attack targets. The unresolved question is what counts.

A testable version asks: what does the operator see, including uncertainty? How long do they genuinely have? Can they reject, delay or suspend? What happens on communications loss? Does approval behaviour collapse under realistic tempo? Those are measurable properties of a system, not assertions in a contract.

SOURCES [2] Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence, UK Ministry of Defence[4] Government response to the House of Lords AI in Weapon Systems Committee report, UK Government

Article 36 review

The legal review every new weapon must pass for compliance with international humanitarian law.

Article 36 of Additional Protocol I obliges states to determine whether a new weapon, means or method of warfare could be used lawfully. The UK conducts these reviews for all new weapons.

The strain point for AI-enabled systems: a review describes the system as tested on a given date. Software that is retrained or updated afterwards may no longer be that system — which is why iterative or triggered re-review matters.

SOURCES [2] Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence, UK Ministry of Defence[4] Government response to the House of Lords AI in Weapon Systems Committee report, UK Government

Traceability & logs

The records that let anyone reconstruct, after the fact, what a system saw, recommended and did.

Reconstruction needs: the model version in use, its inputs, its outputs and confidence, what the operator was shown, and what they did — preserved in tamper-evident form, retained long enough to investigate.

None of this exists by default. Logging is a procurement requirement or it is absent, and its absence is discovered exactly when it matters most: after civilians are dead.

SOURCES [6] Summary of NATO's revised Artificial Intelligence strategy, NATO[7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

The policy fork

2027: Choose a regulatory path

How should the UK govern systems that contribute to the use of lethal force?

Do the systems that find, rank and strike targets need new, testable rules?

“No — existing law, internal review and voluntary commitments are enough. Speed is what matters.”

“Yes — before machine-speed targeting becomes infrastructure.”

What kind of rules?

“One broad approval regime with bright lines. Act now; sort the categories later.”

“Function-by-function tiers with hard prohibitions. Slower to build; harder to break.”

Reading all three paths in sequence. Select a box to follow one.

Path One: Voluntary Assurance. Path Two: Broad Restriction. Path Three: Tiered Assurance. Selecting a path filters the timeline below to that future.

The three paths in detail

Path One

Voluntary Assurance

Primary aim · Field capability at the tempo of the threat, without new institutional friction.Permits · Anything existing law and internal review allow; safeguards are set contract by contract.Prohibits / restricts · Nothing new. Policy language on human involvement is not translated into testable requirements.Principal failure mode · Control claims go untested; records go unkept; after harm, responsibility cannot be assigned in practice.
More detail

Existing law, internal review and voluntary principles, applied contract by contract, while deployment accelerates.

International humanitarian law, Article 36 reviews and the UK's context-appropriate human involvement policy continue as before. Assurance stays inside programmes; suppliers define their own evidence; delivery speed is the measure of success.

Path Two

Broad Restriction

Primary aim · Prevent the worst systems from being deployed at all, and defend a public norm against machine killing.Permits · Systems that pass a single pre-deployment approval regime; narrow emergency defensive exceptions.Prohibits / restricts · Autonomous targeting of people, publicly and clearly — plus, in practice, much that is merely slow to approve.Principal failure mode · One gate for unlike risks: congestion, industry concentration, allied-import dependency and a classified second track.
More detail

Bright-line prohibitions plus a broad licensing regime covering a wide category of military AI.

Legislation drafted after a crisis regulates 'military AI' as one category. The prohibition on autonomous targeting of people is real and durable; the licensing queue around it treats mine-clearers and hunter-killers alike.

Path Three

Tiered Assurance

Primary aim · Match scrutiny to risk: hard limits around lethal force, proportionate speed everywhere else.Permits · Automation by function and envelope — logistics fast, defensive interception bounded, targeting influence deeply tested.Prohibits / restricts · Autonomous selection and engagement of people; systems whose effects cannot be understood, predicted and limited.Principal failure mode · Assurance capacity is the bottleneck: scarce cleared auditors, capture risk, false assurance, costs that favour incumbents.
More detail

Hard prohibitions plus function- and risk-based testing, change control and independent assurance.

Built in stages: interim procurement clauses first, a statutory assurance authority only after pilots and an incident. Human-control claims become acceptance tests at realistic tempo; software changes trigger re-review; responsibility is allocated before deployment.

Follow one path to 2040, or . Your choice is kept in the page address, so a reload or shared link preserves it.

Four shared stress tests

To keep the paths comparable, each faces the same four classes of external pressure. The circumstances differ; the pressure is the same. Chapters responding to one are marked Shared pressure.

  • Tempo · Connected sensors produce more recommendations than humans can investigate.
  • Civilian harm · An AI-supported targeting operation kills civilians and responsibility is contested.
  • Saturation · A mass drone attack exceeds human-by-human defensive authorisation.
  • Escalation · Rival automated systems misread defensive action as preparation for attack.

Shared across all three paths

Who answers when a system contributes to an unlawful strike?

A modern strike is not a button press. It is a chain: procurement choices, software design, certification, mission orders, machine recommendations, a human approval, a weapon release. When civilians die, responsibility is too often passed along that chain until it disappears.

The answer is allocation in advance. Responsibility should follow knowledge, control, contribution and capacity to prevent harm — and a state cannot outsource its responsibility for the use of force to a commander, an operator, a software supplier or a machine. Several forms of responsibility can apply to the same event; none cancels another.

  • State

    Lawful use of force, IHL compliance and reparation

    Was the operation lawful, investigated and remedied?

  • Political / procurement authority

    Authorisation, procurement conditions and oversight

    Was a foreseeable risk accepted without safeguards?

  • Commander

    Mission design, target choice and supervision

    Were precautions and control adequate in context?

  • Operator

    Use within training and approved limits

    Did the operator understand and follow the system's limits?

  • Supplier corporation

    Design claims, testing, disclosure, updates and support

    What did the company know, contribute and have power to prevent?

  • Executives / senior managers

    Decisions within their actual authority

    Did they authorise, conceal or continue conduct despite known risk?

  • Engineers and employees

    Duties proportionate to role, knowledge and control

    Did a person knowingly make a substantial contribution — or report a risk?

  • National assurance authority

    Independent testing, licensing and incident review

    Were claims verified, changes controlled and warnings acted upon?

  • International agency

    Common standards, declarations, inspection and verification

    Can states' and suppliers' compliance claims be independently checked?

Several forms of responsibility can apply to the same event.

Concepts

Can a company commit a war crime?

The International Criminal Court prosecutes people, not corporations — which is why national corporate law carries the weight.

The ICC's jurisdiction covers natural persons only. Corporate criminal liability must ordinarily be pursued under national law, and national systems differ widely.

In practice, cases have more often pursued executives and suppliers as individuals, or reached companies through adjacent offences. That gap is an argument for a clearer national corporate regime; it does not change the legal elements of war crimes themselves — including the required mental element and a sufficient contribution. A bad product, public controversy or a job at a supplier is not enough by itself.

SOURCES [14] Rome Statute of the International Criminal Court, Article 25 (individual criminal responsibility), International Criminal Court[15] International Criminal Court Act 2001, UK Parliament (legislation.gov.uk)[17] Customary IHL Rule 151 — individual criminal responsibility, International Committee of the Red Cross (IHL database)

The knowledge–control–contribution test

Stronger knowledge, closer contribution and greater power to prevent harm justify stronger duties and potential liability.

Liability and regulatory duties should turn on evidence, not on labels like 'dual-use software' or 'advisory output'. The questions are concrete: what did the system do; what did each actor know, and when; what could each actually control; what did each contribute; and what reasonable preventive step was available and not taken.

The test cuts both ways. It stops a supplier hiding behind its customer — and it stops liability landing automatically on a company, or an engineer, whenever a customer commits a crime.

SOURCES [16] Bribery Act 2010, section 7 — failure of commercial organisations to prevent bribery, UK Parliament (legislation.gov.uk)[7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Why certification is not immunity

An approval is evidence of compliance at a moment in time — not a licence for everything done with the system afterwards.

Certification has a scope: a version, an operating envelope, a set of disclosed facts. It cannot excuse concealed limitations, misleading claims, operation outside the approved envelope, or a material update that was never re-reviewed.

Treating a passed test as a shield is the characteristic failure of assurance regimes. Treating it as a floor — with continuing duties of monitoring, disclosure and change control — is what keeps the paper connected to the deployed machine.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Case study · legal hypothetical

If a software supplier is put on notice

documented

In its Q4 2023 Business Update, Palantir reported that, following a 2024 board meeting in Tel Aviv, it 'agreed to a strategic partnership with the Israeli Ministry of Defense to supply Palantir technology to help the country's war effort.'

This case study is hypothetical. The public fact of a defence contract does not establish that Palantir software was used in a particular strike, that the company knew of an unlawful use, or that the company or any employee has committed an offence.

projected · stipulated assumptions

Suppose — as a stipulation, not a fact — that a supplier received credible, specific information that its decision-support software was repeatedly being used outside its approved operating limits, in operations involving possible war crimes.

Suppose it continued to provide operational support and material updates.

What should the law require next?

  1. Credible notice

    A specific allegation, incident report or logged anomaly — not social-media controversy alone.

  2. Preserve and assess

    Protect logs, identify the deployed version, and run a privileged but independently reviewable IHL and technical assessment.

  3. Report and restrict

    Notify the national assurance authority; impose the end-use restrictions that are technically and contractually available.

  4. Escalate

    If misuse persists, use a defined legal process to suspend affected support, updates or licences — without creating an unmanaged battlefield failure.

  5. Investigate and remedy

    Provide protected evidence to competent investigators; apply regulatory, civil or criminal consequences according to proof.

What would investigators need to prove?
  • what functions the software performed in the operation;
  • which version, model, data and configuration were used;
  • what information reached which people, and when;
  • whether warnings showed an unlawful-use risk;
  • what contractual or technical control the company retained;
  • whether continued support made a substantial contribution;
  • whether reasonable preventive steps were available and deliberately not taken.

Ordinary engineers are not criminally liable merely because they worked on a product. Individual liability should follow personal knowledge, conduct and authority — and protected escalation routes should make it safer to report a risk than to sit on one. The burden belongs to institutions, not to whichever employee noticed first.

SOURCES [13] Palantir Q4 2023 Business Update (p. 5), Palantir Technologies

Proposed UK legislation · not current law

Rules that act before the courtroom

War-crime prosecutions are essential — and rare, slow and evidentially demanding. Prevention needs duties and sanctions that operate before, or without, proving a completed international crime.

The following is a proposed UK legislative package. It is a policy recommendation, not a statement of current law.

Working title: the Military AI and Autonomous Systems Accountability Act.

1 · A corporate failure-to-prevent offence

Modelled cautiously on the structure of section 7 of the Bribery Act 2010 — while recognising that serious IHL harm is not bribery. This is not strict liability for war crimes.

  • Applies to suppliers and integrators of defined high-risk military-AI functions.
  • The prosecution must prove a qualifying serious breach connected to the company's system or services, and the statutory connection to persons acting for the company.
  • The company has a defence of reasonable, proportionate prevention procedures.
  • No defence for knowing concealment, falsified evidence or intentional assistance by senior management.
  • Government must publish and maintain prevention guidance.
2 · Individual offences for knowing or reckless assistance

For a director, senior manager or other person acting within their actual role — job title alone proves nothing, and negligent mistakes are not war crimes.

  • Knowingly providing substantial assistance to a prohibited use.
  • Consciously disregarding a clear and substantial risk while authorising continued material support (a demanding, statutorily defined recklessness threshold).
  • Consenting to or conniving in a corporate breach.
  • Deliberately concealing safety limitations, adverse tests or incident evidence.
3 · Regulatory offences below the war-crime threshold

Enforceable without first proving a completed war crime:

  • Unlicensed supply or deployment of a high-risk function.
  • A false or misleading claim about human control, accuracy or IHL compatibility.
  • Failure to disclose a known material limitation.
  • Deploying a material update without required re-certification.
  • Failure to keep or preserve tamper-evident decision and change logs.
  • Falsifying or destroying those records.
  • Failure to report a serious incident, or continuing operation outside the approved envelope.
4 · A continuing supplier due-diligence duty

Proportionate to risk, and continuing for as long as support does:

  • Customer and end-use assessment; a declared operating envelope and prohibited uses.
  • Testing at realistic operational tempo; documented data and model provenance where technically relevant.
  • Monitoring of credible red flags; incident reporting and evidence preservation.
  • Controlled update, rollback and re-certification processes.
  • Contract clauses allowing lawful restriction or suspension; mapped critical subcontractors and third-party models.

The supplier must not become a private foreign-policy authority: restriction and suspension follow pre-agreed contractual triggers and an expedited regulator or court process, with an emergency route where continued support creates an imminent risk of grave harm.

5 · Consequences that change incentives

A graduated set, so enforcement does not depend on the rare criminal verdict:

  • Corrective orders and enhanced monitoring.
  • Turnover-linked civil or administrative fines.
  • Export-licence or product-authorisation suspension.
  • Procurement debarment for serious or repeated breaches.
  • Director disqualification where legally justified.
  • Criminal penalties for proven individual offences.
  • Contributions to victim compensation or reparations mechanisms.
  • Adverse evidential inference where a party with a duty to preserve essential logs cannot account for their absence — subject to fair-trial safeguards.
6 · Whistleblowers and protected evidence
  • Protected reporting routes to cleared regulators and parliamentary or judicial bodies.
  • Anti-retaliation measures with teeth.
  • Secure handling of classified material — protection is not publication.
  • Penalties for suppressing incident evidence.
7 · Procurement as the first enforcement layer

Every high-risk contract gives an independent assurance body access to:

  • System and model cards appropriate to the risk; test plans, adverse results and known limitations.
  • Deployed version identifiers, hashes and change history; data provenance where it affects performance or legality.
  • Human–machine interface and decision-time evidence; incident logs and subcontractor dependencies.
  • Contractual audit rights, incident-notification deadlines, material-change thresholds, remediation, termination and evidence-preservation duties.

Government pays for public assurance through appropriations and a transparent levy on high-risk procurement or exports. Suppliers never select or directly control their own auditor.

These reforms add supplier responsibility; they do not transfer the state's duty to conduct lawful operations or the commander's duty to take feasible precautions.

SOURCES [16] Bribery Act 2010, section 7 — failure of commercial organisations to prevent bribery, UK Parliament (legislation.gov.uk)[15] International Criminal Court Act 2001, UK Parliament (legislation.gov.uk)[7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

What existing cases show

Nuremberg industrialist cases (IG Farben; Krupp)

Historical convictions of individuals

Executives were prosecuted and some convicted as individuals. Commercial role was no shield — but these were not convictions of the corporate entities themselves.

SOURCES [21] The Nuremberg industrialist trials (IG Farben; Krupp) — case summaries, US Holocaust Memorial Museum

Frans van Anraat (Netherlands)

Convicted — complicity in war crimes

A Dutch supplier was convicted of complicity in war crimes for supplying a chemical-weapons precursor with the required knowledge and contribution. A useful supplier analogue; not a software-specific precedent.

SOURCES [22] Public Prosecutor v. Frans van Anraat — judgment summary, ICRC national-practice database / Dutch courts

Lundin (Sweden)

Proceedings ongoing as of July 2026 — no verdict assumed

Swedish proceedings against former executives concern alleged complicity in grave war crimes. As of this site's evidence cut-off, it is not described as a conviction.

SOURCES [23] Sweden v. former Lundin executives (alleged complicity in war crimes), Swedish Prosecution Authority / Stockholm District Court

Lafarge (France / US)

Terrorism-financing conviction; separate complicity litigation

The company's terrorism-financing case is distinct from separate litigation concerning alleged complicity in crimes against humanity. A terrorism-financing conviction is not a corporate war-crimes conviction.

SOURCES [24] Lafarge proceedings — terrorism financing conviction distinguished from separate complicity litigation, French courts / US DOJ

The lesson is not that current law is empty. It is that liability is fragmented, heavily dependent on national rules, and poorly adapted to continuous software support and opaque decision systems.

Proposal · not an existing institution

An IAEA for autonomous weapons — but not a copy of it

The proposed International Autonomous Weapons Agency (IAWA) is a treaty-based verification and assurance body. It is not a world weapons ministry, and it prosecutes no one.

The nuclear analogy earns its place for four things: independent declarations, inspections, common technical methods and confidence-building. Its limits are just as real: software can be copied, updated and hidden; military AI is often dual-use; there is no scarce material like uranium to count; non-state actors matter; and lawfulness usually depends on context of use, not possession.

The IAEA asks whether declared nuclear material and facilities match reality. IAWA would ask what decisions a system is permitted to make, in which conditions, which version was deployed, and whether a harmful use can be reconstructed.

The treaty does not regulate every spreadsheet, logistics model or uncrewed vehicle. It covers defined high-risk functions:

  • selecting or prioritising people or objects for attack;
  • tracking and engaging targets without a new human decision;
  • coordinating lethal swarms;
  • machine-speed coupling of warning, target generation and response;
  • material software updates to certified weapon or targeting functions.
Nine bounded powers
  1. Set minimum international test, logging, change-control and incident-reporting standards.
  2. Maintain a confidential registry of high-risk systems, responsible states, certified functions, approved operating envelopes and deployed version identifiers.
  3. Review whether national weapons-review and assurance processes meet the treaty minimum — without replacing states' Article 36 reviews.
  4. Accredit independent laboratories and national assurance authorities; audit their competence; rotate inspectors.
  5. Conduct agreed routine inspections, and tightly governed challenge inspections where credible evidence suggests an undeclared prohibited system or systematic non-compliance.
  6. Investigate serious incidents with protected access to relevant logs, testing records and version histories.
  7. Support export and end-use verification, emergency incident channels and confidence-building notifications.
  8. Publish sanitised compliance findings while protecting legitimate security information.
  9. Refer evidence to national regulators, prosecutors or competent international courts — any ICC referral concerns natural persons within its jurisdiction, not corporate prosecution.
Governance and payment
  • A conference of states parties adopts the rules; an independent technical secretariat applies them.
  • A multidisciplinary inspectorate: weapons engineers, human-factors experts, IHL lawyers, data and model specialists, military operators, civilian-harm investigators.
  • Inspectors disclose conflicts, rotate assignments, and cannot have recent financial ties to the supplier they inspect.
  • Core inspections are funded by assessed state contributions, so their existence never depends on a supplier or a voluntary donor.
  • A transparent levy on covered procurement or exports may support accredited testing and victim-assistance funds — but a company never hires or chooses its final regulator.
  • Technical findings cannot be vetoed by a permanent Security Council member. Enforcement beyond the agency's own licence and certification consequences remains with treaty parties and competent institutions.
Safeguards against a bad international regime

The nuclear regime also teaches by its failures. The treaty should:

  • avoid permanent classes of 'responsible possessors' and 'prohibited have-nots' — the same functional rules apply to all states;
  • not let secrecy eliminate independent verification;
  • prevent certification from legitimising every use of an approved system;
  • fund assistance so smaller states can comply, rather than making assurance a rich-state monopoly;
  • say openly that non-signatories, proxies and covert software remain serious limits.

IAWA is best understood as a hybrid of nuclear safeguards, chemical-weapons inspection and aviation-style incident investigation — not as a guarantee that abuse disappears.

SOURCES [19] IAEA safeguards and verification, International Atomic Energy Agency[20] IAEA verification and other safeguards activities, International Atomic Energy Agency[18] ICRC position on autonomous weapon systems, International Committee of the Red Cross[2] Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence, UK Ministry of Defence

Legal content on this page is editorial and the reforms described are proposals, not current law. This material requires review by qualified UK and international-criminal-law counsel before publication.

Reading:

Path OneVoluntary Assurance, 2027–2040

Existing law, internal review and voluntary principles, applied contract by contract, while deployment accelerates.

Phase II · 2027–2033The dangerous acceleration

Voluntary Assurance · 2027

Approval becomes throughput

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · TempoConnected sensors produce more recommendations than humans can investigate.

The Targeting Web works: it finds thousands of candidates. Operators approve at interface speed. The record shows a human decision every time.

The Digital Targeting Web delivers on schedule and does what it was built to do: it finds more, faster. A cell that once assessed dozens of candidates a day now receives thousands, each with a confidence score and a recommended weapon.

The arithmetic is unforgiving. Two thousand recommendations across a shift leaves under fifteen seconds each; a real check — imagery, provenance, civilian pattern — takes minutes. The interface adapts to reality, and approval becomes a keystroke.

Nothing here breaks a rule, because no rule specifies what review must involve. Human control is policy; it is nobody's test.

The human remains in the loop, but the loop has become too fast for human judgement.
Operational tempo versus human review time2027202820292030recommendations per shift ↑seconds of review per decision ↓approval becomesa keystroke
Voluntary Assurance path, 2027–2030: machine-generated recommendations rise; seconds of genuine human review per decision fall.Illustration / scenario index — not measured data.A chart with two schematic lines between 2027 and 2030: recommendations per operator shift rising steeply, while review time per decision falls toward a few seconds.

Concepts

Automation bias

The documented human tendency to accept a machine's suggestion rather than check it.

Decades of human-factors research in aviation, medicine and driving show that when a system is right most of the time, people stop verifying it — especially under time pressure, fatigue and workload.

In targeting, automation bias converts 'a human approved each strike' from a safeguard into a formality. It is measurable, which means it can be tested for — or ignored.

Why might this follow?
  1. Connected sensors multiply candidate targets
  2. Review time per decision collapses
  3. Interfaces optimise for throughput
  4. Formal approval survives; deliberation does not
What could prevent or alter this?

Workload caps, minimum review times and uncertainty display would trade tempo for judgement. Under voluntary assurance, nothing requires that trade, and delivery metrics punish it.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)[9] Technological Evolution on the Battlefield, Center for Strategic and International Studies (CSIS)

Voluntary Assurance · 2028

The attribution crisis

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · Civilian harmAn AI-supported targeting operation kills civilians and responsibility is contested.

A recommended strike kills eleven civilians. Four true statements later, no institution can say what the system did or who should answer for it.

A model trained on another conflict classifies civilian vehicles as a re-forming convoy, with high confidence. An operator, ninety approvals into a shift, accepts. Eleven civilians die.

The inquiry hears four accurate statements. Supplier: our output was advisory, our software dual-use. Ministry: a human authorised it. Commander: the system passed review. Operator: the interface never showed that the model was unvalidated for this region. Legal responsibility has not disappeared — the state remains answerable, and negligence claims exist on paper. What is missing is causal traceability: the evidence to establish what the system contributed, and therefore which duty was breached.

No continuing due-diligence duty defines when a supplier on notice must change its support. So, without evidence or duties, institutional accountability defaults to the person with the least actual control: the operator.

The accountability gapSUPPLIER
“We supplied a recommendation, not a strike order.”
GOVERNMENT
“A human authorised the action.”
COMMANDER
“The system was legally reviewed.”
OPERATOR
“The interface concealed uncertainty.”
LETHALDECISIONEveryone participated. Nobody appears wholly responsible.
The accountability gap: every connection to the decision dissolves under examination.Four actors — supplier, government, commander, operator — each connected to a central lethal decision by broken dashed lines. Each actor's statement deflects responsibility elsewhere.

Concepts

Who answers for what

Suppliers, the state, commanders and operators hold different duties — responsibility follows knowledge and control.

Legal responsibility does not vanish when systems are complex; it becomes hard to assign in practice if the evidence and the duties were never allocated.

A workable allocation: suppliers answer for architecture, known limitations and truthful disclosure; the procuring state for testing, integration and deployment conditions; commanders for operational authorisation; operators only for decisions genuinely within their control.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Why might this follow?
  1. Duties were never allocated against evidence
  2. Each actor's account is individually true
  3. Causal reconstruction fails for lack of records
  4. Blame settles on the operator by default
What could prevent or alter this?

Everything the inquiry lacked — logs, disclosed limitations, tested interfaces — was proposed in 2026 as procurement language. It was not required, so it was not priced in.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)[5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross

Voluntary Assurance · 2029

Reconstruction defeated

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.

Parliament asks how the decision was made. Software updates, commercial secrecy and ninety-day log cycles make the question unanswerable.

The follow-up scrutiny runs into the architecture. The model weights are supplier intellectual property. The training data is commercially confidential. Operational logs were overwritten on a ninety-day cycle. The deployed software has been updated eleven times since its legal review, and no rule says which update would have triggered a fresh one.

The select committee records, carefully, that it cannot determine what happened and cannot compel the material that would let it try. Government replies that oversight is robust and detail is classified.

Public confidence settles into cynicism of a specific kind: 'national security' now shields legitimate secrets and institutional failure alike, and from outside no one can tell which is which.

Concepts

Change control

Rules about what happens when the software, model or data behind an approved system changes.

A machine-learning system is not stable the way a rifle is: new training data, retuned thresholds or a swapped model can change behaviour without changing the hardware.

Change control defines which changes trigger re-testing and re-review, who signs off, and what evidence must be produced. Without it, 'the system passed review' describes history, not the thing currently deployed.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Why might this follow?
  1. Evidence is dispersed across contractors and classifications
  2. Updates sever the link between review and reality
  3. Oversight bodies document their own blindness
What could prevent or alter this?

A cleared, technically staffed review body with statutory access could still separate secrets from failures. Creating one now means conceding it was needed before the deaths.

SOURCES [3] Proceed with Caution: Artificial Intelligence in Weapon Systems (HL Paper 16), House of Lords AI in Weapon Systems Committee[4] Government response to the House of Lords AI in Weapon Systems Committee report, UK Government

Voluntary Assurance · 2030–2031

The envelope creeps

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · SaturationA mass drone attack exceeds human-by-human defensive authorisation.

Mass drone attacks force machine-speed defence. The exception is defensible. The boundary of 'incoming threat' is not — and it moves case by case.

Saturation attacks make human-per-engagement defence arithmetic impossible, and autonomous interception within pre-set parameters becomes standard. Against incoming munitions, this is the strongest case for autonomy, and this scenario does not dispute it.

What follows is quieter. Systems authorised against 'incoming threats' meet ambiguous cases — a surveillance drone, a fast boat, a vehicle accelerating at a checkpoint sensor. Each case is resolved operationally, under pressure, by reading the envelope a little wider. No single extension is large enough to trigger political review.

By 2031 the question 'does Britain operate autonomous weapons?' has no clean public answer — and no institution whose job is to keep the answer clean.

Concepts

Defensive autonomy

Machine-speed interception of incoming munitions and drones — the strongest case for autonomy, and the usual door.

A hundred incoming drones cannot be intercepted by a human approving each shot; ship and base defence systems have operated automated modes for decades against clearly inanimate threats.

The governance question is not whether to permit this but how to bound it: defined envelopes, instrumentation, and review of every boundary case — because 'incoming threat' is a category that stretches.

SOURCES [5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross

Operating envelope

The bounded conditions — area, targets, time, weather — within which a system is authorised to act.

An envelope turns 'autonomous' from a property of a machine into a property of a mission: this airspace, these target classes, this duration, these abort conditions.

Envelopes are how defensive autonomy is made lawful and how it quietly expands. Each boundary case resolved under pressure widens the envelope a little; nobody ever decides to widen it a lot.

Why might this follow?
  1. Saturation defeats human-by-human authorisation
  2. Bounded defensive autonomy is authorised
  3. Boundary cases widen the envelope quietly
  4. No one ever decides the cumulative change
What could prevent or alter this?

Envelope discipline is possible: defined target classes, mandatory review of every boundary case, published aggregate statistics. It requires an institution with the standing to say no during an attack's aftermath.

SOURCES [5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross

Voluntary Assurance · 2031

Proliferation, software-style

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.

Autonomy packages — commercial drones, open vision models, cheap compute — spread to smaller states and proxies. This is software diffusion, not an aircraft programme.

What was a leading-edge capability in 2024 is a purchasable package in 2031: airframes from the commercial market, vision models from open repositories, integration from a contractor with a dozen employees. Smaller states field meaningful strike capacity without air forces; armed groups field it without states.

Export control built for platforms struggles with capability that travels as weights and code. Attribution weakens as sourcing diversifies — a strike by a generic airframe running generic software carries no flag.

The precedent problem compounds it. The states that deployed fastest, and asked fewest questions, now find their own practice quoted back as the standard everyone else is entitled to.

Proliferation of low-cost autonomous systems2024leading states2028most militaries2031small states, non-state groupsEach dot: an actor with meaningful autonomous strike capability (schematic).
Voluntary Assurance path: capability diffuses like software, not like combat aircraft.Illustration / scenario index — not measured data.Three clusters of dots for 2024, 2029 and 2034 showing the number of actors with meaningful autonomous strike capability growing from a few leading states to small states and non-state groups.
Why might this follow?
  1. Capability decomposes into commercial components
  2. Cost and skill barriers collapse
  3. Attribution weakens; precedent spreads
What could prevent or alter this?

Component-level controls, intelligence-led interdiction and negotiated norms slow diffusion; none reverses it. The window for setting restrictive precedent by example has largely closed on this path.

SOURCES [5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross[10] Defining Autonomy: Why Software, Not Drones, Will Decide the Next War, Center for Strategic and International Studies (CSIS)

Voluntary Assurance · 2033

A machine-speed near miss

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.
Shared pressure · EscalationRival automated systems misread defensive action as preparation for attack.

One system's defensive posture reads, to another system, as preparation for attack. Leaders inherit machine-framed options on machine-set deadlines. Luck holds — this time.

This chapter is conditional: it requires warning and response systems coupled tightly enough to react to each other, which is a choice, not an inevitability. On this path, that choice has been made by default.

During a regional crisis, an automated air-defence network raises readiness in response to anomalous satellite manoeuvres. A rival's monitoring system classifies the change as strike preparation and recommends dispersal — which a third system reads as mobilisation. Within an hour, three governments hold machine-generated assessments that are individually rational and mutually incompatible.

The crisis defuses through luck and weather. The near miss is classified on all sides; nothing about the architecture changes, because on each side the machines performed as designed.

Machine-speed escalation loopSTATE A
air-defence net raises readiness
STATE B
reads it as strike preparation; disperses forces
STATE C
reads dispersal as mobilisation
each assessment is rationalall are mutually incompatibleTempo is set by machines; diplomacy runs at human speed.
Three automated postures read each other — and each response confirms the others' worst assessment.Three state systems arranged in a triangle. Arrows between them show a defensive readiness change being read as attack preparation, prompting dispersal, which is read as mobilisation.

Concepts

Escalation coupling

When rival automated systems watch and react to each other faster than humans can confer.

One side's automated readiness change is another side's sensor input. Tightly coupled sensing and response systems can drive interaction cycles at machine speed, with each output individually rational and the ensemble destabilising.

This is a known failure pattern of past early-warning systems, made faster. It is conditional — it requires coupling that need not be built — which is why it is labelled speculative in this scenario.

SOURCES [8] The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I, Stockholm International Peace Research Institute (SIPRI)

Why might this follow?
  1. Rival sensing and response systems couple
  2. Defensive signals read as offensive preparation
  3. Assessment tempo outruns diplomatic tempo
  4. Nothing changes, because nothing 'failed'
What could prevent or alter this?

Crisis channels with machine-speed provisions, mandated human pauses at thresholds, and mutual transparency about automated postures. All are buildable; none exists on this path, and near misses stay secret.

SOURCES [8] The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I, Stockholm International Peace Research Institute (SIPRI)

Phase III · 2034–2040Lock-in and long consequences

Voluntary Assurance · 2036

Algorithmic conflict becomes normal

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.

Persistent drone incursions, automated retaliation, deniable sabotage: cheap, continuous, partially autonomous conflict that resembles neither war nor peace.

By the mid-2030s, low-intensity autonomous operations are continuous: incursions probed and intercepted nightly, infrastructure sabotage by systems whose owners are unprovable, retaliation increasingly pre-delegated because response windows are shorter than meetings.

Humans still set objectives. But the translation of objectives into violence — what is targeted, when, at what confidence — is machine-shaped, and the thresholds that once separated war from peace have dissolved into a permanent operational hum.

Democratic accountability, built around identifiable decisions by identifiable people, finds little to grip. Elections change governments; they no longer discernibly change what the systems do.

Why might this follow?
  1. Cheap autonomy lowers the cost of using force
  2. Attribution failure lowers its political price
  3. Pre-delegation spreads as windows shrink
  4. Conflict becomes ambient rather than declared
What could prevent or alter this?

Even now, attribution capability and declared no-automation zones could re-create thresholds. Both require international coordination this path has spent a decade underinvesting in.

SOURCES [8] The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I, Stockholm International Peace Research Institute (SIPRI)[5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross

Voluntary Assurance · 2040

The question stops being asked

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.

The endpoint is institutional, not cinematic: after fourteen years without required records or independent verification, 'why did this strike happen?' has no answerable form.

No robots rose up. Force remains under nominal human authority, and most operations pass without visible failure. What has been lost is specific: the capacity to reconstruct decisions, assign responsibility and correct errors.

Investigations still open after civilian harm. They close without findings, because the evidence was never required to exist. Officials answer questions in the passive voice; inquiries recommend; nothing binds. A late international code of conduct is voluntary — it can inspect nothing and compel no incident data. The operator remains the only person ever named.

The entrenchment is the endpoint. Accountability failure is no longer an incident but an architecture — procured, doctrinal, precedent-setting — and every year of operation makes the retrofit more expensive. The exits are the same as in 2027: traceability, testable control, allocated responsibility. They now cost fourteen years more.

Path TwoBroad Restriction, 2027–2040

Bright-line prohibitions plus a broad licensing regime covering a wide category of military AI.

Phase II · 2027–2033The dangerous acceleration

Broad Restriction · 2027

A bright line, drawn broadly

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.

After an allied AI-assisted strike kills civilians, Parliament prohibits autonomous targeting of people — and wraps most military AI in one licensing regime.

The trigger is a strike abroad, AI-assisted, fatal to civilians, and vivid on every feed. Parliament acts within months. The Act's core is a clear public prohibition: no system may autonomously select and engage people. Around that line, it adds pre-deployment licensing for military systems that use AI to identify, track, prioritise or engage targets.

The prohibition is the Act's lasting achievement — clear, popular, enforceable at its centre, and influential abroad. Clear entity-level offences make obvious violations straightforward to sanction, and several allies adopt versions of both.

The licensing regime is where the trouble will grow: drafted in a crisis, it defines its category broadly, so a mine-clearance classifier and a loitering munition enter the same queue.

Why might this follow?
  1. A visible incident demands rapid action
  2. Acting fast means defining categories broadly
  3. A durable prohibition and a congestible queue are born together
What could prevent or alter this?

A risk-tiered framework prepared before the crisis was the alternative. It existed in think-tank drafts; it had no parliamentary owner when the window opened.

SOURCES [5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross[3] Proceed with Caution: Artificial Intelligence in Weapon Systems (HL Paper 16), House of Lords AI in Weapon Systems Committee

Broad Restriction · 2028

Unlike risks, one queue

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · TempoConnected sensors produce more recommendations than humans can investigate.

The tempo pressure arrives as applications, not targets: protective systems, decision aids and lethal platforms wait in the same overloaded approval channel.

The approval office is competent, small and immediately oversubscribed. Counter-drone interceptors, casualty-evacuation aids, navigation updates and targeting software require substantially the same dossier and the same sign-off chain. Waiting time is measured in months and grows.

The queue distributes delay indiscriminately, so it penalises the benign hardest: protective systems wait behind paperwork while the operational tempo problem they would answer keeps compressing. Units needing capability now find workarounds — indefinite 'trials', experimental designations, allied systems borrowed rather than procured.

The prohibition, meanwhile, works. Nothing in the banned class deploys, and everyone knows it. The Act's defenders and critics are both right, about different parts of it.

Regulatory congestionmine clearancemedical evacuationnavigation updatedecision supportcounter-dronelogistics AIlethal autonomysonar classifierONE GATEmonths ofwaiting timeOne approval process for every risk level — delay penalises the benign systems hardest.
Broad Restriction path, 2028: very different risks queue for the same gate.Eight differently-labelled systems, from mine clearance to lethal autonomy, funnel toward a single narrow approval gate with a long queue.
Why might this follow?
  1. One gate processes every risk level
  2. Assurance capacity cannot scale to volume
  3. Delay hits protective systems hardest
  4. Workarounds move activity off the books
What could prevent or alter this?

Risk-tiering within the Act — fast lanes for protective and non-lethal functions — is proposed annually and deferred annually as 'weakening'. The prohibition's popularity shields the queue's design.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Broad Restriction · 2029

The bright line holds — and reveals a gap

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · Civilian harmAn AI-supported targeting operation kills civilians and responsibility is contested.

Civilians die in an allied operation using an imported decision-support system. Where the Act applies, accountability works. This system sat outside it.

The civilian-harm test arrives twice. A domestic case first: a licensed reconnaissance system strays toward the prohibited boundary, and the regime works — deployment suspended, licence conditions tightened, findings published. Decisive, public, fast.

Then the harder case. A coalition operation kills civilians using targeting software procured through an allied programme under a treaty exemption. The Act's licensing never touched it; its logs sit with a foreign supplier; its legal review happened in another jurisdiction.

The inquiry can state the outcome but not reconstruct the mechanism. The lesson is structural: a national bright line governs what you build, not what you plug into — and coalition warfare runs on what you plug into.

Why might this follow?
  1. Domestic regime handles the domestic case well
  2. Allied exemption routes around the licensing
  3. Foreign evidence lies beyond national compulsion
  4. The gap is structural, not accidental
What could prevent or alter this?

Coalition assurance agreements — shared logging standards, mutual audit access — could close the gap. They require allies to accept scrutiny Britain itself resisted two years earlier.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)[6] Summary of NATO's revised Artificial Intelligence strategy, NATO

Broad Restriction · 2030

Concentration and dependency

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.

Compliance costs consolidate domestic supply into a few primes; allied-procurement exemptions deepen reliance on systems Britain cannot inspect.

Licensing costs what only large companies can carry: specialist lawyers, compliance staff, capital to survive an eighteen-month cycle. Small firms — often the best software houses — sell to the primes or leave defence. Supplier diversity falls; prices rise; iteration slows.

Capability gaps get filled through the Act's allied-procurement exemption: systems certified by treaty partners enter service without domestic licensing. The mechanism is lawful and explicit — foreign systems do not evade the Act; they are exempted by it, because the alternative was going without.

The result is a quiet inversion. Britain now has the strictest rules for what it builds and diminishing insight into what it operates. And a licence, once granted, is rarely revisited — certification drifts toward paper shield, because the Act bought approval, not monitoring.

Why might this follow?
  1. Fixed compliance costs meet unequal firm sizes
  2. Domestic supply consolidates and thins
  3. Exemptions import capability minus scrutiny
  4. Sovereign insight declines behind strict rules
What could prevent or alter this?

Subsidised certification for small firms and audit-access conditions on exempted imports were both available. Each required spending assurance money the Act's design assumed away.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Broad Restriction · 2031

The saturation test: a narrow exception

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · SaturationA mass drone attack exceeds human-by-human defensive authorisation.

A mass drone attack overwhelms human-by-human authorisation. The framework bends — a bounded emergency exception for defensive interception — rather than breaking.

Several hundred drones against a deployed task group; engagement rules requiring human authorisation per interception; commanders who protect their people with every automated mode the systems possess. The facts force the question the Act deferred.

The response is narrower than collapse. Parliament legislates an emergency exception: autonomous interception of uncrewed incoming threats, in defined envelopes, with mandatory logging and quarterly review. The person-targeting prohibition is untouched. The framework's defenders correctly call this bending, not breaking.

But the exception's paperwork lives partly in classified annexes, and 'uncrewed incoming threat' acquires interpretations that never reach the public register. A second, quieter regime has begun to grow inside the first.

Why might this follow?
  1. Saturation defeats per-engagement authorisation
  2. Commanders act; law follows
  3. A bounded exception is legislated
  4. Its interpretations migrate into classified annexes
What could prevent or alter this?

Pre-authorised defensive envelopes, designed in peacetime with published statistics, were the tiered alternative. The Act's single category made designing them politically impossible until the attack forced it.

SOURCES [5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross

Broad Restriction · 2033

Two tracks, and a near miss in the dark

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.
Shared pressure · EscalationRival automated systems misread defensive action as preparation for attack.

Public licensing stays strict; sensitive systems live in classified channels. When rival automated systems misread each other, the review happens where nobody can see it.

By 2033 the regime is visibly two things. The public track: slow, rigorous, covering logistics and training systems that threaten nobody. The classified track: emergency exceptions, allied imports and intelligence systems, reviewed by a small committee, briefed after the fact.

The escalation test lands in the second track. A coalition early-warning system misreads an adversary's defensive posture during a crisis; automated recommendations propagate before humans intervene. The person-targeting prohibition is not implicated — nothing engaged — but the machinery that nearly drove escalation sits exactly where scrutiny is weakest.

The near miss is reviewed, classified, and filed. The public regime's strictness is real; it just governs the systems that matter least.

Concepts

Escalation coupling

When rival automated systems watch and react to each other faster than humans can confer.

One side's automated readiness change is another side's sensor input. Tightly coupled sensing and response systems can drive interaction cycles at machine speed, with each output individually rational and the ensemble destabilising.

This is a known failure pattern of past early-warning systems, made faster. It is conditional — it requires coupling that need not be built — which is why it is labelled speculative in this scenario.

SOURCES [8] The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I, Stockholm International Peace Research Institute (SIPRI)

Why might this follow?
  1. Burden pushes sensitive systems toward classification
  2. Scrutiny inverts: least where risk is greatest
  3. The escalation near miss happens off the books
What could prevent or alter this?

A cleared technical oversight body spanning both tracks could reunify scrutiny. Creating it requires admitting the dualism publicly — which no government managing the classified track wants to do first.

SOURCES [8] The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I, Stockholm International Peace Research Institute (SIPRI)[3] Proceed with Caution: Artificial Intelligence in Weapon Systems (HL Paper 16), House of Lords AI in Weapon Systems Committee

Phase III · 2034–2040Lock-in and long consequences

Broad Restriction · 2036

Contested rollback

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.

A delayed counter-drone system is blamed for battlefield deaths. Reform pressure is real and partially justified — and reform, not repeal, is what follows.

The failure the critics predicted arrives: soldiers die in an engagement a queued protective system — thirty-one months in approval — would plausibly have altered. 'Died waiting for paperwork' is politically unanswerable, and a rollback coalition forms.

What follows is a fight, not a collapse. The prohibition on autonomous person-targeting survives — a decade of observance has made it close to untouchable. The licensing regime is cut back: risk tiers at last, fast lanes for protective systems, sunset clauses on the emergency exceptions.

The reform is what the Act should have been in 2027, built nine years late, under worse conditions, by a parliament that now trusts neither the old regime nor its critics.

Why might this follow?
  1. Queue delay contributes to a visible failure
  2. Rollback pressure gains legitimate grievances
  3. The prohibition survives; the queue is reformed
What could prevent or alter this?

Whether reform stops at tiering or slides further depends on the next incident and who owns the narrative. The scenario assumes contested equilibrium, not a deregulation spiral — that is a judgement, not a law.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Broad Restriction · 2040

Precaution with dependency

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.

The bright line held: fewer autonomous-engagement systems, a norm with real reach, no domestic person-targeting deployments. The costs are sovereignty, concentration and a two-track regime.

The honest audit records durable achievements. Britain fielded fewer systems capable of autonomous engagement than any comparable power. The public prohibition on autonomous targeting of people held for thirteen years, shaped allied practice and anchored a proposed international ban — though states never agreed its definitions or defensive exceptions. Incidents involving prohibited classes are genuinely fewer.

The bill is structural. A defence-software sector of a few primes. Deep dependence on allied systems Britain cannot inspect, exempted into service by treaty. A public regime that is strict and a classified regime that is not, with the gap between them now institutional. Limited influence over adversaries who never accepted the premise.

It is a real form of safety — narrower and more brittle than its founders intended, and purchased with capabilities and insight that will be expensive to rebuild.

Path ThreeTiered Assurance, 2027–2040

Hard prohibitions plus function- and risk-based testing, change control and independent assurance.

Phase II · 2027–2033The dangerous acceleration

Tiered Assurance · 2027

Interim clauses, provisional tiers

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.

No mature regulator appears. Britain starts with what procurement can do now: contract clauses, a provisional function-and-risk framework, and two hard prohibitions.

The tiered path begins unglamorously. New contracts carry interim clauses: logging requirements, disclosure of known limitations, defined behaviour on communications loss. A provisional framework sorts systems by function — logistics, surveillance, decision support, defensive interception, targeting influence — with scrutiny scaled to each.

Two prohibitions anchor it, aligned with the ICRC's direction: no autonomous selection and engagement of people; no systems whose effects cannot be understood, predicted and limited.

Everything else is explicitly temporary. There is no assurance authority yet, no settled test protocol, not enough cleared engineers to review what already exists. The framework's founders publish that list of gaps themselves — the honesty is strategic, because the gaps will surface anyway.

Tiered Assurance · 2028

Control claims meet the clock

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · TempoConnected sensors produce more recommendations than humans can investigate.

Pilot tests run human-control claims at realistic tempo. Most systems fail on the first pass. That is the finding, not the failure.

The tempo pressure hits here as it hits everywhere: the targeting network produces more recommendations than humans can check. The difference is a pilot programme that measures it. Operators are tested under realistic load: what they see, how long they have, whether they can refuse, what happens when the queue doubles.

First results are bad. Approval rates above ninety-five per cent at under ten seconds; uncertainty displays nobody reads; automation bias fully expressed. Three systems fail acceptance outright.

The failures are the product. For the first time, 'meaningful human control' has a measurable shape — and a visible cost: review capacity, interface redesign, slower deployment. So does its absence: the pilot's data quantifies exactly what unreviewed tempo buys and loses.

The human-control test panelMEANINGFUL HUMAN CONTROL — ACCEPTANCE TESTtime to review and challenge the recommendationuncertainty and source provenance displayedcontext on target and civilian environmentauthority to delay, reject or overridedefined behaviour on communications lossautomation bias measured under realistic loaddecision record preserved for investigationTested at operational tempo — not in a controlled demonstration.
Human control as a test: each condition demonstrated at realistic operational tempo, not asserted in a contract.A test panel listing seven conditions of meaningful human control, each with a checkbox, from review time and uncertainty display to preserved decision records.

Concepts

Meaningful human control

Not whether a human pressed the button, but whether they had the information, time and authority to decide.

UK policy requires 'context-appropriate human involvement' in weapons that identify, select and attack targets. The unresolved question is what counts.

A testable version asks: what does the operator see, including uncertainty? How long do they genuinely have? Can they reject, delay or suspend? What happens on communications loss? Does approval behaviour collapse under realistic tempo? Those are measurable properties of a system, not assertions in a contract.

SOURCES [2] Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence, UK Ministry of Defence[4] Government response to the House of Lords AI in Weapon Systems Committee report, UK Government

Automation bias

The documented human tendency to accept a machine's suggestion rather than check it.

Decades of human-factors research in aviation, medicine and driving show that when a system is right most of the time, people stop verifying it — especially under time pressure, fatigue and workload.

In targeting, automation bias converts 'a human approved each strike' from a safeguard into a formality. It is measurable, which means it can be tested for — or ignored.

Why might this follow?
  1. Tempo pressure is measured, not assumed
  2. Systems fail tests that previously didn't exist
  3. Control acquires a price and a definition
What could prevent or alter this?

Testing to the test begins immediately: suppliers tune for the scenario library. Rotating adversarial test design, red-teamed by operators, is the counter — and needs the scarce experts everyone is short of.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)[2] Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence, UK Ministry of Defence

Tiered Assurance · 2028

The supplier-notice case

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.

A UK supplier learns its exported decision-support tool is being used outside its declared envelope. Nobody — company, regulator or contract — knows whose move it is.

The interim clauses meet their first hard case. A British supplier receives credible, specific reports that an export customer is using its decision-support software outside the declared operating envelope, in operations drawing IHL allegations.

The company improvises: an internal review, a quiet warning to the customer, continued support. The contract names no trigger for restriction; the regulator-to-be has no jurisdiction abroad; no law says what notice obliges. When the case becomes public, every actor can honestly say no rule told them what to do next.

The scandal does what argument had not: it makes a continuing due-diligence duty — notice, assessment, reporting, restriction, escalation — the centrepiece of the statutory proposals now being drafted for 2030.

Why might this follow?
  1. Interim clauses cover new contracts, not conduct after notice
  2. A credible misuse report finds no defined obligation
  3. Improvisation and publicity expose the gap
  4. A statutory due-diligence duty enters the 2030 bill
What could prevent or alter this?

Suppliers lobby to keep duties voluntary; customers resist end-use conditions. What carries the duty into statute is the political cost of the phrase 'no rule told them what to do'.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)[16] Bribery Act 2010, section 7 — failure of commercial organisations to prevent bribery, UK Parliament (legislation.gov.uk)

Tiered Assurance · 2029

Civilians still die. Evidence exists.

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · Civilian harmAn AI-supported targeting operation kills civilians and responsibility is contested.

A certified system contributes to a strike that kills nine civilians. The logs work, the inquiry works — and the certification itself is found wanting.

The civilian-harm test is not prevented. A decision-support system, certified the previous year, ranks a civilian convoy as a military target under sensor conditions its test scenarios never covered. The strike kills nine people.

What differs is what follows. Logs exist and are intact. The supplier's disclosed limitations are on file. The interface's behaviour can be replayed. Within weeks the inquiry can say what the system contributed, what the operator saw, and which test scenario should have existed and did not.

The uncomfortable finding: certification produced false assurance. The system passed because the test was inadequate, and passing had been read as safety. Requirements change; the test library grows; the dead are still dead. The difference is a system that learns, not one that doesn't kill.

Failures still occur; they produce evidence.

Concepts

Traceability & logs

The records that let anyone reconstruct, after the fact, what a system saw, recommended and did.

Reconstruction needs: the model version in use, its inputs, its outputs and confidence, what the operator was shown, and what they did — preserved in tamper-evident form, retained long enough to investigate.

None of this exists by default. Logging is a procurement requirement or it is absent, and its absence is discovered exactly when it matters most: after civilians are dead.

SOURCES [6] Summary of NATO's revised Artificial Intelligence strategy, NATO[7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Why might this follow?
  1. An unmodelled condition defeats a certified system
  2. Preserved evidence enables real reconstruction
  3. Certification itself is found deficient
  4. Requirements change because findings exist
What could prevent or alter this?

False assurance is tiered regulation's characteristic failure: a passed test read as a guarantee. The counter is institutional humility — treating certification as a floor, never a defence.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Tiered Assurance · 2030

The authority, born of failure

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.

Pilot data plus a public inquiry produce what neither could alone: a statutory assurance authority, change-control rules and responsibility allocated before deployment.

The Defence AI Assurance Authority is legislated three years after the interim clauses — built from the pilot's methods and the 2029 inquiry's findings, not from a blank page. It is small, cleared, and reports outside the delivery chain. Its central power is mundane: material changes to models, data or operating context trigger reassessment.

Responsibility is allocated in the same statute, following knowledge and control. The state answers for testing, integration and deployment conditions; commanders for operational authorisation; suppliers for architecture, data practices and truthful disclosure; operators only for what the tested system genuinely left in their hands. Alongside it: a corporate failure-to-prevent offence with a reasonable-procedures defence, the due-diligence duty the 2028 notice case demanded, and regulatory offences that bite below the war-crime threshold.

The costs are immediate: cleared ML auditors are scarcer than the plan assumed, the authority's first budget is fought over annually, and industry's compliance spending favours exactly the large firms tiering was meant not to.

Allocation of responsibilitySUPPLIER
architecture · known limitations · data practices · truthful disclosure
PROCURING AUTHORITY
acquisition · integration · testing · deployment conditions
MILITARY COMMAND
operational authorisation within tested conditions
OPERATOR
only decisions genuinely within the operator's control
Allocated in advance — so an answer exists afterwards.
Tiered Assurance path: responsibility follows knowledge, control, contribution and capacity to prevent harm.Four stacked rows allocating responsibility: suppliers for architecture and disclosure, the procuring authority for acquisition and deployment conditions, command for operational authorisation, operators only for what they genuinely control.

Concepts

Independent assurance

Technically competent testing and audit by people whose incentives are not the programme's incentives.

Self-certification fails predictably where delivery is rewarded and delay is punished. An assurance function needs three scarce things: technical staff who can read a model card and a test log; security clearance; and institutional separation from the delivery chain.

All three are expensive, and cleared ML auditors are genuinely rare — which is why this scenario treats assurance capacity, not rules, as the binding constraint.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Change control

Rules about what happens when the software, model or data behind an approved system changes.

A machine-learning system is not stable the way a rifle is: new training data, retuned thresholds or a swapped model can change behaviour without changing the hardware.

Change control defines which changes trigger re-testing and re-review, who signs off, and what evidence must be produced. Without it, 'the system passed review' describes history, not the thing currently deployed.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Who answers for what

Suppliers, the state, commanders and operators hold different duties — responsibility follows knowledge and control.

Legal responsibility does not vanish when systems are complex; it becomes hard to assign in practice if the evidence and the duties were never allocated.

A workable allocation: suppliers answer for architecture, known limitations and truthful disclosure; the procuring state for testing, integration and deployment conditions; commanders for operational authorisation; operators only for decisions genuinely within their control.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Why might this follow?
  1. Evidence from pilots and inquiry converges
  2. Statute fixes duties to knowledge and control
  3. Change control ends the reviewed-once fiction
  4. Capacity, not rules, becomes the constraint
What could prevent or alter this?

Capture and starvation are the standing threats: an authority staffed from the industry it audits, or defunded at the first quiet spending review. Neither is hypothetical; both nearly happen.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)[4] Government response to the House of Lords AI in Weapon Systems Committee report, UK Government

Tiered Assurance · 2031

The saturation test: bounded autonomy, visible cost

projected · medium confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as medium-confidence.
Shared pressure · SaturationA mass drone attack exceeds human-by-human defensive authorisation.

Pre-authorised defensive envelopes absorb the mass attack. One interceptor, still in review, sits out the fight — and the cost of caution becomes public.

When the saturation attack comes, the framework has an answer it prepared in peacetime: autonomous interception of uncrewed threats inside defined, instrumented envelopes, pre-authorised and logged. The engagement is machine-speed, lawful and reconstructable. Envelope discipline holds because an institution exists whose job is holding it.

The cost arrives the same day. A new counter-swarm system, its certification incomplete, stays grounded at one location; the older system there performs worse; equipment is lost that the newer system would likely have saved. No one dies, this time, but the ledger is public and the criticism is fair.

Tiered assurance means some lawful capability arrives later than the threat. The framework survives because it says so out loud, before the attack rather than after.

Concepts

Defensive autonomy

Machine-speed interception of incoming munitions and drones — the strongest case for autonomy, and the usual door.

A hundred incoming drones cannot be intercepted by a human approving each shot; ship and base defence systems have operated automated modes for decades against clearly inanimate threats.

The governance question is not whether to permit this but how to bound it: defined envelopes, instrumentation, and review of every boundary case — because 'incoming threat' is a category that stretches.

SOURCES [5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross

Operating envelope

The bounded conditions — area, targets, time, weather — within which a system is authorised to act.

An envelope turns 'autonomous' from a property of a machine into a property of a mission: this airspace, these target classes, this duration, these abort conditions.

Envelopes are how defensive autonomy is made lawful and how it quietly expands. Each boundary case resolved under pressure widens the envelope a little; nobody ever decides to widen it a lot.

Why might this follow?
  1. Peacetime envelopes absorb wartime saturation
  2. Discipline holds because an institution owns it
  3. A delayed system makes caution's price visible
What could prevent or alter this?

Every delayed deployment builds pressure for emergency certification. The framework's answer — provisional certificates with enhanced logging — works only while the authority can actually process them.

SOURCES [5] Autonomous Weapon Systems and International Humanitarian Law — position paper, International Committee of the Red Cross[7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Tiered Assurance · 2033

Mutual recognition, limited

projected · low confidence
Basis: A future development extrapolated from identified incentives and capabilities. A projection is a causal proposition, not evidence. Confidence: the scenario treats this causal step as low-confidence.

Five allies accept each other's test evidence for defined system classes. It is real, useful — and far short of seamless NATO-wide convergence.

Assurance regimes attract partners the way standards do: slowly, partially, and for commercial reasons as much as principled ones. By 2033 five allied states accept one another's test evidence for logistics, navigation and defensive-interception classes. A supplier certifying once can sell five times; mid-sized firms benefit most.

On those foundations, the UK and four partners table a treaty proposal first floated in 2031: an International Autonomous Weapons Agency. By 2035 it exists in pilot form — a confidential registry, accredited test laboratories, incident-investigation teams — with major powers pointedly outside.

The limits are as important as the achievement. Targeting-influence systems remain nationally reviewed; two major allies still use incompatible definitions of 'human control'. NATO-wide convergence is a communiqué aspiration, not a fact.

What the cluster changes is the market: inside it, accountability stops being a competitive disadvantage. Outside it, systems built to no comparable standard compete on price — and win sales this cluster refuses.

Why might this follow?
  1. Shared evidence formats cut certification costs
  2. Partial recognition forms among five states
  3. Targeting systems stay national; definitions diverge
What could prevent or alter this?

Divergence is the pressure that never resolves: adversaries outside the regime, allies half-inside it, and a price gap that tempts every procurement ministry in between.

SOURCES [6] Summary of NATO's revised Artificial Intelligence strategy, NATO[7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Phase III · 2034–2040Lock-in and long consequences

Tiered Assurance · 2034

The escalation test: brakes that reach one side

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.
Shared pressure · EscalationRival automated systems misread defensive action as preparation for attack.

Rival systems misread a defensive posture. Mandated human pauses and incident channels help manage the crisis — and bind nobody outside the regime.

The escalation test arrives as it does on every path: an automated warning system reads a defensive readiness change as attack preparation, and recommendation loops begin to couple. The conditional nature of this chapter matters — it requires coupled architectures that this framework has spent years limiting but cannot abolish.

On the regime's side, the brakes designed in peacetime engage: automated escalatory recommendations pause at mandated human checkpoints; a standing incident channel carries a clarification within minutes; bounded automation stays bounded because its envelopes were audited.

The adversary's systems observe no such rules. The crisis is managed, narrowly, and the honest lesson is recorded in public: national assurance regimes govern one side of a coupling. Escalation controls that bind only yourself are a hedge, not a solution.

Concepts

Escalation coupling

When rival automated systems watch and react to each other faster than humans can confer.

One side's automated readiness change is another side's sensor input. Tightly coupled sensing and response systems can drive interaction cycles at machine speed, with each output individually rational and the ensemble destabilising.

This is a known failure pattern of past early-warning systems, made faster. It is conditional — it requires coupling that need not be built — which is why it is labelled speculative in this scenario.

SOURCES [8] The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I, Stockholm International Peace Research Institute (SIPRI)

Why might this follow?
  1. Coupled warning systems misread defence as attack
  2. Mandated pauses and channels slow one side
  3. Asymmetry limits what any national regime can do
What could prevent or alter this?

The missing instrument is bilateral: reciprocal transparency about automated postures, agreed pause protocols. This scenario does not assume adversaries sign; it records the cost of their not signing.

SOURCES [8] The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I, Stockholm International Peace Research Institute (SIPRI)

Tiered Assurance · 2036

Sustained conflict strains the machinery

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.

Two years of continuous operations: emergency certificates accumulate, recertification backlogs grow, and permissions granted in crisis prove hard to unwind.

The framework was tested by a crisis and passed. A long war is a different test. By the second year of sustained operations, emergency certificates outnumber standard ones in some system classes. Recertification of updated systems runs months behind. A third of the authority's cleared staff have been poached by the suppliers they audited.

Unwinding proves hardest. Each emergency permission was time-limited on paper; each expiry meets an operational argument for renewal, and renewals are granted more often than not. IAWA's first serious incident investigation does allocate responsibility across a state and two suppliers — while exposing delays, secrecy disputes and uneven access. The envelope discipline of 2031 frays at exactly the rate the war continues.

The framework bends without breaking — so far. Its designers' question is now empirical: whether review capacity or the war runs out first.

Why might this follow?
  1. Sustained tempo outruns review capacity
  2. Emergency measures accumulate faster than they expire
  3. Staffing drains toward better-paying suppliers
  4. Discipline frays at the rate of the war
What could prevent or alter this?

Surge capacity — reserve auditors, allied load-sharing, pre-agreed triage rules — was planned for a crisis, not a decade. Building wartime-scale assurance in wartime is the unsolved problem.

SOURCES [7] Responsible Procurement of Military Artificial Intelligence, Stockholm International Peace Research Institute (SIPRI)

Tiered Assurance · 2040

Accountable, expensive, contested, incomplete

speculative · low confidence
Basis: A lower-confidence possibility included because its consequences could be extreme. Confidence: the scenario treats this causal step as low-confidence.

High autonomy in bounded roles; tested human judgement where people and escalation are at stake. Failures still occur. They produce evidence, findings and change.

By 2040 Britain operates machine-speed defence inside audited envelopes — interception, mine countermeasures, logistics — while decisions about people, ambiguous environments and escalation still pass through humans whose actual control is tested, not asserted.

The ledger is mixed and public. The 2029 and 2037 incidents killed civilians despite certification. Assurance costs more than its founders promised and still favours large firms. Two allies' regimes remain incompatible; adversaries observe none of it; emergency permissions from the war years are still being unwound. Some lawful strikes were slowed or forgone, and some of those decisions were probably wrong.

What the money bought is specific: fourteen years on, 'why did this strike happen?' still has an answerable form — preserved evidence, an institution obliged to answer it, and requirements that change when the answer is bad. Not a safe world. A correctable one.

Comparison · 2040

Three end states, side by side

None of these futures is frictionless, and none is completely safe. The comparison below is the honest ledger: what each path buys, what it costs, and what it leaves unresolved.

The three 2040 end states, as scenario-index profiles VOLUNTARY ASSURANCEautonomydecision timetraceabilitysupplier dutyproliferationcouplingscrutinysystemic risk: EXTREME BROAD RESTRICTIONautonomydecision timetraceabilitysupplier dutyproliferationcouplingscrutinysystemic risk: HIGH TIERED ASSURANCEautonomydecision timetraceabilitysupplier dutyproliferationcouplingscrutinysystemic risk: GUARDED
The three 2040 end states as indicator profiles. Illustrative scenario indices, not forecast probabilities.
Comparison of the three 2040 end states
2040 Voluntary Assurance Broad Restriction Tiered Assurance
Military effectivenessHighest tempo and mass; brittle against manipulated inputs, with failure modes nobody can examine.Fewer systems in restricted classes; capability gaps filled by allied imports; reformed after 2036, years behind.Slower at the lethal frontier by design; fast in bounded roles; some lawful strikes delayed or forgone.
Civilian protectionErodes as review collapses and envelopes creep; harm recurs without producing correction.Strong where the bright line applies; weakest around imported, exempted and classified systems.Anchored by prohibitions and tested control; civilians still die (2029, 2037) — with findings and changed requirements.
Human controlFormal approval persists; substantive judgement dissolves into throughput.Robust in the public track; untested in the classified track and in allied systems.Defined, tested and evidenced; narrowest where machine speed genuinely governs.
Democratic scrutinyParliament receives conclusions it cannot test; secrecy and failure merge.Visible strict rules; a second classified regime whose existence is the real scrutiny problem.Cleared technical oversight spans classified work; wartime emergency measures strain but don't escape it.
ProliferationFastest diffusion; the leading users' own practice becomes everyone's precedent.Domestic restraint with real norm influence on allies; little effect on adversaries or the commercial pathway.Diffusion continues globally; the five-state assurance cluster sustains an accountable market inside it.
Crisis resilienceNo designed limits to erode; escalation couples at machine speed and near misses stay secret.Bent at the 2031 saturation test into a narrow exception; the classified track absorbed the 2033 near miss unseen.Absorbed 2031 within pre-built envelopes; strained badly by sustained war after 2035.
Catastrophic-risk exposureHighest: coupled machine-speed systems, no shared brakes, degraded attribution.Moderate but poorly observed: the riskiest machinery sits in the least-scrutinised track.Lowest of the three, not low: one-sided brakes cannot bind adversaries.
Principal unresolved danger'Why did this strike happen?' loses its answerable form — and stops being asked.Norm leadership without technical sovereignty; a two-track regime nobody will admit to first.Assurance capacity in a long war, and the standing temptation to trade tested control for tempo.

Accountability at 2040

The same reforms proposed above — allocated responsibility, corporate duties and international verification — arrive differently on each path. Each column contains benefits, failures and unresolved problems.

Accountability outcomes of the three 2040 end states
Accountability Voluntary Assurance Broad Restriction Tiered Assurance
State responsibilityFormally intact; practically diluted — inquiries close without findings and reparation is discretionary.Asserted strongly in public law; blurred where allied and exempted systems did the work.Explicit and non-delegable in statute; still contested when coalition systems are involved.
Command & operator accountabilityThe operator is the only person ever named; command decisions dissolve into tempo.Clear where the bright line applies; opaque inside the classified track.Tested control defines what each genuinely decided; operators stop being the default defendant.
Corporate dutiesVoluntary policies that differ by company and can be waived in crisis.Strong licence conditions at approval; weak continuing duties after it.Statutory failure-to-prevent offence, continuing due diligence and notice-to-action obligations.
Senior-manager & individual liabilityNearly unreachable: fragmented evidence defeats complicity cases.Reachable for obvious domestic violations; rare in practice.Defined offences for knowing or reckless assistance; the first convictions are slow and hard-fought.
Incident evidence & victim remedyLogs optional and overwritten; victims receive conclusions, not explanations.Good records for licensed systems; gaps for imports; remedy depends on venue.Preserved evidence by default; a reparations contribution mechanism exists — underfunded.
Independent national assuranceNone; assurance stays inside programmes.A licensing office, congested and approval-focused.A statutory authority with change control — chronically short of cleared auditors.
International verificationA late voluntary code with no inspection or compulsion.A proposed prohibition treaty stalled on definitions.IAWA pilots: registry, accredited labs, incident teams — with major powers outside.
Material software updatesUnreviewed; the deployed system diverges silently from the reviewed one.Re-approval required on paper, rarely enforced for exempt systems.Updates trigger reassessment; wartime backlogs strain the rule.
Smaller suppliers & innovationEasy entry, no duties; incumbency by shipping speed.Compliance moat; small firms exit or sell to primes.Published standards and shared test infrastructure help — costs still favour scale.
Enforcement against non-parties & proxiesNone; precedent set by the least careful users.Little beyond export control; proxies untouched.Named openly as the regime's limit; managed, not solved, through export and end-use verification.

The risk ladder

From present harm to existential risk

These five levels are not a single slope, and the higher rungs are not the inevitable destination of the lower ones. Each belongs to a different order of evidence, and the honest way to present them is separately.

  1. 1Direct humanitarian harmpresent now

    Misidentification, civilian casualties, unpredictable behaviour and automation bias. These harms are immediate, documented in analogous systems, and researchable today.

  2. 2Failure of legal and democratic accountabilitydeveloping

    Responsibility distributes across secret technical systems, contractors and military institutions until Parliament cannot reconstruct how a lethal decision was made — the central subject of this study.

  3. 3Structural proliferation and cheaper conflictdeveloping

    Inexpensive autonomous systems make persistent surveillance, political repression and low-level conflict easier, and may make states readier to use force when fewer of their own soldiers are exposed.

  4. 4Global catastrophic escalationconditional

    AI embedded in early warning, cyber operations and strategic command creates pathways by which false information, adversarial manipulation or machine-speed interaction could drive large-scale — including nuclear — escalation.

  5. 5Conditional existential risklow certainty, extreme consequence

    Plausible only through additional assumptions: highly agentic frontier systems; access to cyber capability, military logistics or autonomous forces; loss of practical human supervision; rivalry that rewards deployment despite warnings; systems able to acquire resources or conceal behaviour.

The direct humanitarian danger is already present. Human extinction is not the inevitable endpoint of autonomous weapons. It is a low-certainty but extreme-consequence pathway that becomes more plausible if increasingly capable AI is connected to cyber systems, military logistics, strategic command or autonomous physical force while human control deteriorates.
Conditional pathway to catastrophic riskincreasing AI capabilitymilitary integrationcompressed human oversightstrategic dependenceadversarial manipulation or loss of controlcatastrophic escalation
A conditional pathway, not a prediction: every arrow is a policy choice that can be refused.Six steps connected by arrows, from increasing AI capability through military integration, compressed oversight, strategic dependence and loss of control to catastrophic escalation. Each arrow widens, indicating growing uncertainty.

A methodological warning

This is not a prediction. The shared timeline rests on documented events and clearly labelled reported claims; the three futures are causal propositions — if these incentives and institutions, then plausibly these consequences. The dangerous dynamics arrive in the late 2020s; 2040 is an institutional horizon, not the expected date of first harm. Indicators are illustrative scenario indices, and the extreme outcomes carry the heaviest uncertainty warnings.

The scenario has a point of view: it argues that accountability architecture is a core component of defence capability. That view deserves the same scrutiny the scenario applies to the policies it criticises — the Method page sets out the assumptions, and the ways this exercise could be wrong.

Sources and further reading

  1. The Strategic Defence Review 2025 — Making Britain Safer: secure at home, strong abroad UK Government, 2025 (external link)
  2. Ambitious, Safe, Responsible: our approach to the delivery of AI-enabled capability in Defence UK Ministry of Defence, 2022 (external link)
  3. Proceed with Caution: Artificial Intelligence in Weapon Systems (HL Paper 16) House of Lords AI in Weapon Systems Committee, 2023 (external link)
  4. Government response to the House of Lords AI in Weapon Systems Committee report UK Government, 2024 (external link)
  5. Autonomous Weapon Systems and International Humanitarian Law — position paper International Committee of the Red Cross, 2025–2026 (external link)
  6. Summary of NATO's revised Artificial Intelligence strategy NATO, 2024 (external link)
  7. Responsible Procurement of Military Artificial Intelligence Stockholm International Peace Research Institute (SIPRI), 2026 (external link)
  8. The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I Stockholm International Peace Research Institute (SIPRI), 2019 (external link)
  9. Technological Evolution on the Battlefield Center for Strategic and International Studies (CSIS), 2025 (external link)
  10. Defining Autonomy: Why Software, Not Drones, Will Decide the Next War Center for Strategic and International Studies (CSIS), 2026 (external link)
  11. Germany funds 50,000 strike drones for Ukraine with terminal autonomous tracking (report) Reuters, 12 July 2026 (external link)
  12. First deaths from 'killer drones' New Scientist (International Edition, issue of 20 June 2026), 20 June 2026 (external link)
  13. Palantir Q4 2023 Business Update (p. 5) Palantir Technologies, 2024 (external link)
  14. Rome Statute of the International Criminal Court, Article 25 (individual criminal responsibility) International Criminal Court, 1998 (2024 ed.) (external link)
  15. International Criminal Court Act 2001 UK Parliament (legislation.gov.uk), 2001 (external link)
  16. Bribery Act 2010, section 7 — failure of commercial organisations to prevent bribery UK Parliament (legislation.gov.uk), 2010 (external link)
  17. Customary IHL Rule 151 — individual criminal responsibility International Committee of the Red Cross (IHL database), n.d. (external link)
  18. ICRC position on autonomous weapon systems International Committee of the Red Cross, 2021 (external link)
  19. IAEA safeguards and verification International Atomic Energy Agency, n.d. (external link)
  20. IAEA verification and other safeguards activities International Atomic Energy Agency, n.d. (external link)
  21. The Nuremberg industrialist trials (IG Farben; Krupp) — case summaries US Holocaust Memorial Museum, n.d. — URL and corroboration required before publication
  22. Public Prosecutor v. Frans van Anraat — judgment summary ICRC national-practice database / Dutch courts, 2005–2009 — URL and corroboration required before publication
  23. Sweden v. former Lundin executives (alleged complicity in war crimes) Swedish Prosecution Authority / Stockholm District Court, ongoing as of July 2026 — URL and corroboration required before publication
  24. Lafarge proceedings — terrorism financing conviction distinguished from separate complicity litigation French courts / US DOJ, 2022–ongoing — URL and corroboration required before publication

Full entries, notes and verification status: Sources.