Papers
Topics
Authors
Recent
Search
2000 character limit reached

Assurance 2.0: Modern Continuous System Assurance

Updated 14 July 2026
  • Assurance 2.0 is a modern framework that continuously integrates rigorous assurance into system lifecycles, supporting innovation in AI-driven and autonomous systems.
  • It enhances traditional assurance by embedding structured argumentation, evidence management, and defeater analysis throughout development to reduce bias.
  • The methodology employs quantitative confidence assessment and dynamic tool support to balance rapid innovation with robust, certifiable safety standards.

Assurance 2.0 is a modernized framework for system assurance proposed by Robin E. Bloomfield and John Rushby in response to autonomous systems with major functions driven by machine learning and AI, ultra-rapid system development, and the persistent need for efficient, effective, and timely assurance. It recasts assurance from a post-hoc, costly appendage into an enabler that supports innovation and continuous incremental assurance, while also making assurance more rigorous through increased focus on the reasoning and evidence employed, and through explicit identification of defeaters and counterevidence (Bloomfield et al., 2020).

1. Origins, motivation, and problem setting

The motivating claim of Assurance 2.0 is that traditional assurance is often seen as a brake on innovation and is often costly and time consuming. In the traditional picture, assurance is primarily an after-the-fact exercise, often documenting what was done and retroactively justifying decisions. Assurance 2.0 instead places assurance throughout the system life cycle, from earliest design explorations, so that it offers constructive input, helps identify issues early, and continuously guides development toward trustworthy, certifiable systems (Bloomfield et al., 2020).

The framework was explicitly motivated by two kinds of pressure. One is new: autonomous systems with major functions driven by machine learning and AI, together with ultra-rapid system development and deployment. The other is familiar but persistent: the demand that assurance be efficient, effective, and timely. The central response is not to relax rigor, but to increase it, especially in the treatment of reasoning, evidence, counterevidence, and doubt (Bloomfield et al., 2020).

Later discussion of AI assurance extends this motivation by arguing that critical-systems practice should ask four questions: what is the system, how good does it have to be, what is the impact of criticality on system development, and how much should we trust it. That discussion also expresses concern that system boundaries are not broad enough, that the tolerability and nature of the risks are not sufficiently elaborated, and that current assurance methods lack theories that would allow behaviours to be adequately assured. It further emphasizes orders of magnitude difference in confidence needed in critical rather than everyday systems, and argues that everyday techniques do not scale in rigour (Bloomfield et al., 7 Jan 2025).

2. Argument structure and reasoning discipline

Assurance 2.0 retains structured argument approaches such as Claims-Arguments-Evidence (CAE) and Goal Structuring Notation (GSN), but strengthens them. The argument is a tree in which top-level claims are decomposed into subclaims, supported through evidence at the leaves and reasoning steps at interior nodes. The framework constrains argumentation through a limited set of rigorously defined building blocks: evidence incorporation, calculation, decomposition, substitution, and concretion. The point of this restriction is to avoid ad-hoc or error-prone argumentation and to clarify expectations at each step (Bloomfield et al., 2020).

A central discipline is explicit, as-deductive-as-possible reasoning. Deductive steps are those in which the subclaims logically entail the parent claim, often written as subclaims    parent claim\text{subclaims} \implies \text{parent claim}. If a step is inductive rather than deductive, the gap must be marked and addressed with side claims and additional evidence. Assurance 2.0 therefore treats argument quality not simply as good presentation, but as a matter of explicit logical form, explicit conditions for inference, and explicit treatment of the places where entailment does not hold directly (Bloomfield et al., 2020).

The framework also distinguishes between the development of models or theories and the assurance case that leverages them as evidence or argument. This separation becomes important when, for example, an operational profile, a probabilistic model, or a domain theory is used to support an assurance claim: the model is not the assurance argument itself, but an input to it (Bloomfield et al., 2020).

Subsequent work generalizes this structure with reusable assurance theories, side claims, and explicit counterarguments. In a scalable framework for self-driving vehicles, Assurance 2.0 is instantiated through a structured template and a three-tiered decomposition strategy. At the top level, development is divided into Requirements Engineering, Verification and Validation, and Post-Deployment; each phase is further decomposed according to its Product Development Lifecycle; and each lifecycle stage is then analyzed using an adapted 5M1E model: Man, Machine, Method, Material, Measurement, and Environment. That paper presents this as a way to support fine-grained traceability of safety claims, evidence, and potential defeaters (Chen et al., 30 Sep 2025).

3. Defeaters, counterevidence, and eliminative argumentation

A defining feature of Assurance 2.0 is its insistence that assurance should not consist only of a convincing positive case. It emphasizes explicit, systematic search and documentation of defeaters and counterevidence, partly because a purely positive case creates a breeding ground for confirmation bias. Defeaters are objections or sources of doubt that could undermine a claim. The basic distinction is between undercutting defeaters, where the evidence is insufficient or inadequate, and rebutting defeaters, where there is explicit evidence against the claim. The corresponding indefeasibility criterion is that a claim is justified only if all identified defeaters have been addressed and there are no outstanding doubts that could change the conclusion (Bloomfield et al., 2020).

This dialectical orientation is not auxiliary. It is one of the framework’s primary evaluation criteria, and it is closely connected to the claim that assurance development should resemble vigorous and skeptical debate. The framework also encourages counterclaims and countercases, namely arguments supporting the negation of the main claim, as a way to stress-test the main case for overlooked vulnerabilities or weaknesses (Bloomfield et al., 2020).

A later formalization distinguishes exploratory defeaters from exact defeaters. Exploratory defeaters express doubts about a node without necessarily stating its logical negation. Exact defeaters are the precise logical negation of the targeted claim and support eliminative argumentation, a style in which assurance is sought by refuting all reasons why the top claim could be false. The same work describes multiple levels of defeaters, defeater-on-defeater structures, and propagation rules in Clarissa/ASCE. In that semantics, an unsustained defeater makes the affected node unsupported; exact defeaters are handled by logical negation; and completed cases can retain defeaters and their subcases as documentation for future assessors (Bloomfield et al., 2024).

The practical significance of this mechanism is twofold. First, it formalizes the recording of dissent, doubt, and challenge inside the assurance case rather than outside it. Second, it provides a structured way to reduce confirmation bias by making skeptical scrutiny part of case construction rather than merely part of external review (Bloomfield et al., 2024).

4. Confidence, confirmation, and quantitative assessment

Assurance 2.0 treats confidence as central, but not reducible to a single number. One line of work argues that confidence should be assessed from three perspectives: positive, negative, and residual doubts. A later refinement organizes the assessment around four perspectives: logical soundness, probabilistic assessment, dialectical examination, and residual risks. Across these formulations, the common target is indefeasible confidence: confidence strong enough that credible doubts have been systematically considered and either resolved or consciously managed (Bloomfield et al., 2022, Bloomfield et al., 2024).

The positive perspective begins with logical soundness. Assurance cases are interpreted using Natural Language Deductivism, under which each reasoning step is treated as deductive. For evidence, the framework draws on confirmation theory. Let cc be a claim and ee the supporting evidence. Assurance 2.0 emphasizes not only P(ce)P(c|e), but also the discriminatory power of the evidence, namely how well ee supports cc versus ¬c\neg c. One confirmation measure advocated for this purpose is the Kemeny-Oppenheim measure:

P(ec)P(e¬c)P(ec)+P(e¬c).\frac{P(e|c) - P(e|\neg c)}{P(e|c) + P(e|\neg c)}.

The same body of work also discusses probabilistic valuation and propagation through the argument tree, for example by a product rule for deductive claims under independence,

Confidence in parent claim=iConfidence in subclaimi,\text{Confidence in parent claim} = \prod_i \text{Confidence in subclaim}_i,

and, for more conservative non-deductive settings, by a sum-of-doubts style bound (Bloomfield et al., 2020, Bloomfield et al., 2022).

More recent work develops a compositional probabilistic method for Assurance 2.0 arguments using elementary probabilistic constructions and Frechet bounds. It distinguishes different patterns of decomposition according to the relationships among subclaims and the way they eliminate concerns. For diversity or independence, doubts combine multiplicatively; for partitioned or disjoint concerns, parent confidence is a weighted sum; for containment or complete overlap, the most conservative estimate is bounded by the maximum subclaim confidence; and for cumulative chains, confidence is expressed by conditional products. That work is explicit that the primary evaluation criteria for Assurance 2.0 remain logical indefeasibility and dialectical examination, while probabilistic assessment is a complementary tool for cost/confidence tradeoffs and for judging the balance of confidence across a structured argument (Bloomfield et al., 21 Mar 2026).

A separate line of research proposes a compositional semantics for quantitative confidence assessment using Subjective Logic. In that approach, argument elements are represented as opinions (b,d,u,a)(b,d,u,a), and relations between elements are mapped to Subjective Logic operators so that confidence flows through the argument as an analyzable confidence network. The proposal emphasizes explicit warrants, principled handling of context, preserved provenance, and compatibility with GSN (Herd et al., 21 May 2026).

5. Dynamic assurance and lifecycle integration

Assurance 2.0 is explicitly continuous, early, and incremental. One concrete realization links safety performance measurement to the core assurance artifacts of a safety case by associating safety indicators, metrics, and measures with hazards, barriers, safety claims, and related artifacts. In that framework, the safety architecture and its underlying risk assessment model provide the bridge from safety measurement to operational risk. The dynamic measurement layer is gathered under the Safety Measurement Basis,

cc0

and operational data are used to revise the risk assessment iteratively through Bayesian inference:

cc1

The same framework characterizes change in safety risk through a risk ratio,

cc2

so that practical drift can be detected by monitoring its trend. The tool support described for this framework is AdvoCATE (Denney et al., 2024).

This dynamic perspective is consistent with the original claim that assurance should evolve iteratively and keep pace with modern, rapidly changing systems. It also makes the safety case explicitly responsive to operational evidence rather than static approval-time assumptions (Denney et al., 2024).

A related continuous-assurance framework for autonomous systems integrates design-time, runtime, and evolution-time assurance within a traceable, model-driven workflow. In its design-time phase it uses RoboChart for functional correctness and PRISM for probabilistic risk analysis; it further proposes a model-driven transformation pipeline, implemented as an Eclipse plugin, that automatically regenerates structured assurance arguments whenever formal specifications or their verification results change. That framework introduces placeholders such as trace_expr, monitor_id, confidence_threshold, and evidence_cost, together with stereotypes such as <<RuntimeAssumptionMonitor>>, <<ConfidenceMonitor>>, <<Reopened>>, and <<ImpactAnalysis>>, as machine-interpretable hooks linking runtime monitoring and evolution-time change analysis back into the assurance case (Abeywickrama et al., 17 Nov 2025).

6. Tool support, automation, and current debates

Assurance 2.0 has been accompanied by a sustained effort to make assurance artifacts amenable to automation without reducing them to purely structural checks. One approach formalizes assurance statements as Object-Property-Environment triples and translates claims, evidence, side claims, theories, and defeaters into answer set programming predicates. Using goal-directed ASP via s(CASP), that work analyzes semantic properties such as logical consistency, adequacy, completeness, theory-application correctness, harmonious coexistence, and indefeasibility. The stated aim is to move beyond tools that automate only syntax or document structure and toward tools that assess semantic coherence and correctness of the assurance argument itself (Murugesan et al., 2024).

Another line addresses review rather than construction. An LLM-as-a-judge framework for GSN-compliant assurance cases formalizes well-established review criteria with predicate-based rules and prompt templates. The four main criteria are argument comprehension, well-formedness (syntax), expressive sufficiency, and argument criticism and defeat. Experiments across GPT-4o, GPT-4.1, DeepSeek-R1, and Gemini 2.0 Flash report that DeepSeek-R1 and GPT-4.1 demonstrate superior performance, but also conclude that human reviewers are still needed to refine the reviews LLMs yield (Yu et al., 4 Nov 2025).

Automation has therefore become both an opportunity and a source of concern. A recent reframing argues that safety and assurance cases risk becoming detached from the understanding needed for responsible engineering and governance decisions, especially under pressures for increased tempo, reduced scrutiny, software complexity, and growing use of AI generated artefacts. To operationalize understanding, that work proposes two linked artefacts: an Understanding Basis, defined as a documented body of epistemic commitments, models, and evidence that forms the foundation for a justified, defensible and revisable human grasp of system behavior; and a Personal Understanding Statement, through which participants make their grasp explicit and challengeable. The same paper warns that automation may improve artefact production while weakening understanding (Bloomfield, 7 Apr 2026).

Current debate therefore centers less on whether Assurance 2.0 should use automation than on how automation should be constrained. Published discussions identify remaining challenges in confidence measurement, residual doubt management, automation support, the practical handling of defeaters and confirmation bias, system boundaries, and the lack of theories that would allow behaviours to be adequately assured, particularly for AI-intensive systems (Chen et al., 30 Sep 2025, Bloomfield et al., 7 Jan 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Assurance 2.0.