---
title: Biconditional Correctness Criterion
url: https://www.emergentmind.com/topics/biconditional-correctness-criterion
type: topic
---

# Biconditional Correctness Criterion

The biconditional correctness criterion provides an "if and only if" (biconditional) formalization of system or artifact correctness, ensuring that observed effects and declared approvals are in precise correspondence. It emerges independently in agent runtime security and in mathematical logic (notably, logic programming and linear logic), where it serves as a definitive test for the agreement between implementation behavior and specification, between proof structure and sequent derivability, or between skill execution and approved operational effects. Recent work (“Skills as Verifiable Artifacts…” [2605.00424]) treats it as a linchpin in trust schemas for agent skills, while in logic it underpins classic notions of program realized models and proof net sequentialization [1412.8739, 2506.21678].

## 1. Formal Definitions in Agent Runtimes and Logic

In agent runtimes, the biconditional criterion is defined over an agent run starting with corpus state $s_0$ and ending in $s_1$. Let $D = \delta(s_0, s_1)$ be the multiset of observed side-effects ("delta"), and let $S = \{ r \in L \mid r.\mathrm{type} = \mathit{irreversible.executed}\ \wedge\ r.ok = \mathit{true} \}$ be the set of audit log entries corresponding to approved and executed irreversible actions. The audit log $L$ passes the biconditional criterion if and only if the multiset projection of $D$ onto $(\mathrm{operation},\ \mathrm{target})$ equals that of $S$:
$$
\mathrm{proj}_{(op,\,t)}(D) = \mathrm{proj}_{(op,\,t)}(S)
$$
This ensures, simultaneously, that every actual side-effect is matched by an approval and vice versa—no ghost effects, and no illusory (unrealized) approvals [2605.00424].

In logic programming, the biconditional takes the form $\forall A \in HB.\ [P \models A \leftrightarrow A \in S]$, i.e., every provable atom in $P$ must be in the intended model $S$ and vice versa. This is equivalently expressed as $M_P = S$, where $M_P$ is the least Herbrand model [1412.8739].

In linear logic, the biconditional criterion characterizes proof-structure correctness via necessary and sufficient conditions: for suitable fragments, a proof structure $R$ is sequentializable if and only if its correctness graph satisfies (i) all switchings are acyclic, (ii) the number of connected components matches $|w|+1$, and (iii) a geometric $(–w)$-condition preventing erasing subnets from connecting to tensor/cut nodes [2506.21678].

## 2. Underlying Assumptions and Key Definitions

The agent runtime formulation is predicated on the explicit separation of observed reality ($D$) and the system’s claim of action ($S$). Skills are tuples $(M, \mathrm{content}, \sigma)$, where $M$ is a manifest assigning capabilities, a quoted verification level (unverified, declared, tested, formal), and other metadata. A verification procedure must pass the biconditional on adversarial exercises to elevate a skill to “tested” [2605.00424].

Logic programming places the criterion in the context of model-theoretic semantics: $S$ is the specification, $M_P$ is the program’s least model. Correctness ($M_P \subseteq S$) and completeness ($M_P \supseteq S$) correspond precisely when $M_P = S$ [1412.8739].

Linear logic proof-nets use correctness graphs, switchings, and erasing nodes connecting the formalization to geometric graph properties—in particular, preventing the formation of “erasing threads” that violate syntactic proof properties [2506.21678].

## 3. Rationale and Distinction from Soundness/Completeness

A biconditional criterion transparently unites “soundness” and “completeness.” Soundness ensures that all claimed (approved) actions or proofs actually occur or hold (no false positives); completeness ensures all real effects or legitimate derivations are claimed (no false negatives). The biconditional is unique in catching both failure types in runtime logging (e.g., unauthorized changes or missing actions), as well as in logic (missing or spurious derivations) [2605.00424, 1412.8739].

This is operationalized in agent skill verification by forbidding “gate bypasses” (altering state without audit) and “audit forgeries” (claiming to execute, but nothing happens). In logic, it eliminates programs or proofs that undershoot (failure to derive intended facts) or overshoot (deriving unintended facts) the specification/model.

## 4. Verification Workflows and Evaluation Methodologies

In agent systems, adversarial ensemble exercises are used: agents (e.g., Cleaners, Auditors, Critics) propose destructive or risky actions on a fixed corpus, driving the runtime under offensive input conditions to expose mismatches between $D$ and $S$. Pass/fail is binary and deterministic. The biconditional’s mechanical verification is the acceptance test for a skill to reach “tested” level and bypass human-in-the-loop (HITL) gating [2605.00424].

In logic programming, the biconditional is evaluated by showing $S \models P$ (correctness) and $S \subseteq M_P$ (completeness), with sufficient conditions including “recurrent coverage,” semi-completeness with recurrence, or acceptable level mappings. Pruned SLD-trees (csSLD-trees) require compatibility to maintain biconditional correctness under restricted computation [1412.8739].

In linear logic, combinatorial/graph-theoretic algorithms (Danos–Regnier connectivity, $(–w)$-conditions) are checked on proof-structure graphs. Only when all switchings satisfy the necessary and sufficient conditions, and the geometric constraints on erasing nodes are met, does sequentialization (i.e., existence of a corresponding sequent proof) hold [2506.21678].

## 5. Failure Modes and Detection

Enumerating possible failures is essential for operational assurance:

| Failure Mode                  | Manifestation                                                | Detected by Biconditional |
|-------------------------------|-------------------------------------------------------------|---------------------------|
| Gate bypass                   | $D \not\subseteq S$: side-effect with no log                | Yes                       |
| Audit forgery                 | $S \not\subseteq D$: log entry with no actual change        | Yes                       |
| Silent host failure           | Approved log, no state change                               | Yes                       |
| Wrong-target execution        | Log refers to $t_1$, effect observed on $t_2 \ne t_1$       | Yes                       |

In all cases, $\mathrm{proj}_{(op,t)}(D) \ne \mathrm{proj}_{(op,t)}(S)$, causing an immediate failure of the biconditional check. This yields high assurance in both sandboxed skill certification and runtime operation, elevating security and trust in automated or semi-automated agent systems [2605.00424].

## 6. Comparative Structures in Logic Programming and Linear Logic

The biconditional criterion in logic programming, $M_P = S$, has a close parallel to the model-equality verification in runtime systems. Approximate specifications ($S_{\it comp} \subseteq M_P \subseteq S_{\it corr}$) further allow partial biconditionality where exact specification is impractical [1412.8739].

In linear logic proof-theory, sequentializability becomes an “iff” property only when both Danos–Regnier acyclicity/connectivity and the geometric $(–w)$-condition are imposed. The $(–w)$-condition ensures that erasing subnets cannot attach in a way that would escape the control of the sequentializable proof calculus, structurally guaranteeing the biconditional property for wide fragments of MELL [2506.21678].

## 7. Integration with Trust Schemas and Broader Implications

In agent skill ecosystems, the biconditional criterion anchors the trust schema: only those skills that pass adversarial ensemble evaluation and the biconditional, achieving “tested” or “formal” verification status, are permitted to bypass continuous HITL gating. This approach both eliminates rubber-stamping risk and provides a reproducible, mechanical benchmark for skill certification. Failures cause session aborts for “trusted” skills and immediate operator alerts [2605.00424].

In logical and proof-theoretic contexts, biconditional criteria provide the foundation for both theoretical completeness results and practical verification workflows for programs and proof artifacts [1412.8739, 2506.21678].

A plausible implication is that biconditional correctness frameworks reveal and tightly capture the boundary between trusted automation and operator mediation, enforceable through both runtime instrumentation and formal methods in diverse computer science subfields.

Source: https://www.emergentmind.com/topics/biconditional-correctness-criterion