---
title: Question Completeness in Formal Systems
url: https://www.emergentmind.com/topics/question-completeness
type: topic
---

# Question Completeness in Formal Systems

Question completeness denotes whether a formalism, analysis, dataset, or answer artifact contains enough structure to settle the questions posed within a specified language, semantics, or use case. Across the cited literature, the notion is explicitly context-relative rather than uniform: in query answering it means that returned answers coincide with all answers that exist in reality; in logic it means that a formula or proof system decides every statement semantically fixed by it; in static analysis it means that an abstraction or algorithm finds the required invariant whenever one exists; and in attributed question answering it means that a gap-free deductive path connects the question to the final answer [1604.08377] [1411.3015] [1605.01004] [1110.1614] [2211.09572] [2601.15050]. This suggests that question completeness is best understood as a family of domain-indexed completeness guarantees.

## 1. Core structure of completeness claims

A recurrent starting point is the distinction between **soundness** and **completeness**. In the logical formulation recalled by Monniaux, a proof system is sound with respect to a semantics if it never proves a formula that is false in that semantics, and complete if it can prove every formula that is true in that semantics; a decision procedure is both sound and complete [2211.09572]. The same two notions reappear in static analysis, query answering, modal logic, and proof calculi, but the relevant objects vary: domains, algorithms, formulas, queries, and models.

A second recurrent distinction is between **global** and **local** completeness. Monniaux separates completeness of an abstraction/domain from completeness of the analysis procedure that searches for invariants [2211.09572]. Drabent distinguishes completeness of a logic program with respect to a specification from completeness **for a query** \(Q\), where every ground instance required by the specification must be produced as an answer [1411.3015]. Achilleos studies completeness of a **single modal formula**, not of the ambient proof system, by asking whether that formula decides every formula over its own propositional vocabulary [1605.01004].

A third distinction concerns **binary** versus **quantitative** notions. In scenario-based testing for automated driving, completeness of a scenario concept is binary—“a scenario concept is sufficiently complete for a use case if all relevant driving situations are adequately captured”—whereas coverage is “the quantifiable extent to which a set of scenarios or parameters represent a defined ODD or predefined set of scenarios” [2404.01934]. LogicScore adopts the same binary pattern for long-form QA: completeness is \(1\) exactly when a minimal sufficient reasoning path exists, and \(0\) otherwise [2601.15050]. This suggests that completeness often marks a threshold property, while related notions such as coverage, conciseness, or degree of answer loss provide graded refinements.

## 2. Proof-theoretic and semantic completeness

In modal logic, a formula \(\varphi\) is complete for a logic \(l\) when for every \(\psi \in L(P(\varphi))\), either \(\vdash_l \varphi \rightarrow \psi\) or \(\vdash_l \varphi \rightarrow \neg \psi\) [1605.01004]. Model-theoretically, this is equivalent to saying that all pointed \(l\)-models satisfying \(\varphi\) are bisimilar modulo \(P(\varphi)\). The associated decision problem has the same complexity as validity, except for trivial cases: for \(K\), \(K4\), \(D4\), and \(S4\) it is PSPACE-complete, while for \(KD45\), \(S5\), and more generally \(l+5\), it is coNP-complete; for \(KD\) and \(T\) with nonempty propositional vocabulary, satisfiable complete formulas do not exist in general [1605.01004].

For intuitionistic first-order logic, the relevant semantic notion is **uniform validity** under BHK semantics. The main theorem for minimal logic states that for any \(\psi \in MF\),
\[
[\forall M:S.\, M \models \psi] \;\Leftrightarrow\; \vdash_{ML}\psi,
\]
and Friedman’s \(A\)-transformation lifts this to intuitionistic first-order logic:
\[
[\forall M:S.\, M \models \psi^A] \;\Leftrightarrow\; \vdash_{IL}\psi.
\]
The procedure \(\mathsf{Prf}\) converts a uniform evidence term into a formal proof, and the strongest termination argument uses the Fan Theorem [1110.1614]. Here question completeness is not ordinary validity over all models, but the existence of a **single** evidence term inhabiting the intersection \(\bigcap_{M\in S} M(\psi)\).

Graphical calculi use the same syntax–semantics schema. Backens shows that the ZX-calculus is complete for pure state stabilizer quantum mechanics and for the single-qubit Clifford+T group [1602.08954]. Jeandel, Perdrix, and Vilmart strengthen this line by proving completeness for the Clifford+T fragment, for linear diagrams with Clifford+T constants, and—after adding axiom (A)—for the unrestricted ZX-calculus as a whole [1903.06035]. In these settings, completeness means that every semantically valid equality between diagrams is derivable graphically. A plausible implication is that question completeness in proof calculi is the strongest possible form of deductive adequacy: no semantically valid question about equality remains outside the rewrite system.

## 3. Query-answer completeness in data systems

In incomplete database theory, question completeness is formulated directly as equality between answers over available and ideal data. An incomplete relational database is a pair \(D=(D^i,D^a)\) with \(D^a \subseteq D^i\), and a query completeness statement \(\Compl(Q)\) is satisfied exactly when
\[
Q(D^a)=Q(D^i)
\]
under the chosen semantics [1411.2855]. The thesis develops metadata-based reasoning from table completeness statements \(\Compl(R(\bar{t});G)\) to query completeness, reduces key entailment problems to query containment, extends the framework to null values, RDF, spatial data, and process models, and shows that restricted compactness and design-time/runtime verification can be obtained in process settings [1411.2855].

For RDF under the open-world assumption, completeness is handled via explicit completeness statements \(\compl{P_C}\) and entailment \(\mathcal{C},G \models \compl{Q}\), meaning that every admissible completion \(G'\) of the graph that respects \(\mathcal{C}\) yields the same answers for \(Q\) as \(G\) itself [1604.08377]. The paper defines a transfer operator \(T_{\mathcal{C}}(G)\), shows that the general completeness-entailment problem is \(\Pi^P_2\)-complete, and identifies a practically important fragment of **SP-statements** of the form \(\compl{\{(s,p,?v)\}}\) that supports scalable indexing on Wikidata-scale graphs [1604.08377]. Here question completeness is local and entity-centric: one does not close the world globally, but only on the subject–predicate regions explicitly declared complete.

Logic programming gives an especially direct query-level formulation. Drabent defines completeness for a query \(Q\) with respect to a specification \(S\) by requiring that for every ground instance \(Q\theta\),
\[
S \models Q\theta \;\Rightarrow\; Q\theta \text{ is an answer for } P.
\]
Coverage, semi-completeness, recurrence, and acceptability provide sufficient conditions, while pruning via cut is analyzed using csSLD-trees and adjustable coverage [1411.3015]. In this formulation, question completeness means that the program produces **all** answers required by the specification for the question actually asked.

Ontology-based data access sharpens the same idea against incomplete reasoners. Glimm, Horrocks, Lutz, and Sattler define a reasoner to be \((Q,T)\)-complete if, for every ABox \(A\), it detects unsatisfiability whenever \(T \cup A\) is unsatisfiable, and otherwise returns all certain answers to \(Q\) with respect to \(T\) and \(A\) [1401.4604]. Their framework introduces abstract reasoners, monotonicity and faithfulness conditions, and finite test suites exhaustive for \((Q,T)\)-completeness under suitable assumptions. For UCQ-rewritable cases, full or injective instantiations of rewritings yield finite certification suites; for more general recursive settings, first-order reproducible reasoners and datalog\({}^{\pm,\vee}\) rewritings recover positive results, although strong impossibility theorems show that no finite suite can work in full generality [1401.4604]. This is one of the clearest operational readings of question completeness: a reasoner may be incomplete in principle yet complete for the specific questions and ontology used by an application.

## 4. Operational completeness under approximation

In abstract interpretation, Monniaux distinguishes several levels of completeness: exactness of the abstraction for a specific problem, completeness of the analysis method with respect to a chosen abstract domain, and decidability of the invariant-existence problem for that domain [2211.09572]. A method is complete with respect to its domain when it “would compute always such invariants if they exist.” For finite-height domains this can coincide with ordinary ascending fixpoint iteration; for richer settings, exact solving via quantifier elimination or policy iteration may recover method-level completeness. At the same time, widening is identified as a major source of incompleteness: in the circular-array interval example the ideal interval invariant \([0,42]\) exists in the domain, yet widening plus narrowing yields \([0,999]\), which does not prove the assertion [2211.09572].

Monniaux also exhibits **problem-specific completeness**. In simplified LRU cache analysis, a chain of exact abstractions—addresses only, then per-set decomposition, then antichain abstraction for “block \(a\) is present and the set of blocks younger than \(a\) is \(P\)”—produces a complete/exact analysis of cache behavior with respect to the control-flow-only program model [2211.09572]. This suggests that question completeness in analysis is often modular: exactness can be achieved for a subsystem even when global completeness for full program semantics is abandoned.

Scenario-based testing for automated vehicles frames the same issue in an open context. A scenario concept is sufficiently complete for a use case if all relevant driving situations are adequately captured, whereas coverage is the quantifiable extent to which scenarios or parameters represent the ODD or a predefined scenario set [2404.01934]. The methodology uses Goal Structuring Notation, counter-hypotheses, knowledge-based evidence, and data-driven evidence. In the inD case study, rules derived from the concept detected **59,253 base scenarios**, and “there is no second in the recordings where no base scenario is assigned” at the chosen abstraction level [2404.01934]. The paper therefore argues not for absolute completeness of an open traffic world, but for **sufficient completeness** relative to a use case, an ODD, and an abstraction boundary.

## 5. Completeness in attributed question answering

LogicScore operationalizes question completeness for long-form attributed QA as a property of the reasoning chain rather than of isolated factual statements. Given a question \(\mathcal{Q}\), retrieved documents \(\mathcal{D}\), a long-form answer \(\mathcal{LA}\), and a short answer \(\mathcal{SA}\), the framework transforms \(\mathcal{LA}\) into atomic propositions \(\mathbb{P}=\{P_1,\dots,P_n\}\) and then applies backward verification from the gold answer entity \(e_{gold}\) to the question entity \(e_q\) [2601.15050]. If a connected chain \(\xi\) exists, it is the minimal sufficient set \(\mathbb{P}_{min}\); otherwise \(\mathbb{P}_{min}=\emptyset\). Completeness is then defined by
\[
\mathrm{Completeness}=
\begin{cases}
1 & \text{if } \mathbb{P}_{min}\neq \emptyset,\\
0 & \text{otherwise.}
\end{cases}
\]

This formulation is explicitly global. It is not enough that individual sentences be supported by citations; the propositions must form a gap-free deductive path from question to answer. Conciseness and Determinateness are then layered on top:
\[
\mathrm{Conciseness}=\frac{|\mathbb{P}_{min}|}{|\mathbb{P}|},
\qquad
\mathrm{Determinateness}=\mathbb{I}(\hat{\mathcal{SA}}\equiv\mathcal{SA}).
\]
The empirical results expose a large attribution–reasoning gap. Over more than 20 LLMs and three multi-hop datasets, leading systems can achieve high attribution precision while remaining weak on global logic; the paper reports **92.85\% precision for Gemini-3 Pro** but only **35.11\% Conciseness** [2601.15050]. Completeness on MusiQue remains around 60% even for top proprietary models, with Gemini-3-Pro at **59.21**, GPT-5.1 at **60.11**, GPT-o3 at **60.83**, and Claude-4.5 at **63.88** [2601.15050]. Human evaluation yields a Jaccard similarity of about **94.39\%** between LogicScore completeness labels and human labels [2601.15050]. In this setting, question completeness is a binary structural property of reasoning continuity, not a synonym for citation correctness.

## 6. Open theories, causal models, and the limits of completeness

In the philosophy of physics, the paper on quantum theory defines completeness of a physical theory through the existence of a **formal (continuous) causal model** whose laws are complete, consistent, and reality conformal [1512.08720]. A formal causal model is a collection of conditional state-transition rules
\[
L_i:\ \text{IF } c_i(s_0)\ \text{THEN } s_1=f_i(s_0),
\]
possibly using \(\text{RANDOM}(\text{valuerange},\text{probabilitydistribution})\) for nondeterministic transitions. A physical theory is complete if such a model can be constructed for the theory [1512.08720].

Within that framework, “question completeness” becomes the demand that every physically meaningful question about state evolution, including measurement, be answerable by the model using only theory-internal predicates and operations. The paper argues that standard quantum theory is not complete in this sense because the measurement problem, the interference-collapse rule, and the lack of a fully specified causal account of interactions prevent construction of a complete formal causal model [1512.08720]. Questions such as exactly when a measurement occurs or exactly when interference is lost cannot be mapped to the \(c_i,f_i\) schema without extra-theoretical notions such as “observer” or “capable of determining” [1512.08720].

A plausible implication, reinforced by the static-analysis and scenario-testing literature, is that question completeness is frequently bounded by undecidability, open-world assumptions, or irreducible modeling choices. Monniaux notes undecidability of invariant existence for sufficiently rich abstract domains and leaves the convex-polyhedral case as an open problem [2211.09572]. Scenario-based testing replaces absolute completeness by sufficiently complete concepts relative to an ODD [2404.01934]. Quantum theory, in the analyzed formulation, lacks a fully internal causal account of some of its own central questions [1512.08720]. Across these cases, completeness remains a target of formalization, but not always an attainable global property.

Source: https://www.emergentmind.com/topics/question-completeness