Strength Inconsistency Explanations
- Strength inconsistency explanations are formal methods that diagnose variability in model attributions under fixed assumptions and perturbations.
- They utilize metrics such as variance, inter-quantile width, and the Explanation Reliability Index to measure non-uniqueness and reliability across different models and settings.
- By revealing structural non-uniqueness, perturbation-induced instability, and stochastic training effects, these explanations guide enhancements in model design and interpretability.
Strength inconsistency explanations are explanatory formalisms in which the central object is not only an explanation itself, but the magnitude, direction, or structure of its variation under fixed modeling assumptions, perturbations, or updates. In recent work, this idea appears in intrinsically interpretable neural additive models, perturbation-based saliency analysis, explanation reliability metrics, dynamic quantitative bipolar argumentation, inconsistency-tolerant logical reasoning, and graded analyses of contextuality and Dutch Bookability. Across these settings, inconsistency strength is treated either as a reliability diagnostic, a signal of non-uniqueness, or the explanandum to be traced back to specific causes (Kim et al., 2024, Sengupta et al., 4 Feb 2026, Kampik et al., 21 Sep 2025).
1. Core idea and semantic scope
In neural additive models (NAMs), explanation inconsistency is the variability of learned shape functions or per-instance attributions across independently trained model instances that are equally accurate. With and centered attribution , inconsistency is explicitly defined across model instances as variation in or . In text models, the analogous notion is “strength inconsistency”: small, imperceptible perturbations can change attribution magnitudes, token ranks, or the salient-token set, even when the predicted label and confidence remain roughly unchanged. In dynamic QBAFs, the same phrase denotes a change in the partial order over final strengths of topic arguments after an update (Kim et al., 2024, Marjanović et al., 2024, Kampik et al., 21 Sep 2025).
The literature also uses graded strength scales. In political inconsistency detection, the spectrum is explicit: “Surface contradiction” is “the strongest degree of inconsistency,” “Factual inconsistency” is moderate, and “Indirect (Value) inconsistency” is weaker and more nuanced. In quantum foundations, the Abramsky–Brandenburger hierarchy yields strong, logical, and probabilistic contextuality, and these are mapped to correspondingly graded violations of subadditivity, additivity, and convexity, each producing a Dutch Book of a different strength. In probabilistic argument-strength modeling, strength is itself decomposed into conclusion probability and precision of the probability interval, with
so weakness can arise either from a low conclusion probability or from high imprecision (Sagimbayeva et al., 25 May 2025, Steeger et al., 2017, Pfeifer et al., 2017).
2. Quantification and metrics
The most direct formalization appears in BayesNAM. For a feature and input , if 0 is the set of attributions across trained NAM instances, inconsistency strength can be measured by the variance
1
the inter-quantile width
2
and sign disagreement through
3
BayesNAM extends the same logic to posterior samples, adding shape-function uncertainty
4
and credible intervals
5
Global indicators such as 6, 7, 8, and 9 summarize feature-level unreliability over the evaluation distribution (Kim et al., 2024).
A more general reliability formalization is the Explanation Reliability Index (ERI). If 0 is an explanation vector and 1 is a non-adversarial transformation drawn from 2, explanation drift is
3
Reliability is then
4
The framework instantiates this along four axes—small perturbations, redundancy collapse, model evolution, and distributional shift—yielding ERI-S, ERI-R, ERI-M, and ERI-D, and adds ERI-T for sequential models:
5
Large 6 means strong inconsistency; low ERI means low reliability (Sengupta et al., 4 Feb 2026).
Other literatures operationalize the same problem differently. In healthcare XAI, explanation consistency is defined as the complement of separability across training variations:
7
with binary-classifier separability
8
In removal-based explanations, the central quantity is interpretation error
9
supplemented by the truthful gap and the spectral distance
0
For perturbation-based saliency in NLP, robustness is summarized by average correlation and cosine similarity between clean and perturbed saliency vectors,
1
while predictive entropy, mutual information, and ECE characterize uncertainty and calibration (Watson et al., 2021, Zhang et al., 2022, Marjanović et al., 2024).
3. Why inconsistency arises
One influential explanation is structural non-uniqueness. In BayesNAM’s toy model, 2 is a strong single predictor while 3 are weakly correlated with the label but collectively powerful. When multiple features are predictive, many additive decompositions 4 can realize high accuracy, so one NAM run can place most explanatory mass in 5 while another places it in 6, even though predictions remain stable. The paper’s Case-II with 7, 8, 9, and 0 exhibits dramatically different shapes for 1 and 2 across seeds, while test accuracy remains above 3. BayesNAM’s feature-dropout theorem formalizes the complementary point that dropout discourages over-reliance on a single weakly correlated feature and pushes the model toward multiple plausible decompositions (Kim et al., 2024).
A second mechanism is perturbation-induced path instability. In NLP saliency maps, masking with [MASK]/[UNK] and obfuscation such as l33t perturb tokenization and embedding content abruptly, producing local nonlinearity, sharp shifts in gradient paths, and major reorderings of attribution magnitudes and token ranks. By contrast, realistic typos and synonym replacements preserve semantics and token identity sufficiently to keep gradients aligned, so explanations remain relatively stable. In removal-based XAI, inconsistency is even stronger: the “Impossible Trinity Theorem” states that interpretability, efficiency, and consistency cannot hold simultaneously for removal-based explanations. Exact fidelity under all masks conflicts with a single globally consistent interpretable surrogate, and practical methods therefore incur either inefficiency or inconsistency (Marjanović et al., 2024, Zhang et al., 2022).
A third explanation links inconsistency to stochastic training itself. In deep neural network training, inconsistency and instability of model outputs appear directly in an upper bound on the expected generalization gap. With 4, where 5 is inconsistency across runs on the same training set and 6 is instability across resampled training sets, the paper proves
7
Empirically, inconsistency is strongly predictive of generalization gap and is more reliable than sharpness across varied settings; algorithmic reduction of inconsistency improves performance and provides a theoretical basis for co-distillation and ensemble methods (Johnson et al., 2023).
4. Turning inconsistency into an explanatory signal
BayesNAM is the clearest case in which inconsistency is elevated from defect to diagnostic. It combines Bayesian neural networks with feature dropout, approximates the posterior over NAM parameters by Bayes by Backprop, and samples posterior shape functions and masks to obtain attribution distributions, credible intervals, and per-feature inconsistency strengths. High 8, 9, 0, or wide 1 indicate that the model class admits multiple plausible explanations in that region. On COMPAS, BayesNAM shows that juv_other_count has high variance in contributions and wide shape-function credible intervals beginning at counts 2, matching sparsity and skew in that region. On California Housing, high variance in Longitude between 3 and 4 reveals a structural limitation of a pure additive longitude effect; adding the interaction term Latitude 5 Longitude reduces RMSE from 6 to 7 and narrows the credible interval (Kim et al., 2024).
In quantitative bipolar argumentation, CE-QArg treats mismatches between actual and desired argument strength as counterfactual recourse problems. Given a focal argument 8 and target 9, it searches for a modified base-score function 0 minimizing an 1 distance
2
subject to a target band on 3. Its two core modules are polarity, which sets update direction from path parity, and priority, which scales updates by inverse shortest-path length. The framework proves sign conditions for difference quotients, alteration-existence results, and robust-validity transformations such as Nullified-Validity and Related-Validity (Yin et al., 2024).
A related but more general program appears in dynamic QBAFs. “Strength inconsistency explanations” are defined for changes in the partial order over topic arguments after an update. A set of changed arguments can be a sufficient strength inconsistency explanation (SSI), a counterfactual strength inconsistency explanation (CSI), or a minimally necessary strength inconsistency explanation (NSI), depending on whether its retained changes are sufficient, whether reverting it restores consistency, and whether it intersects every sufficient explanation. The paper proves that non-empty SSI, CSI, and NSI explanations exist if and only if the update produces strength inconsistency. “Strength change explanations” then generalize this idea from diagnosis to recourse: a strength change 4 is an explanation if changing the initial strengths of a mutable set 5 makes the modified graph satisfy a desired preorder, and inverse problems and strong counterfactual problems reduce to this setting (Kampik et al., 21 Sep 2025, Kampik et al., 26 Jan 2026).
5. Graded hierarchies outside model attribution
In political inconsistency detection, the explanatory problem is explicitly taxonomic. The benchmark defines a three-way classification task over statement pairs—{Unrelated, Consistent, Inconsistent}—and, for inconsistent pairs, a subtype prediction over {Surface contradiction, Factual inconsistency, Indirect (Value) inconsistency}. The taxonomy is ordered by severity: surface contradiction requires no external knowledge and is the strongest degree of inconsistency; factual inconsistency requires world knowledge; indirect inconsistency captures value divergence even when both statements could be true. The dataset contains 698 human-annotated pairs, and 237 pairs have at least one explanation, for 334 explanations in total. Models and humans approach the bootstrapped upper bound in the 3-class setting, but none reaches the upper bound for the 5-class subtype setting, especially on Factual and Indirect types (Sagimbayeva et al., 25 May 2025).
In inconsistency-tolerant logical reasoning, dialogue-based explanations make strength a dialectical property. The proposed approach translates maximal-consistent-subset semantics into structured argumentation with collective attacks, then computes explanations as dialectical proof trees. These trees expose the internal inference steps of arguments, the attacks and counter-attacks, the defence set 6, and the culprits 7. Strength is not numerical here; it is reflected in acceptance status. Credulous acceptance, grounded acceptance, and sceptical acceptance correspond to progressively stronger robustness against counterarguments, and the paper provides soundness and completeness theorems for dialogue trees satisfying defensive, finite, or ideal conditions (Ho et al., 16 Feb 2025).
In quantum foundations, strength is formalized as a hierarchy of formal incoherence. Strong contextuality yields maximal subadditivity violation, 8, corresponding to a finite null cover and a Dutch Book with sure-loss margin 9. Logical contextuality yields a non-maximal but strictly positive subadditivity defect, 0. Probabilistic contextuality yields failure of 1-convexity and additivity only upon monotone extension, and Dutch Books arise through separating linear functionals with margins scaling with the contextuality gap. The resulting explanation is explicitly graded: stronger contextuality implies stronger Dutch Bookability (Steeger et al., 2017).
6. Interpretation, practice, and limitations
The practical literature converges on a common prescription: explanations should be accompanied by explicit strength measures. BayesNAM recommends publishing 2 with credible bands 3 and global 4, and showing example-level attributions with error bars and the probability of positive attribution 5. ERI gives a bounded interpretive scale—6 as high reliability, 7 as moderate reliability, and 8 as low reliability—and emphasizes component-wise reporting across perturbation, redundancy, model evolution, distributional shift, and temporal axes. In perturbation-based NLP audits, robustness under realistic noise and calibration via ECE are recommended as deployment checks, especially when predictive and epistemic uncertainty are over-confident (Kim et al., 2024, Sengupta et al., 4 Feb 2026, Marjanović et al., 2024).
Several recurring misconceptions are explicitly rejected. High inconsistency strength does not necessarily mean that a feature is unimportant; in BayesNAM it can indicate that attribution is important but non-uniquely distributed among correlated features. Reliability is necessary but not sufficient: ERI notes that trivially invariant explainers can achieve 9 while remaining uninformative. Uncertainty is not a monotone proxy for explanation quality: in text models, high uncertainty does not necessarily imply low plausibility, and the correlation can be moderately positive when noise has been learned during training. Theoretical guarantees are also local to their assumptions: BayesNAM’s feature-dropout theorem is proved only for a simplified linear proxy under a toy data model, and CE-QArg’s guarantees depend on semantics satisfying directionality and monotonicity (Kim et al., 2024, Sengupta et al., 4 Feb 2026, Marjanović et al., 2024, Yin et al., 2024).
Deployment results underscore why these distinctions matter. In healthcare imaging, explanations from identical deep architectures trained under orthogonal variations have low consistency, approximately 0 on average, whereas kernel methods reach 1 explanation consistency, indicating that explanation instability can persist even when predictive performance is similar. In large reasoning models, efficient reasoning strategies such as No-Thinking and Simple Token-Budget consistently increase inconsistency across task settings, inconsistency between training objectives and learned behavior, and inconsistency between internal reasoning and self-explanations; correlated increases in withholding scores show that compressed reasoning verbalizes fewer influential factors. A plausible implication is that strength inconsistency explanations are becoming a general auditing interface: they quantify where an explanation is unstable, characterize the form of the instability, and sometimes identify the smallest structural or parametric change needed either to diagnose it or to reverse it (Watson et al., 2021, Yang et al., 24 Jun 2025).