Characterize REALLY-KNOW Specifications as Causal Graphs

Fully characterize the validity and applicability of REALLY-KNOW specifications as causal graphs for estimating concept-level causal effects in vision-language models.

Background

DiaVLo uses the relational structure of REALLY-KNOW specifications as a proxy for a causal graph and applies refutation tests to assess the resulting causal model and estimands. The authors report mixed or inconclusive outcomes: structural refutation tests succeed only when enough counterfactual samples are available, and placebo refutation tests are inconclusive.

The paper consequently leaves unresolved whether and under what conditions REALLY-KNOW specifications provide reliable causal graphs for multimodal model analysis. Resolving this requires a fuller theoretical and empirical characterization of the assumptions, identifiability, and robustness of this representation.

References

Nonetheless, further research is required to fully characterise the use of REALLY-KNOWs as causal graphs.

— DiaVLo: Diagnosing Behaviours of Vision-Language Models  (2609.22008 - Corti et al., 18 Sep 2026) in Section 5, “Results & Discussion,” subsection “REALLY-KNOWs as Causal Graphs”