Close the implementation, replay, semantic, generalization, and causal-evidence gaps

Establish Java–Lean refinement, complete-trace replay, semantic completeness, industrial generalization, and causal attribution for the typed flexible-arity slotted e-graph construction and its Alloy evaluation pipeline.

Background

The paper distinguishes its abstract Lean-checked certificate and finite-unfolding results from the Java implementation and the bounded Alloy experiments. The implementation was not shown to refine the formal construction, the corpus was not independently replayed through complete exported traces, and the bounded evaluation does not establish semantic completeness or broader industrial validity. The authors also report that the measurements do not isolate causal sources of the observed performance and behavior.

These limitations are explicitly presented as unresolved rather than merely as scope restrictions: Java–Lean refinement, full-trace replay, semantic completeness, industrial generalization, and causal attribution are all identified as open. Resolving them would connect the formal specification to the concrete producer and strengthen the empirical claims beyond bounded structural-consolidation evidence.

References

Consequently, the evaluation supports bounded structural-consolidation and transformation-recognition observations only; Java--Lean refinement, full-trace replay, semantic completeness, industrial generalization, and causal attribution remain open.

Typed Flexible-Arity Slotted E-Graphs: A Soundness Construction and an Alloy Case Study  (2609.03998 - Wu et al., 3 Sep 2026) in Section 5, paragraph “Independent checks and limits”