Determine overlap of failure clusters across methods

Determine whether the failure clusters discovered by Teach-to-Crash, PAFOT, and ChatScene are shared across methods or represent disjoint regions of the executable scenario space.

Background

The evaluation measures within-run Jaccard novelty, the number of failure clusters, and the unique-failure ratio. These metrics show how many structurally distinct failures each method discovers internally but do not compare the identities of clusters discovered by different methods.

Cross-method cluster overlap would reveal whether the methods find the same failure modes with different frequencies or explore substantially different regions of the scenario space.

References

These summaries do not establish whether failure clusters are shared or unique across methods, an analysis left for future work (Section~\ref{sec:discussion}).

— Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation  (2609.27296 - Ghazal et al., 23 Sep 2026) in Section 6.3, RQ3: Diversity of Collision Scenarios