Determine whether excluded SAE features encode omitted topic semantics

Determine whether lower-validation SAE features excluded from the MonoTM descriptor vocabulary encode meaningful, independent topic semantics that are absent from the validated feature set.

Background

MonoTM estimates document–topic mixtures using all active SAE features but constructs topic descriptors only from features whose labels pass an LLM-based interpretability threshold. The authors audit lower-validation features that are highly associated with inferred topics and compare them with validated same-topic descriptors in document-activation space.

The audit finds that many excluded features appear to be covered by validated descriptors or capture auxiliary cues such as reporting format. However, the analysis does not establish that every excluded feature is unimportant or fully understood. Consequently, whether some excluded features represent independent, meaningful topic dimensions remains unresolved, particularly for corpus-specific regularities, pragmatic cues, formatting patterns, or culturally specific concepts that are difficult to summarize with short labels.

References

Overall, while this audit does not prove that no excluded feature ever captures a meaningful omitted semantic dimension, it alleviates the concern that MonoTM arbitrarily omits a large class of independent topic-defining features.

Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features  (2609.09575 - Joh et al., 9 Sep 2026) in Section 3.3.1, “Audit: do excluded features hide important topic semantics?”; see also Section 4, “Remaining uncertainty about excluded features”