Cross-model generalization of findings
Determine whether the empirical findings reported for SERA using the Qwen 3-32B base model and GLM-4.5-Air/GLM-4.6 teachers generalize to other language model families when evaluated thoroughly.
References
While we have some experiments with Claude 3.7 Sonnet and CLaude 4.0 Sonnet that hints at generalization of our method, we do not know whether our findings generalize to other model families when evaluated thoroughly.
First, although we cover dense Transformer, dense hybrid, and MoE architectures, all evaluated models belong to the Qwen family; whether the observed pruning patterns generalize to other model families remains to be studied.
Absolute accuracies would change with a stronger reader; whether the relative gains hold across reader and encoder families is left to future work.
We therefore state cross-backbone generalization as untested rather than answer it with an underpowered arm.
Fourth, all experiments use a single generative model (Qwen3-8B via Ollama in a capped, no-think configuration) and a single adjudicator (Llama 3). Generalization across model families is untested, and two consequences of these choices bear on the results: a capped, no-think configuration constrains reasoning depth and may compress differences between context qualities that a stronger generator would exploit, and the reviewed-correctness adjudicator is itself an unvalidated LLM, with no human agreement study performed against its labels.
It remains unclear whether the same findings generalize to substantially larger models or other model families.
And no significance test was run between the families, so the honest summary of this appendix is that the asymmetry, its conditioning, the restatement-trace collapse and the pipeline's lead over monolithic judging all reproduce on a second family, that single-note detection levels move with the family, and that a full cross-family comparison of the eight designs remains future work.
Several axes remain open. Loss-based selection across model families inherits the fit-to-text-versus-behaviour gap of \S\ref{sec:downstream} and still needs a downstream check.
Because a prior is plain text, one model's library could in principle be handed to another; such cross-model inheritance remains untested.
The ablations run three of the six conditions, so they cannot say whether the residue still sits in the partner notes on another model or at twenty episodes.