Evaluate safety-oriented behavior under depth compression

Investigate hallucination, refusal, and safety behavior in XMerge-compressed language models using instruction-tuned or safety-tuned variants, which were not evaluated for the base decoder-only models studied here.

Background

The paper evaluates XMerge primarily through downstream task performance, perplexity, and a limited calibration probe on one Llama-3-8B backbone. Although XMerge exhibits the smallest calibration degradation among the evaluated operators in that probe, the authors explicitly distinguish this result from a general safety guarantee.

The unresolved issue is whether depth compression with XMerge preserves safety-relevant behaviors, including hallucination control, refusal behavior, and other safety properties. The paper does not assess these properties because the experiments use base models rather than instruction- or safety-tuned variants, leaving their evaluation for future work.

References

We do not evaluate the remaining safety-oriented aspects (hallucination, refusal/safety behaviour), which require instruction- or safety-tuned variants rather than the base models studied here, and we leave them to future work.

— XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression  (2609.02083 - Hu et al., 2 Sep 2026) in Section 6, Discussion and Limitations, paragraph “Reliability checks”

We have not evaluated whether distillation preserves the safety behavior of the base models.

— PDMD: Projected Distribution Matching Distillation for Video Diffusion Models  (2609.35768 - Wang et al., 28 Sep 2026) in Ethics Statement