- The paper introduces MERIT, a training-free layer-selective merging framework that restores temporal reasoning in VLMs while preserving visual perception.
- It employs an evolutionary search to optimize layer-specific self-attention merging recipes, effectively grafting causal reasoning from text LLMs.
- Empirical evaluations demonstrate notable TR improvements (up to +23.9%) across multiple benchmarks, with consistent performance on out-of-distribution tasks.
Restoring Temporal Reasoning in Video-LLMs through Layer-Selective Merging
Motivation: Degradation of Temporal Reasoning in Multimodal Adaptation
Video-LLMs (VLMs) are typically constructed by attaching a visual encoder to a pretrained LLM. While this multimodal adaptation bolsters perceptual performance, it often compromises temporal reasoning (TR) capacities that originate from the language backbone. Two forms of reasoning degradation are identified: (1) VLMs frequently fail TR tasks in pure text that the unimodal LLMs solve correctly; (2) VLMs, although visually competent, are unable to infer causal event sequences in videos, revealing a clear decoupling between perception and TR. This phenomenon is attributed to current adaptation regimes emphasizing object-level perception with insufficient temporal abstraction and causal supervision.

Figure 1: Multimodal adaptation can erode intrinsic temporal reasoning—base LLMs answer text TR tasks correctly, but their VLM counterparts do not; on video tasks, VLMs detect objects but miss event causality.
MERIT: Task-Driven Layer-Selective Model Merging
The paper introduces MERIT, a training-free model merging framework designed to restore TR in VLMs by leveraging their paired text-only backbones. Unlike prior approaches—with full-model or random layer parameter averaging—MERIT conducts an evolutionary search over possible layerwise self-attention "merging recipes," optimizing an objective that balances TR improvements with constraints on temporal perception (TP) degradation. Layer selection and interpolation weights are thus optimized to graft reasoning-related mechanisms from the text LLM into the VLM’s backbone while minimizing perceptual interference.

Figure 2: MERIT framework overview: evolutionary search identifies a layer-selective self-attention merging recipe, rewarding TR gains and penalizing TP loss to achieve a merged model with improved TR and preserved perception.
Empirical Evaluation and Numerical Results
MERIT is evaluated on three representative VLMs—LongVA-7B, InternVL3-8B, and Qwen3-VL-4B—across a diverse suite of challenging video benchmarks. Relative TR improvements on Video-MME are significant: +23.9% (LongVA-7B), +10.8% (InternVL3-8B), +3.8% (Qwen3-VL-4B), with no perceptual regression (TP maintained or improved). Critically, MERIT’s discovered recipes generalize robustly beyond the search set, delivering TR gains on multiple OOD benchmarks (LongVideoBench, LVBench, MMBench-Video, Video-Holmes), confirming that the approach is not narrowly overfitting.
Three claims are substantiated:
- Selective layer merging is necessary for reliable reasoning recovery—uniform full-model or random-k layer merging is consistently less effective and often degrades perception.
- Layer selectivity is not arbitrary—the specific layers chosen by MERIT are consistently responsible for TR, as proven by interventional masking.
- MERIT’s merging shifts temporal grounding behavior—models shift away from misleading, locally salient cues and toward integration of temporally and causally relevant video evidence.
Interventional and Attributional Analysis
The authors conduct controlled layer-masking experiments, demonstrating that ablating the MERIT-selected layers in the VLM induces a substantially sharper decline in reasoning accuracy than in overall or perception tasks. For instance, masking these layers in LongVA-7B reduces reasoning by up to 51.7%, contrasted with only 16.8% for overall accuracy.

Figure 3: Layer-masking on LVBench—masking MERIT-selected layers in the base VLM causes sharply disproportionate reasoning degradation, establishing their criticality.
Frame-level attribution via gradient-activation products evidences that merged models rely less on spurious local events and more on temporally distributed, causally informative cues to generate final answer tokens and reasoning traces. Case studies in complex video tasks show that MERIT enables causal inference even when decisive evidence is neither contiguous nor explicitly depicted, and that MERIT-integrated models better abstract over event repetition or chain multi-step dependencies.

Figure 4: MERIT links temporally distant, causally informative evidence to provide correct causal reasoning, whereas the base model anchors on a misleading local event.

Figure 5: Both models predict the dream-state, but only MERIT comprehensively chains relevant temporal events, demonstrating superior multi-event reasoning.

Figure 6: The base model latches on an isolated unlocked door, while MERIT integrates a full event sequence to correctly attribute the cause of death.

Figure 7: MERIT captures recurrence and cyclic event structure, in contrast to the base model’s focus on single-instance transitions.
Theoretical and Practical Implications
MERIT directly challenges the notion that perceptual competence and reasoning in VLMs are irreconcilably entangled across all model layers. The findings establish that temporal reasoning resides disproportionately in a localized subnetwork within the backbone, such that reasoning can be restored scalably by targeted, perception-aware parameter merging—without retraining or additional supervision.
Practically, this facilitates post-hoc adaptation of VLMs: one can efficiently recover lost reasoning capacities with trivial additional compute, preserving (or even boosting) perception, and adapting to new LLM or VLM architectures as backbones evolve. The method’s success on diverse benchmarks suggests a pathway to more modular, capability-driven multimodal model construction.
Theoretically, MERIT’s results reinforce recent proposals that model layers are functionally heterogeneous and that reasoning “circuitry” can be isolated or reconstituted via architectural surgeries. These findings encourage further research in mechanistic interpretability, automated model editing, and compositional transfer learning—toward the automated recovery, enhancement, or control of emergent behaviors in LLM and VLM systems.
Future Prospects
Potential avenues include extending MERIT’s recipe search to settings with weaker or noisier capability supervision, richer merging parameterizations (e.g., continuous interpolation beyond discrete gating and fixed weights), and systemic analyses for identifying generally transferable reasoning-critical substructures. The model specificity of current recipes prompts investigation into universal principles for reasoning recovery applicable across architectures.
Conclusion
MERIT provides a principled, training-free framework for restoring temporal reasoning in VLMs via targeted, layer-selective model merging. By optimizing parameter interpolation at the self-attention layer level under explicit TR and TP objectives, MERIT consistently improves reasoning, preserves perception, and fundamentally clarifies the architecture-behavior relationship underlying reasoning in multimodal transformers. The results motivate model surgery as a general strategy for efficient, interpretable, and customizable capability recovery in future AI systems.