Mechanism-level repair through outcome-aligned fine-tuning

Establish whether outcome-aligned fine-tuning of Qwen3-4B can repair the evidence-acquisition and temporal-updating mechanisms underlying research-attention forecasting, rather than merely improving end-to-end forecast scores.

Background

The paper fine-tunes Qwen3-4B on realised RAP outcomes and observes improved Forecast Spearman correlation on later, dependency-disjoint test fields. This demonstrates a learnable within-task signal and improves the overall forecasting policy.

However, the intervention does not resolve the diagnosed mechanisms: Search use and recent-state anchoring increase, while pooled residual-direction alignment and six-month revision-alignment gains do not show resolved improvement. The authors therefore leave open whether outcome supervision can specifically repair the acquisition and future-updating failures identified by the benchmark.

References

This establishes learnability across fields and later origins, but not a mechanism-level repair: Search use and recent-state anchoring increase, whereas pooled residual-direction and six-month revision-alignment gains remain unresolved.

RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases  (2609.10092 - Wu et al., 9 Sep 2026) in Section 4.5, “Outcome Supervision Provides a Learnable Within-Task Signal”