Evaluate reinforcement learning and selective model merging for difficult exploit development

Determine whether applying post-supervised-fine-tuning reinforcement learning and selective MOPD-based model merging to the existing Feyospace checkpoints can improve difficult-task performance while preserving the capabilities recovered by the current data and supervision pipeline.

Background

The paper trains Feyospace checkpoints primarily through supervised fine-tuning on execution-verified coding, security, exploit-development, firmware, and device-backed trajectories. Although this improves vulnerability reproduction and some target-specific exploit capabilities, the evaluated checkpoints do not reach the generic-primitive or full-control tiers of the ExploitBench diagnostics.

The authors identify post-SFT reinforcement learning and selective model merging as prospective methods for advancing difficult-task performance. The unresolved issue is whether either intervention can improve these capabilities without undoing the target-domain behavior recovered through the current data and supervision pipeline.

References

These experiments will test whether post-SFT RL and selective model merging can improve difficult-task performance while preserving the capabilities recovered by the present data and supervision pipeline.

Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models  (2609.08418 - Li et al., 8 Sep 2026) in Section 6, “Future Work”

We also plan to strengthen the RL experts using broader training data, more accurate task-specific verifiers and rewards, and improved optimization algorithms, and to investigate whether iterative expert upgrading followed by repeated MOPD can transfer new capabilities without forgetting existing ones.

Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding  (2609.09300 - Qin et al., 8 Sep 2026) in Section Future Work