Evaluate reinforcement learning and selective model merging for difficult exploit development
Determine whether applying post-supervised-fine-tuning reinforcement learning and selective MOPD-based model merging to the existing Feyospace checkpoints can improve difficult-task performance while preserving the capabilities recovered by the current data and supervision pipeline.
References
These experiments will test whether post-SFT RL and selective model merging can improve difficult-task performance while preserving the capabilities recovered by the present data and supervision pipeline.
We also plan to strengthen the RL experts using broader training data, more accurate task-specific verifiers and rewards, and improved optimization algorithms, and to investigate whether iterative expert upgrading followed by repeated MOPD can transfer new capabilities without forgetting existing ones.