Design reliable rubric generator and verifier for large-scale instruction-following training
Design a reliable rubric generator and a reliable rubric verifier for large-scale reinforcement learning pipelines aimed at improving large language models’ instruction-following capabilities, where the generator synthesizes rubrics for each user prompt from raw training data and the verifier determines whether a given model response satisfies each rubric criterion, so that dependable rubrics and judgments can be provided for training when human labeling is impractical.
References
How to design a good generator and verifier to provide reliable rubrics and judgments for training is still an open problem.
Two open problems follow from the exploitation gap. First, training verifiers calibrated for open-ended evaluation: $\hat{\rho}_v$ provides a measurable target.