Develop multi-dimensional reward models for RL in native image generation
Develop reward models for reinforcement learning in native image generation that effectively capture and balance multiple dimensions, including image quality, instruction following, and alignment with human preferences, to enable reliable optimization of generation policies.
References
The key open problem is to develop reward models that can effectively capture and balance multiple dimensions, including image quality, instruction following, and human preference alignment.
There are avenues to explore alternate reward signals that might overcome drawbacks of CLIP reward (Kang et al., 2025), which we leave as future work.
This work leaves two directions open. First, our reward and acceptance rules focus on physical validity, relation consistency, and functionality. Incorporating richer aesthetic or style-level feedback could further improve visual quality, but may also increase evaluation cost.