Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy (2403.04283v1)

Published 7 Mar 2024 in cs.CL, cs.AI, and cs.LG

Abstract: Reinforcement Learning from Human Feedback (RLHF) is the prevailing approach to ensure LLMs align with human values. However, existing RLHF methods require a high computational cost, one main reason being that RLHF assigns both the generation and alignment tasks to the LLM simultaneously. In this paper, we introduce Proxy-RLHF, which decouples the generation and alignment processes of LLMs, achieving alignment with human values at a much lower computational cost. We start with a novel Markov Decision Process (MDP) designed for the alignment process and employ Reinforcement Learning (RL) to train a streamlined proxy model that oversees the token generation of the LLM, without altering the LLM itself. Experiments show that our method achieves a comparable level of alignment with only 1\% of the training parameters of other methods.

PDF HTML Abstract

Summarize Bookmark Chat (Pro)

References (19)

Authors (11)

Yu Zhu (123 papers)
Chuxiong Sun (12 papers)
Wenfei Yang (18 papers)
Wenqiang Wei (5 papers)
Bo Tang (111 papers)
Tianzhu Zhang (60 papers)
Zhiyu Li (69 papers)
Shifeng Zhang (46 papers)
Feiyu Xiong (53 papers)
Jie Hu (187 papers)
Mingchuan Yang (10 papers)

Citations (3)

View on Semantic Scholar

Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy (2403.04283v1)

Related Papers