Papers
Topics
Authors
Recent
Search
2000 character limit reached

HC-GRPO: Cost-Aware Group Policy Optimization

Updated 28 December 2025
  • The paper introduces HC-GRPO, an RL algorithm that uses group-relative performance to eliminate the need for a learned value critic in cost-aware policy adaptation.
  • HC-GRPO integrates heterogeneous costs from navigation, queries, and memory retrieval to balance operational efficiency in high-dimensional, partially observable environments.
  • Empirical outcomes show HC-GRPO reduces task costs and improves success rates, outperforming standard PPO-based methods in simulated embodied search tasks.

HC-GRPO (Heterogeneous Cost-Aware Group Relative Policy Optimization) is a reinforcement learning (RL) algorithm designed for optimizing multimodal LLM (MLLM) agents engaged in complex embodied search tasks. Unlike traditional Proximal Policy Optimization (PPO), HC-GRPO operates by grouping trajectory rollouts per instruction, exploiting relative performance among these rollouts to eliminate the necessity of a learned value critic. This mechanism facilitates efficient, cost-aware policy adaptation, focusing on the optimal navigation of heterogeneous operational costs—including physical movement, social interaction via queries, and cognitive memory retrieval—within high-dimensional, partially observable environments (Zhou et al., 21 Dec 2025).

1. Problem Formulation

HC-GRPO addresses the challenge of reasoning under ambiguous instructions by integrating heterogeneous actions and their associated costs into a unified RL framework. The state space is implicitly defined via the MLLM’s internal context vector; the multimodal history at timestep tt is ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1}). The action space A\mathcal{A} comprises:

  • Navigate⁡(ℓ)\operatorname{Navigate}(\ell): Physical relocation to location ℓ\ell,
  • Ask⁡(q)\operatorname{Ask}(q): Clarification question to the user,
  • GetMemory⁡(k)\operatorname{GetMemory}(k): Episodic memory retrieval,
  • Found⁡(target)\operatorname{Found}(\text{target}): Terminal "I found it" action.

The cost function C(at)C(a_t), reflecting the heterogeneity of physical, cognitive, and social acts, is as follows:

C(at)={cnav d(pt,pt+1)if at=Navigate, cask(1+αNask(t))if at=Ask, cmemif at=GetMemory, 0otherwise.C(a_t) = \begin{cases} c_{\mathrm{nav}}\,d(p_t, p_{t+1}) & \text{if }a_t = \mathrm{Navigate}, \ c_{\mathrm{ask}} (1 + \alpha N_{\mathrm{ask}}(t)) & \text{if }a_t = \mathrm{Ask}, \ c_{\mathrm{mem}} & \text{if }a_t = \mathrm{GetMemory}, \ 0 & \text{otherwise}. \end{cases}

Here, ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})0 is the physical distance, ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})1 counts prior queries, and cost coefficients (ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})2) reflect operational trade-offs. The core objective maximizes the expected net return over trajectories, combining sparse success rewards ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})3 with the weighted sum of cumulative costs:

ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})4

with hyperparameter ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})5 mediating efficiency-versus-success.

2. HC-GRPO Algorithmic Framework

Distinguishing itself from single-trajectory RL paradigms, HC-GRPO samples ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})6 trajectories per instruction, forming a group-based performance context. For a fixed query ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})7, ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})8 trajectories ht=(observations0:t,CoT0:t−1,actions0:t−1)h_t = (\text{observations}_{0:t}, \text{CoT}_{0:t-1}, \text{actions}_{0:t-1})9 yield rewards A\mathcal{A}0, with group mean A\mathcal{A}1 and standard deviation A\mathcal{A}2. The relative advantage for each sample is:

A\mathcal{A}3

ensuring A\mathcal{A}4 for unbiasedness.

The policy optimization objective uses the PPO-style clipped surrogate, where the importance sampling ratio is A\mathcal{A}5:

A\mathcal{A}6

A\mathcal{A}7 denotes the frozen supervised fine-tuning (SFT) policy, which regularizes KL divergence to maintain proximity to the initialization.

Pseudocode Overview

The algorithm iterates through RL epochs; for each instruction in a batch, A\mathcal{A}8 rollouts are generated and their group-relative advantages computed. Policy parameters are updated with gradient ascent on the loss A\mathcal{A}9, with no requirement for a separate value function network.

3. Theoretical Properties

Key theoretical attributes of HC-GRPO include:

  • Unbiased Baseline: The empirical group mean provides a baseline such that Navigate⁡(ℓ)\operatorname{Navigate}(\ell)0, preserving the unbiasedness of policy gradients.
  • Variance Reduction: Using per-group normalization (relative to single-trajectory critic estimates), the method attains lower variance gradient estimates, which is especially beneficial in the high-dimensional chain-of-thought (CoT) spaces engendered by MLLMs.
  • KL-Regularization: Explicit KL divergence between the current policy and the SFT reference (Navigate⁡(ℓ)\operatorname{Navigate}(\ell)1) ensures stable policy updates within a trust region, following the theoretical underpinnings of PPO with KL-constraint.
  • Critic Elimination: HC-GRPO dispenses with the value network Navigate⁡(ℓ)\operatorname{Navigate}(\ell)2, traditionally employed in PPO, substituting it with group empirical baselines.

4. Implementation Protocols and Hyperparameters

The backbone MLLM is Qwen2.5-VL-7B. The training protocol comprises two stages:

  • SFT Stage: AdamW optimizer, learning rate Navigate⁡(ℓ)\operatorname{Navigate}(\ell)3, batch size 16, 1 epoch, cosine learning rate decay (min 0.1).
  • HC-GRPO Stage: Learning rate Navigate⁡(ℓ)\operatorname{Navigate}(\ell)4, batch size 8, 3 epochs, discount factor Navigate⁡(ℓ)\operatorname{Navigate}(\ell)5, KL penalty Navigate⁡(ℓ)\operatorname{Navigate}(\ell)6, PPO clip Navigate⁡(ℓ)\operatorname{Navigate}(\ell)7, group size Navigate⁡(ℓ)\operatorname{Navigate}(\ell)8, cost tradeoff Navigate⁡(ℓ)\operatorname{Navigate}(\ell)9.

Cost parameters are ℓ\ell0, ℓ\ell1, ℓ\ell2, query fatigue ℓ\ell3. Rewards are ℓ\ell4, ℓ\ell5. Additional scheme includes an entropy bonus ℓ\ell6 and a format-penalty cost ℓ\ell7.

5. Empirical Outcomes

HC-GRPO’s efficacy is substantiated via extensive experiments in the AI2-THOR simulated environment. ESearch-R1, trained using HC-GRPO, achieves a success rate of 61.5%, surpassing the best ReAct baseline (60.0%). The mean total task cost (TTC) is halved (from approximately 3.3 to 1.6), and the success-weighted-by-cost (SwC) metric increases from 0.36 to 0.59, demonstrating marked improvements in operational efficiency.

Ablative analyses indicate that omitting dialogue reduces the success rate (SR) to 10.5%, while excluding memory components yields SR 52.0% and raises TTC to 2.3. Training without HC-GRPO (SFT only) results in SR 59.2% and TTC 2.3.

Sensitivity studies confirm that the policy retains superior cost-weighted performance across broad ranges of ℓ\ell8 and ℓ\ell9, highlighting meta-policy generalization. Qualitatively, emergent strategies prioritize minimal, targeted disambiguation (one Ask or memory lookup) before movement—reflecting an efficient, human-like cost-aware search heuristic.

6. Context and Significance

HC-GRPO constitutes a substantial departure from critic-based on-policy RL in high-cost, multimodal domains. By aligning optimization with the relative efficacy of reasoning-action trajectories under explicit cost structures, the method advances the ability of MLLM agents to operate strategically under real-world constraints. This innovation is particularly salient given the operational asymmetry and cost diversity inherent in embodied instruction-following tasks.

Validations in the ESearch-R1 system demonstrate considerable practical gains, underscoring the robustness of group-relative baseline techniques and their suitability for RL fine-tuning of large, generative multimodal models in interactive, physical contexts (Zhou et al., 21 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HC-GRPO (Heterogeneous Cost-Aware Group Relative Policy Optimization).