---
title: Majority Voting Reward Function
url: https://www.emergentmind.com/topics/majority-voting-reward-function
type: topic
---

# Majority Voting Reward Function

A majority voting reward function is a decision-theoretic or learning-theoretic construct in which collective outcomes, pseudo-labels, or direct rewards are determined by aggregating multiple votes—each corresponding to an agent’s recommendation, model output, or sampled solution—via majority rule. This paradigm underlies core methods in reinforcement learning with pseudo-labels, collective decision theory, and self-supervised reward shaping, and is central to current test-time reinforcement learning for large language models, stochastic voting in societal environments, and mechanism design for truthful aggregation. Both static and dynamically-weighted variants exist, as well as extensions overcoming canonical limitations of naïve majority aggregation.

## 1. Mathematical Implementation of the Majority Voting Reward

The classic majority voting reward is defined by aggregating a set of outputs or votes $\{o_i\}_{i=1}^N$ for a given input using frequency or confidence counts. For each candidate answer $y\in\mathcal{Y}$, the vote count is
$$
V(y) = \sum_{i=1}^N \mathbf{1}\bigl[\mathrm{Ans}(o_i) = y\bigr].
$$
The majority-voted pseudo-label is then
$$
y^*_\mathrm{MV} = \arg\max_{y\in\mathcal{Y}} V(y).
$$
For each output $o_i$, the reward is assigned as
$$
r_\mathrm{MV}(o_i) = \mathbf{1}\bigl[\mathrm{Ans}(o_i) = y^*_\mathrm{MV}\bigr].
$$
This reward function is widely used in pseudo-labeling weakly supervised learning and test-time reinforcement learning, including language modeling and complex problem solving [2512.15146, 2508.00410].

## 2. Limitations and Pathologies of Majority Voting Rewards

Several works have identified inherent limitations of the majority voting reward:

- **Sparse Supervision and Confirmation Bias**: Majority-based rewards provide a single binary feedback per datapoint, often failing to reward minority-but-correct outputs and amplifying confirmation bias in overrepresented, low-quality solutions [2512.15146, 2508.00410].

- **The Pit of Losses Paradox in Stochastic Environments**: Within the ViSE (Voting in Stochastic Environment) model, majority voting can destroy capital in both hostile ($\mu<0$) and highly favorable ($\mu>0$) environments, yielding expected one-step rewards less than those from trivially rejecting or accepting all proposals. This is expressed as
  $$
  E_X[\Delta_i] < 0 \quad \text{for } \mu/\sigma < -0.25 \quad (\text{Gaussian } X)
  $$
  and analogously in highly favorable regimes [2401.00592].

- **Majority Is Not Always Correct**: Empirically, hard aggregation tasks frequently exhibit cases where the true answer is a minority or even singleton solution, especially as the complexity of the reasoning problem increases [2509.06870].

These issues motivated the development of refined and alternative reward functions.

## 3. Advanced Extensions: Confidence-Weighted and Subgroup Rewards

To address these weaknesses, several designs extend the majority voting reward:

- **Stepwise Confidence-Weighted Voting**: Decomposes outputs into reasoning steps and assigns token-level confidences. The confidence-weighted majority pseudo-label is
  $$
  y^*_\mathrm{CW} = \arg\max_{y\in\mathcal{Y}} \sum_{i=1}^N C_i\,\mathbf{1}[\mathrm{Ans}(o_i)=y],
  $$
  where $C_i$ measures average step confidence per output [2512.15146].

- **Local Subgroup Consensus (SCOPE Framework)**: The output pool is dynamically partitioned into subgroups $S_j$, and local weighted majorities $y^*_j$ are computed per subgroup. The reward is
  $$
  r_j(o_i) = \mathbf{1}[\mathrm{Ans}(o_i) = y^*_j], \qquad o_i \in S_j,
  $$
  and subgroup size is chosen via Pareto optimization to balance local consensus strength and answer diversity. This produces multiple supervision signals per example and mitigates signal sparsity [2512.15146].

A summary of these variants:

| Reward Variant            | Formula                                                          | Key Feature                        |
|--------------------------|------------------------------------------------------------------|------------------------------------|
| Majority-Voting ($r_\mathrm{MV}$)       | $\mathbf{1}[\mathrm{Ans}(o_i)=y^*_\mathrm{MV}]$              | Binary per global majority         |
| Confidence-Weighted ($r_\mathrm{CW}$)   | $\mathbf{1}[\mathrm{Ans}(o_i)=y^*_\mathrm{CW}]$              | Incorporates stepwise confidence   |
| Subgroup SCOPE ($r_\mathrm{SCOPE}$)     | $\mathbf{1}[\mathrm{Ans}(o_i)=y^*_j],\ o_i\in S_j$           | Dense, multi-label supervision     |

## 4. Collective Decision-Theoretic Formulations

In stochastic collective choice—exemplified by the ViSE model—rewards from majority voting are linked to societal capital increments. Each proposal is an $n$-vector $\bm{\xi}$ of i.i.d. gains, and symmetrized majority approves proposals with $k>n/2$ yes votes, with tie-breaking at $k=n/2$ using probability $1/2$:
$$
b(k, n) = \begin{cases} 1 & k>n/2 \text{ or } k<n/2 \\ \tfrac12 & k=n/2 \end{cases}
$$
The expected reward for agent $i$ is
$$
E_X[\Delta_i] = \frac{1}{n} \sum_{k=0}^n (k \mu^+ + (n-k) \mu^-) b(k, n) {n \choose k} p^k q^{n-k},
$$
where $p = P(X>0)$ and $q = 1-p$ [2401.00592].

A notable identity is the mirror-symmetry:
$$
E_X[\Delta_i] - E_{-X}[\Delta_i] = \mu
$$
for all $n$ and $X$ with mean $\mu$. This underlies the mirrored performance and the emergence of "twin pits of losses"—regimes on both sides of neutrality where majority rewards are strictly inferior to always rejecting or always accepting proposals.

## 5. Mechanism Design and Strategic Incentives

Mechanism-design approaches implement majority rule and associated rewards via structured games:

- **Bloc-Formation Mechanism:** Agents submit a vote and a cooperation set; outcomes are determined by majority blocs. Off-equilibrium, a lottery $\eta(m)$ is enforced, assigning outcome probabilities based on coalition nominations. Truthful bloc formation yields Nash equilibrium, as any deviation decreases a voter's expected reward [2302.09548].

- **Random-Confirmations Mechanism:** Voting is followed by random sampling of confirmers. If no confirmers approve, a lottery $\beta(v)$, proportional to vote shares, is executed. Equilibrium is reached with sincere majority voting and confirmation. This maintains subgame perfect implementation of the majority decision [2302.09548].

In both cases, the "reward" is the probabilistic outcome—either deterministic (upon majority) or randomized (in off-equilibrium scenarios)—and is strictly increasing in an agent’s truthful support for her favorite alternative.

## 6. Majority Voting Reward in Reinforcement Learning and Language Models

Majority voting reward functions are fundamental in large language model RL:

- **Self-Consistency and Pseudo-Labeling:** In test-time RL, sampling $n$ rollouts, aggregating answers, and using the majority as a reward enables unsupervised or weakly supervised policy improvement. For rollout $y_i$, reward is $r(y_v, y_i) = \mathbf{1}[\mathrm{ans}(y_i) = \mathrm{ans}(y_v)]$ [2508.00410].

- **Contrastive and Cross-Referencing Rewards (Co-Reward):** By constructing semantically analogous prompts, computing majority pseudo-labels on each, and cross-referencing rewards across analogues, one enforces consistency and improves reward robustness, mitigating reward collapse [2508.00410].

- **Aggregator Learning (AggLM):** Instead of static voting, an aggregator model is trained to select or synthesize the final answer, rewarded only if it matches ground truth, enabling minority-correct answer recovery [2509.06870].

## 7. Implications, Symmetries, and Tuning

The fundamental symmetry $E_X[\Delta] - E_{-X}[\Delta] = \mu$ in majority voting reward functions demonstrates that adding environmental drift simply translates expected reward curves. As a consequence:

- In moderately neutral environments ($|\mu/\sigma| < 0.2$), majority voting outperforms always-accept or always-reject strategies.
- For $|\mu/\sigma|$ above pit-thresholds, extremal deterministic rules (“accept all” or “reject all”) are superior.
- Tuning voting quotas, or introducing reward/penalty adjustments, can keep aggregation systems out of “pit of losses” regimes [2401.00592].

Mechanism-design perspectives, confidence weighting, and subgroup partitioning are active research directions to improve the fidelity, density, and reliability of majority-based reward schemes in both artificial and collective-intelligence systems.

---

**References:**  
- [2401.00592] Majority voting is not good for heaven or hell, with mirrored performance  
- [2512.15146] Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning  
- [2509.06870] The Majority is not always right: RL training for solution aggregation  
- [2508.00410] Co-Reward: Self-supervised Reinforcement Learning for Large Language Model Reasoning via Contrastive Agreement  
- [2302.09548] Legitimacy of collective decisions: a mechanism design approach

Source: https://www.emergentmind.com/topics/majority-voting-reward-function