---
title: Posterior-Tempered Group Sampling (PTGS)
url: https://www.emergentmind.com/topics/posterior-tempered-group-sampling-ptgs
type: topic
---

# Posterior-Tempered Group Sampling (PTGS)

Posterior-Tempered Group Sampling (PTGS) is a training-time rollout-sampling method for reinforcement learning of language-model agents. It assigns a distinct decoding temperature to each prompt by estimating that prompt’s latent probability of producing a successful trajectory. Difficult prompts are sampled at higher temperatures to increase behavioral diversity, whereas easy prompts are sampled at lower temperatures to consolidate reliable behavior. PTGS changes the distribution of rollout groups without changing the PPO, GRPO, or other reinforcement-learning objective, and is designed to mitigate the loss of repeated-sampling coverage associated with post-training “sharpening” [2610.01509].

## 1. Motivation and empirical context

Post-training can improve single-shot accuracy while reducing solution coverage under repeated sampling. The relevant distinction is between $\mathrm{pass}@1$, which measures the probability that one rollout succeeds, and $\mathrm{pass}@k$, which measures the probability that at least one of $k$ rollouts succeeds. The metric $\mathrm{pass}^{k}$ measures the probability that all $k$ rollouts succeed and therefore characterizes repeated-attempt consistency.

For a prompt with per-rollout success probability $p_x$, the idealized relationships are

$$
\mathrm{pass}@k=\mathbb E_x\!\left[1-(1-p_x)^k\right],
$$

and

$$
\mathrm{pass}^{k}=\mathbb E_x[p_x^k].
$$

Prompts with $p_x$ near zero or one have little retry value. Prompts with intermediate $p_x$ have substantial solution coverage because some attempts fail while others succeed. The paper argues that post-training often moves prompts toward the extremes: tasks become either “always solved” or “never solved.” This increases immediate reliability for some prompts but eliminates recoverable failures that could have been solved through additional sampling.

The paper quantifies this phenomenon with the raw scalability measure

$$
A(K)=\sum_{k=1}^{K-1}\left[\mathrm{pass}@K-\mathrm{pass}@k\right].
$$

It also defines calibrated scalability,

$$
S(K)=\frac{1}{K-1}\sum_{k=1}^{K-1}
\frac{\mathrm{pass}@K-\mathrm{pass}@k}{1-\mathrm{pass}@1}
=
\frac{A(K)}{(K-1)(1-\mathrm{pass}@1)}.
$$

The Sharpening Tax compares a base model with its post-trained counterpart:

$$
\mathrm{Tax}_X(K)=X_{\mathrm{Base}}(K)-X_{\mathrm{Post}}(K),
\qquad X\in\{A,S\}.
$$

A positive tax indicates reduced test-time scalability after post-training. Under the per-task success-probability model, a task’s contribution is

$$
a_K(p)=\sum_{k=1}^{K-1}\left[(1-p)^k-(1-p)^K\right],
$$

which satisfies $a_K(0)=a_K(1)=0$ and is positive for $0<p<1$. Thus, sharpening reduces the value of retries by moving probability mass toward tasks that are nearly always solved or nearly always failed.

The motivating evaluation used 14 base/post-trained checkpoint pairs from four model families—Gemma-4, Ministral-3, Qwen2.5, and Qwen3.5—with effective model sizes from 3B to 35B parameters. It examined BFCL v4 multi-turn base split, WebShop, and ACEBench using rollout budgets $K\in\{1,2,4,8,16,32,64,128\}$. At $K=128$, calibrated Sharpening Tax was positive in 36 of 42 model–benchmark combinations. On WebShop with Gemma-4-31B, the base model exceeded $85\%$ $\mathrm{pass}@128$, compared with $56\%$ for the post-trained model.

## 2. Bayesian model of prompt difficulty

PTGS treats each prompt $x$ as having a latent current success probability

$$
p_x\in[0,1].
$$

Low values indicate difficulty for the current policy; high values indicate that the prompt is easy or mastered. For a group of $n$ rollouts, the number of successful trajectories is modeled as

$$
s_x\mid p_x\sim\mathrm{Binomial}(n,p_x).
$$

PTGS maintains discounted success and failure counts rather than an indefinitely accumulating historical total:

$$
(\tilde{s}_x,\tilde{f}_x)
\leftarrow
\left(
\gamma\tilde{s}_x+s_x,\;
\gamma\tilde{f}_x+n-s_x
\right),
\qquad 0\leq\gamma<1.
$$

The forgetting factor allows the estimate to track the current policy rather than an obsolete earlier policy.

At each training stage, PTGS specifies a target success rate $\tilde p$. It uses a Beta prior with total pseudo-count two:

$$
p_x\sim\mathrm{Beta}\bigl(2\tilde p,2(1-\tilde p)\bigr).
$$

After incorporating discounted observations, the posterior is

$$
p_x\mid\text{history}
\sim\mathrm{Beta}(a_x,b_x),
$$

where

$$
a_x=2\tilde p+\tilde{s}_x,
\qquad
b_x=2(1-\tilde p)+\tilde{f}_x.
$$

The posterior mean is $a_x/(a_x+b_x)$, but PTGS does not use this mean directly for temperature selection. Instead, it draws

$$
\hat p_x\sim\mathrm{Beta}(a_x,b_x)
$$

using Thompson sampling. The sampled value represents a plausible current success probability and retains posterior uncertainty.

The target $\tilde p$ has two roles. It centers the Beta prior and defines the boundary between hard and easy prompts. In the main experiments, $\tilde p$ was gradually increased from $0.25$ to $0.5$. The final value is motivated by the fact that Bernoulli variance $p(1-p)$ is maximized at $p=0.5$, where groups are most likely to contain both successful and unsuccessful trajectories.

## 3. Prompt-adaptive temperature rule

Let $\tau>1$ denote the temperature spread. PTGS assigns prompt $x$ the temperature

$$
T_x=\tau^{h(\hat p_x)},
$$

where

$$
h(p)=
\begin{cases}
\dfrac{\tilde p-p}{\tilde p},
& p\leq\tilde p,\\[8pt]
\dfrac{\tilde p-p}{1-\tilde p},
& p>\tilde p.
\end{cases}
$$

This mapping yields

$$
T_x\in[1/\tau,\tau].
$$

Its anchor points are

$$
T_x=\tau\quad\text{when }p=0,
$$

$$
T_x=1\quad\text{when }p=\tilde p,
$$

and

$$
T_x=1/\tau\quad\text{when }p=1.
$$

Consequently, hard prompts are heated, prompts near the target success rate retain the reference temperature, and easy prompts are cooled. The method can also use a reference temperature:

$$
T_x=T_{\mathrm{ref}}\tau^{h(\hat p_x)},
$$

with range

$$
T_x\in
\left[
\frac{T_{\mathrm{ref}}}{\tau},
T_{\mathrm{ref}}\tau
\right].
$$

When $\tilde p=1/2$, the rule simplifies to

$$
T_x=T_{\mathrm{ref}}\tau^{1-2\hat p_x}.
$$

The temperature policy is heuristic rather than the solution of an explicitly stated global optimization problem. Its decision principle is difficulty-guided exploration and exploitation: increase temperature where successful behavior is rare and decrease temperature where success is already reliable.

A global temperature cannot express this distinction. Increasing one temperature for all prompts may improve exploration on hard prompts while unnecessarily disrupting easy prompts and reducing $\mathrm{pass}@1$. PTGS instead generates a distribution of temperatures within a training step. Its mean temperature need not exceed the fixed-temperature baseline.

Thompson sampling provides uncertainty-sensitive exploration. When little evidence is available, the Beta posterior is broad and temperature choices vary stochastically. As discounted evidence accumulates, the posterior becomes more concentrated and the temperature assignment becomes more stable.

## 4. Group-level learning mechanism

For a group of $n$ conditionally independent rollouts with success probability $q$, PTGS uses

$$
I_n(q)=1-(1-q)^n,
$$

the probability that the group contains at least one success, and

$$
J_n(q)=1-q^n-(1-q)^n,
$$

the probability that the group is mixed, containing at least one success and at least one failure.

If a hard prompt’s success probability increases from $p$ to $p'>p$, then

$$
I_n(p')>I_n(p)
\qquad
\text{for }0\leq p<p'\leq1,
$$

and

$$
J_n(p')>J_n(p)
\qquad
\text{if }p<p'\leq\frac12.
$$

These properties are relevant to group-based reinforcement learning. PPO-like actor–critic methods can reinforce a successful trajectory when a group contains at least one success. GRPO-like group-contrast methods require mixed groups to produce nonzero relative advantages. Heating hard prompts therefore increases the probability of obtaining informative groups.

For mastered prompts with

$$
\frac12\leq p<p',
$$

the mixed-group probability decreases:

$$
J_n(p')<J_n(p).
$$

The paper gives the bound

$$
J_n(p)-J_n(p')
\leq
(p')^n-p^n.
$$

The interpretation is that cooling a mastered prompt reduces mixed groups primarily by converting them into all-success groups rather than all-failure groups.

This mechanism distinguishes PTGS from simply increasing entropy globally. Hard prompts are heated to avoid all-failure groups and expose successful behavior, while easy prompts are cooled so that sampling concentrates on reliable trajectories. The result is intended to improve both the learning signal available to the optimizer and the eventual trade-off between single-shot accuracy and repeated-sampling coverage.

## 5. Training algorithm and integration with reinforcement learning

PTGS operates per prompt group at every reinforcement-learning training step. The prompt-specific posterior state is maintained across visits to the prompt.

Initially,

$$
\tilde{s}_x=0,
\qquad
\tilde{f}_x=0
$$

for every task $x$ in the training pool. The principal hyperparameters are the group size $n$, target success rate $\tilde p$, temperature spread $\tau$, and forgetting factor $\gamma$.

For each training step, a prompt batch $\mathcal B_t$ is sampled. For every $x\in\mathcal B_t$, PTGS:

1. Forms the Beta posterior parameters
   $$
   a_x=2\tilde p+\tilde{s}_x,
   \qquad
   b_x=2(1-\tilde p)+\tilde{f}_x.
   $$

2. Draws
   $$
   \hat p_x\sim\mathrm{Beta}(a_x,b_x).
   $$

3. Computes
   $$
   T_x=\tau^{h(\hat p_x)}
   $$
   or
   $$
   T_x=T_{\mathrm{ref}}\tau^{h(\hat p_x)}.
   $$

4. Generates $n$ rollouts for $x$ at $T_x$, using the same temperature for every turn of every rollout in that group.

5. Scores the trajectories and counts the number $s_x$ of successful rollouts.

6. Updates the discounted counts:
   $$
   \tilde{s}_x\leftarrow\gamma\tilde{s}_x+s_x,
   $$
   $$
   \tilde{f}_x\leftarrow\gamma\tilde{f}_x+n-s_x.
   $$

The ordinary PPO, GRPO, or other RL update is then applied to the sampled trajectories. Policy-gradient log probabilities are evaluated at the temperature used to generate the trajectories, preserving on-policy training.

PTGS does not modify the RL loss, reward, critic, advantage estimator, or optimizer. It changes how rollout groups are sampled. In the main experiments it is removed at test time; evaluation uses a specified test temperature rather than the prompt-adaptive training mechanism.

## 6. Results, comparisons, and limitations

PTGS was evaluated during task-specific RL fine-tuning of Qwen2.5-7B-Instruct on Sokoban and slippery FrozenLake. Both PPO and GRPO were tested with and without PTGS. The training configuration used 200 StarPO/RAGEN training steps, a pool of 200 tasks, eight tasks per step, 16 rollouts per task, and 128 trajectories per step. The fixed-temperature baseline used $T=1.0$.

For PPO, the reported PTGS settings were $\tau=1.5$, temperature range $[0.667,1.5]$, $\gamma=0.95$, and $\tilde p$ increasing from $0.25$ to $0.5$. For GRPO, Sokoban used $\tau=1.4$ with fixed $\tilde p=0.5$, while FrozenLake used $\tau=1.2$ with $\tilde p$ increasing from $0.25$ to $0.5$. Evaluation used 64 unseen tasks per environment, 128 ancestral rollouts per task, test temperature $T=0.5$, and five independent runs.

| Environment and method | $\mathrm{pass}@1$ | $\mathrm{pass}@128$ | $\mathrm{Tax}_S(128)$ |
|---|---:|---:|---:|
| Sokoban, PPO | 46.5 | 55.0 | 0.094 |
| Sokoban, PPO + PTGS | **61.1** | **69.7** | **0.081** |
| Sokoban, GRPO | 36.5 | 55.3 | 0.081 |
| Sokoban, GRPO + PTGS | **39.1** | **72.5** | **0.025** |
| FrozenLake, PPO | 63.7 | 74.1 | 0.039 |
| FrozenLake, PPO + PTGS | **65.0** | **80.0** | **0.020** |
| FrozenLake, GRPO | 63.4 | 77.8 | 0.029 |
| FrozenLake, GRPO + PTGS | **67.3** | **81.2** | **0.028** |

Relative to fixed-temperature RL, PTGS improved both $\mathrm{pass}@1$ and $\mathrm{pass}@128$ in the reported PPO and GRPO comparisons while reducing the measured tax in every comparison. Posterior estimates correlated with realized group success at approximately $0.81$ on Sokoban and $0.71$ on FrozenLake. PTGS also reduced the frequency of zero-success groups.

The results were not attributed simply to globally hotter sampling. On FrozenLake, the mean temperature eventually fell below one as the model mastered prompts. On Sokoban, the mean remained hotter because the environment stayed difficult. Within a training step, the standard deviation of prompt temperatures was approximately $0.2$–$0.3$.

A fixed-temperature ablation compared PPO at $T\in\{0.5,0.7,1.0,1.3,1.5\}$. On Sokoban, PPO at $T=1.5$ achieved $\mathrm{pass}@1=57.4$, $\mathrm{pass}@128=64.4$, and tax $0.102$, whereas PTGS achieved $61.1$, $69.7$, and $0.081$, respectively. On FrozenLake, PPO at $T=1.5$ achieved $62.8$, $68.4$, and $0.055$, compared with PTGS at $65.0$, $80.0$, and $0.020$.

PTGS has several limitations. Its Bayesian model assumes binary success feedback and conditional independence within rollout groups. It is most natural for verifiable agentic environments and does not directly represent open-ended quality or partial credit. The method also assumes that heating can improve a hard prompt’s success probability; the paper does not claim that higher temperature always helps.

Prompt-specific statistics require recurring prompts or a finite task pool. The method is less informative for entirely novel prompts unless an inference-time pilot phase is used. Its behavior depends on $\tau$, $\tilde p$, $\gamma$, and group size, and the reported studies do not provide systematic ablations of alternative Beta prior masses, posterior families, group sizes, forgetting factors, or learned temperature-selection objectives.

The main training procedure reports no additional model-forward computational overhead beyond sampling and scoring the same rollout groups, apart from bookkeeping and Beta sampling. An inference-time PTGS variant does require pilot rollouts, which consume sampling budget. PTGS may also fail when higher temperature destroys structured action validity, when failures are fundamentally impossible rather than underexplored, when the difficulty posterior is miscalibrated, or when the policy contains no latent successful behavior for exploration to uncover.

PTGS is therefore a prompt-adaptive exploration–exploitation mechanism for RL rollout generation. Its central operation is the composition

$$
\text{discounted success statistics}
\;\longrightarrow\;
\text{Beta posterior}
\;\longrightarrow\;
\text{Thompson sample}
\;\longrightarrow\;
\text{prompt-level temperature}
\;\longrightarrow\;
\text{grouped rollouts}.
$$

The reported evidence supports the conclusion that PTGS can improve single-shot accuracy and repeated-sampling coverage simultaneously in the studied PPO and GRPO settings, while reducing Sharpening Tax relative to fixed-temperature training. The broader claim that it universally preserves rare capabilities or moves the post-training accuracy–coverage frontier outward remains dependent on the task distribution, success feedback, policy state, and temperature response of individual prompts.

Source: https://www.emergentmind.com/topics/posterior-tempered-group-sampling-ptgs