---
title: 'r1k: A Reverse Reasoning Dataset Study'
url: https://www.emergentmind.com/topics/r1k
type: topic
---

# r1k: A Reverse Reasoning Dataset Study

Searching arXiv for the cited paper and related references.
r1k is a reverse reasoning dataset introduced to study bidirectional reasoning in multi-stage fine-tuning. It is constructed by inverting 1,000 forward reasoning examples from s1k into naturally occurring reverse reasoning instances, so that a model must recover an original question from a newly posed question whose answer corresponds to that original prompt. In the underlying study, r1k is used to compare supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) under forward, reverse, and mixed-direction training regimes, with the central finding that SFT on reverse data alone outperforms SFT on the corresponding forward data, while naive forward–reverse mixing produces conflicting supervision signals and directional interference [2509.13079].

## 1. Conceptual definition and formalization

The study distinguishes **forward reasoning** and **reverse reasoning** as two related but directionally distinct objectives. Forward reasoning corresponds to s1k, a set of 1,000 high-quality “question → chain-of-thought → answer” examples from Muennighoff et al. (2025). Each forward instance is written as $(x_f, y_f)$, where $x_f$ is the prompt or question and

$$
y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].
$$

A teacher model $f$ implements the mapping

$$
f: x_f \to \mathrm{CoT}_f \to y_f.
$$

Reverse reasoning is defined through an inversion operator $g$ applied to each forward example:

$$
g(x_f, y_f) = (x_r, y_r),
$$

where $x_r$ is a new question whose answer is the original question $x_f$, and

$$
y_r = [\mathrm{CoT}_r;\ \texttt{<ANS>};\ \text{new final answer}].
$$

The paper characterizes this as an approximate inversion of the teacher’s forward mapping,

$$
g \approx f^{-1},
$$

so that solving $x_r$ requires “running the chain of thought in reverse” [2509.13079].

This formulation makes r1k more than a simple paraphrase set. The intended distinction is directional: forward examples and reverse examples encode different reasoning trajectories even when they remain semantically coupled. A plausible implication is that the dataset probes whether current fine-tuning methods preserve reasoning orientation or collapse distinct paths into a weaker averaged behavior.

## 2. Construction of the dataset

r1k is derived from **all 1,000 examples in s1k**. No additional human filtering was applied; DeepSeek-R1 both generated and validated each reverse instance. The construction pipeline is described as a three-step inversion methodology for each $(x_f, y_f) \in \mathrm{s1k}$ [2509.13079].

First, **Prompt template \#1** is used to “swap roles” of question and answer. Its input is the original question $x_f$ together with the final answer in $y_f$, and its output is a new question $x_r$ and a new answer $y_r$ expressed as a statement. Second, **Prompt template \#2** generates a detailed chain-of-thought and boxed final answer for $x_r$. Third, a **quality filter** applies automatic sanity checks on numeric consistency, while token counts are plotted in Figure A.3 of the paper.

The resulting dataset statistics are concise and specific. The dataset size is

$$
|r1k| = 1{,}000
$$

examples, with **no train/val/test split; all examples are used for SFT**. Domain coverage matches s1k, including **math word problems, logic puzzles, physics, …**. The **median token length is approximately 600 tokens**, with the full distribution reported in Figure A.3. In addition, a **small probe set of 50 forward + 50 reverse examples** is held out for learning-dynamics analysis [2509.13079].

These design choices matter because they isolate the effect of directional inversion while holding data scale fixed. The absence of additional human filtering suggests that quality control is delegated to model-based generation and validation plus automatic consistency checks, rather than manual curation.

## 3. Relationship to s1k and the role of inversion

The relation between s1k and r1k is one of paired directional transformation rather than independent corpus construction. s1k supplies high-quality forward exemplars, and r1k is built as their inverted counterpart. In forward form, the answer terminates the reasoning process; in reverse form, the new prompt is constructed so that the original question becomes the answer target. The study therefore treats the two datasets as aligned but directionally opposed supervision sources [2509.13079].

This pairing is central to the paper’s analysis of **bidirectional reasoning objectives**. The authors do not frame reverse reasoning as merely data augmentation. Instead, reverse instances are treated as a distinct supervision regime with its own chain-of-thought pattern. The experimental results then test whether training on one direction, the other direction, or both simultaneously changes downstream behavior.

A common misconception would be to assume that forward and reverse examples are interchangeable because they are derived from the same underlying semantic content. The paper argues against this view indirectly through its empirical findings: mixed training does not preserve clean directional preferences, and performance can deteriorate markedly relative to pure-direction training. This suggests that equivalence at the level of semantic content does not guarantee equivalence at the level of optimization dynamics.

## 4. Fine-tuning protocols and optimization setup

The experimental setup comprises an SFT stage and, in some conditions, a subsequent DPO stage. For SFT, the models are **Qwen2.5-Instruct (7B and 14B)**. The training objective is standard cross-entropy:

$$
L_{\mathrm{SFT}} = - \sum_{(x,y)} \sum_t \log p_\theta(y_t \mid x, y_{<t}).
$$

Parameter-efficient adaptation uses **LoRA** with **rank $r=256$** and **$\alpha=512$** on **q, k, v, o, gates, down/up\_proj, lm\_head**. The hyperparameters are: **initial learning rate $3 \times 10^{-4}$ with cosine schedule + restarts; weight decay $10^{-6}$; AdamW with $(\beta_1=0.9,\beta_2=0.95)$; clip = 1.0; warmup = 5% epochs; 10 epochs total; batch size = 1; mixed-precision BF16; max length = 20k tokens** [2509.13079].

For DPO, preference pairs $(x, y^+, y^-)$ are collected from the combined $\,\mathrm{s1k} \cup \mathrm{r1k}\,$ set. For a forward prompt $x_f$, the preferred response is $y_f$ and the dispreferred response is $y_r$; for a reverse prompt $x_r$, the preference relation is reversed, so $y^+=y_r$ and $y^-=y_f$. The DPO objective is

$$
L_{\mathrm{DPO}}(\theta)
=
-\mathbb{E}_{(x,y^+,y^-)}
\left[
\log \sigma \left(
\beta \big(
\log p_\theta(y^+|x) - \log p_\theta(y^-|x)
\big)
\right)
\right],
$$

together with an implicit KL penalty that keeps $p_\theta$ close to the SFT reference. The reported DPO hyperparameters are **$\beta=0.6$ for training and $\beta=0.2$ for probe analyses; learning rate $5\times 10^{-7}$ with cosine decay; gradient accumulation = 4; 200 steps (approximately 3 h)**. The same memory optimizations are retained: **FlashAttention, DeepSpeed ZeRO-3, and BF16** [2509.13079].

The two-stage design enables a separation between **instruction-following adaptation** and **preference-based directional correction**. In the paper’s framing, this is necessary because mixed SFT alone does not maintain a strong distinction between forward and reverse chains.

## 5. Empirical performance and learning dynamics

The paper reports downstream zero-shot accuracy using **lm-eval-harness** on **AIME24, Math500, and GPQA**. The central quantitative result is that **SFT on r1k outperforms SFT on s1k** for both model sizes [2509.13079].

| Data | Model | AIME / Math / GPQA / Average |
|---|---:|---|
| s1k | 7B | 16.7% / 77.0% / 34.0% / 42.6% |
| r1k | 7B | 20.0% / 77.4% / 42.4% / 46.6% (+4.0) |
| s1k | 14B | 20.0% / 83.2% / 48.4% / 50.6% |
| r1k | 14B | 33.3% / 86.0% / 53.0% / 57.4% (+6.8) |

The paper summarizes this pattern as **a 1.6%–6.8% accuracy improvement over s1k across evaluated benchmarks**. The reported table specifically shows average gains of **+4.0** for the 7B model and **+6.8** for the 14B model, with especially large movement on AIME24 for 14B [2509.13079].

By contrast, **mixing forward and reverse data degrades performance**. The paper considers a **half-mix (500 forward + 500 reverse)** and a **full mix (2k)**, and reports **average accuracy approximately 31–38%**, substantially below pure r1k training. The stated explanation is that **directional consistency collapses**: the model cannot strongly favor the correct chain in either direction.

DPO offers only a partial recovery. Applied to the mixed-trained **7B** model, it yields a **+7.1% gain (31.5 → 38.6)**, but this remains inferior to SFT on r1k alone. To analyze the dynamics, the paper tracks the margin

$$
\Delta = \mathrm{ALP}(y^+) - \mathrm{ALP}(y^-),
$$

where

$$
\mathrm{ALP}(y)
=
\frac{1}{N}\sum_i \frac{1}{|y_i|}\sum_t \log p(y_{i,t}\mid x_i, y_{<t}).
$$

Under mixed-data SFT, the margin remains **very small, approximately 0.05–0.10**. DPO widens $\Delta$ moderately, but also shifts probability mass away from even the correct reverse chain toward **generic safe outputs** [2509.13079].

These results establish the paper’s main contrast: reverse-only SFT is beneficial, mixed-direction SFT is harmful, and DPO only incompletely restores directional separation.

## 6. Interpretation, failure modes, and direction-aware alignment

The study’s interpretation is that **mixed reasoning data introduce conflicting supervision signals**. In the paper’s wording, forward and reverse examples carry conflicting gradient signals, and under an **NTK view** the interaction term **$K(x_u,x_v)\cdot G$ cancels out**, producing what the paper calls **directional interference** [2509.13079].

The observed failure modes are described in two parts. First, models trained on mixed data exhibit **higher hallucination**, specifically **in-distribution $y^-$ spikes**. Second, they assign **lower correct reasoning likelihood**, in that **$y^+$ underperforms pure-direction SFT**. DPO can partially restore the preference margin, but the paper cautions that it does so by shifting probability mass toward outputs that are not the desired reverse chains, characterizing these as **overly guarded outputs** or **generic safe outputs** depending on the context [2509.13079].

The paper therefore recommends several **direction-aware alignment strategies**. It proposes keeping **separate SFT stages for each reasoning direction**, followed by a **targeted DPO that respects direction tags**. It also recommends introducing **explicit direction tokens or prefixes**, such as **“[FORWARD]” vs. “[BACKWARD]”**, so the model conditions on the desired reasoning orientation. Further recommendations are to use **curriculum or weighted sampling** to avoid simultaneous conflicting updates, and to explore **alternative preference objectives that penalize only “off-direction” hallucinations rather than lumping them together** [2509.13079].

A plausible implication is that r1k functions as a stress test for alignment objectives in settings where semantically related examples encode mutually incompatible reasoning trajectories. In that sense, the dataset is relevant not only to reverse reasoning per se, but also to the broader question of how fine-tuning pipelines should represent and preserve latent task orientation.

## 7. Significance within multi-stage fine-tuning research

Within the paper’s scope, r1k serves as a compact, controlled benchmark for examining whether limited but high-quality reverse reasoning data can outperform corresponding forward data in downstream evaluation. The answer given by the experiments is affirmative for pure-direction SFT: **r1k provides a compact, high-quality signal that, when used alone for SFT, yields consistent gains over s1k** [2509.13079].

Its broader significance lies in what it reveals about the limitations of naive data mixing. The paper does not claim that more data is inherently worse; rather, it shows that **mixed directionality** can be worse when the optimization procedure does not explicitly encode the intended reasoning orientation. This reframes the problem from one of raw dataset size to one of supervision compatibility.

Accordingly, r1k occupies a specific position in the study of reasoning alignment: it is a reverse reasoning corpus built by inversion from s1k, a testbed for bidirectional reasoning objectives, and an empirical basis for the conclusion that **robust and direction-aware alignment strategies** are needed in multi-stage fine-tuning pipelines.

Source: https://www.emergentmind.com/topics/r1k