Papers
Topics
Authors
Recent
Search
2000 character limit reached

r1k: A Reverse Reasoning Dataset Study

Updated 12 July 2026
  • r1k is a reverse reasoning dataset that inverses 1,000 forward examples (from s1k) to study directional reasoning and optimization dynamics.
  • The dataset is constructed via a three-step inversion process including role-swapping, chain-of-thought generation, and automatic quality filtering.
  • Empirical findings reveal that SFT on r1k outperforms s1k by up to 6.8% accuracy, while mixing forward and reverse data leads to conflicting supervision signals.

Searching arXiv for the cited paper and related references. r1k is a reverse reasoning dataset introduced to study bidirectional reasoning in multi-stage fine-tuning. It is constructed by inverting 1,000 forward reasoning examples from s1k into naturally occurring reverse reasoning instances, so that a model must recover an original question from a newly posed question whose answer corresponds to that original prompt. In the underlying study, r1k is used to compare supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) under forward, reverse, and mixed-direction training regimes, with the central finding that SFT on reverse data alone outperforms SFT on the corresponding forward data, while naive forward–reverse mixing produces conflicting supervision signals and directional interference (Deng et al., 16 Sep 2025).

1. Conceptual definition and formalization

The study distinguishes forward reasoning and reverse reasoning as two related but directionally distinct objectives. Forward reasoning corresponds to s1k, a set of 1,000 high-quality “question → chain-of-thought → answer” examples from Muennighoff et al. (2025). Each forward instance is written as (xf,yf)(x_f, y_f), where xfx_f is the prompt or question and

yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].

A teacher model ff implements the mapping

f:xf→CoTf→yf.f: x_f \to \mathrm{CoT}_f \to y_f.

Reverse reasoning is defined through an inversion operator gg applied to each forward example:

g(xf,yf)=(xr,yr),g(x_f, y_f) = (x_r, y_r),

where xrx_r is a new question whose answer is the original question xfx_f, and

yr=[CoTr; <ANS>; new final answer].y_r = [\mathrm{CoT}_r;\ \texttt{<ANS>};\ \text{new final answer}].

The paper characterizes this as an approximate inversion of the teacher’s forward mapping,

xfx_f0

so that solving xfx_f1 requires “running the chain of thought in reverse” (Deng et al., 16 Sep 2025).

This formulation makes r1k more than a simple paraphrase set. The intended distinction is directional: forward examples and reverse examples encode different reasoning trajectories even when they remain semantically coupled. A plausible implication is that the dataset probes whether current fine-tuning methods preserve reasoning orientation or collapse distinct paths into a weaker averaged behavior.

2. Construction of the dataset

r1k is derived from all 1,000 examples in s1k. No additional human filtering was applied; DeepSeek-R1 both generated and validated each reverse instance. The construction pipeline is described as a three-step inversion methodology for each xfx_f2 (Deng et al., 16 Sep 2025).

First, Prompt template #1 is used to “swap roles” of question and answer. Its input is the original question xfx_f3 together with the final answer in xfx_f4, and its output is a new question xfx_f5 and a new answer xfx_f6 expressed as a statement. Second, Prompt template #2 generates a detailed chain-of-thought and boxed final answer for xfx_f7. Third, a quality filter applies automatic sanity checks on numeric consistency, while token counts are plotted in Figure A.3 of the paper.

The resulting dataset statistics are concise and specific. The dataset size is

xfx_f8

examples, with no train/val/test split; all examples are used for SFT. Domain coverage matches s1k, including math word problems, logic puzzles, physics, …. The median token length is approximately 600 tokens, with the full distribution reported in Figure A.3. In addition, a small probe set of 50 forward + 50 reverse examples is held out for learning-dynamics analysis (Deng et al., 16 Sep 2025).

These design choices matter because they isolate the effect of directional inversion while holding data scale fixed. The absence of additional human filtering suggests that quality control is delegated to model-based generation and validation plus automatic consistency checks, rather than manual curation.

3. Relationship to s1k and the role of inversion

The relation between s1k and r1k is one of paired directional transformation rather than independent corpus construction. s1k supplies high-quality forward exemplars, and r1k is built as their inverted counterpart. In forward form, the answer terminates the reasoning process; in reverse form, the new prompt is constructed so that the original question becomes the answer target. The study therefore treats the two datasets as aligned but directionally opposed supervision sources (Deng et al., 16 Sep 2025).

This pairing is central to the paper’s analysis of bidirectional reasoning objectives. The authors do not frame reverse reasoning as merely data augmentation. Instead, reverse instances are treated as a distinct supervision regime with its own chain-of-thought pattern. The experimental results then test whether training on one direction, the other direction, or both simultaneously changes downstream behavior.

A common misconception would be to assume that forward and reverse examples are interchangeable because they are derived from the same underlying semantic content. The paper argues against this view indirectly through its empirical findings: mixed training does not preserve clean directional preferences, and performance can deteriorate markedly relative to pure-direction training. This suggests that equivalence at the level of semantic content does not guarantee equivalence at the level of optimization dynamics.

4. Fine-tuning protocols and optimization setup

The experimental setup comprises an SFT stage and, in some conditions, a subsequent DPO stage. For SFT, the models are Qwen2.5-Instruct (7B and 14B). The training objective is standard cross-entropy:

xfx_f9

Parameter-efficient adaptation uses LoRA with rank yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].0 and yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].1 on q, k, v, o, gates, down/up_proj, lm_head. The hyperparameters are: initial learning rate yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].2 with cosine schedule + restarts; weight decay yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].3; AdamW with yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].4; clip = 1.0; warmup = 5% epochs; 10 epochs total; batch size = 1; mixed-precision BF16; max length = 20k tokens (Deng et al., 16 Sep 2025).

For DPO, preference pairs yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].5 are collected from the combined yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].6 set. For a forward prompt yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].7, the preferred response is yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].8 and the dispreferred response is yf=[CoTf; <ANS>; final answer].y_f = [\mathrm{CoT}_f;\ \texttt{<ANS>};\ \text{final answer}].9; for a reverse prompt ff0, the preference relation is reversed, so ff1 and ff2. The DPO objective is

ff3

together with an implicit KL penalty that keeps ff4 close to the SFT reference. The reported DPO hyperparameters are ff5 for training and ff6 for probe analyses; learning rate ff7 with cosine decay; gradient accumulation = 4; 200 steps (approximately 3 h). The same memory optimizations are retained: FlashAttention, DeepSpeed ZeRO-3, and BF16 (Deng et al., 16 Sep 2025).

The two-stage design enables a separation between instruction-following adaptation and preference-based directional correction. In the paper’s framing, this is necessary because mixed SFT alone does not maintain a strong distinction between forward and reverse chains.

5. Empirical performance and learning dynamics

The paper reports downstream zero-shot accuracy using lm-eval-harness on AIME24, Math500, and GPQA. The central quantitative result is that SFT on r1k outperforms SFT on s1k for both model sizes (Deng et al., 16 Sep 2025).

Data Model AIME / Math / GPQA / Average
s1k 7B 16.7% / 77.0% / 34.0% / 42.6%
r1k 7B 20.0% / 77.4% / 42.4% / 46.6% (+4.0)
s1k 14B 20.0% / 83.2% / 48.4% / 50.6%
r1k 14B 33.3% / 86.0% / 53.0% / 57.4% (+6.8)

The paper summarizes this pattern as a 1.6%–6.8% accuracy improvement over s1k across evaluated benchmarks. The reported table specifically shows average gains of +4.0 for the 7B model and +6.8 for the 14B model, with especially large movement on AIME24 for 14B (Deng et al., 16 Sep 2025).

By contrast, mixing forward and reverse data degrades performance. The paper considers a half-mix (500 forward + 500 reverse) and a full mix (2k), and reports average accuracy approximately 31–38%, substantially below pure r1k training. The stated explanation is that directional consistency collapses: the model cannot strongly favor the correct chain in either direction.

DPO offers only a partial recovery. Applied to the mixed-trained 7B model, it yields a +7.1% gain (31.5 → 38.6), but this remains inferior to SFT on r1k alone. To analyze the dynamics, the paper tracks the margin

ff8

where

ff9

Under mixed-data SFT, the margin remains very small, approximately 0.05–0.10. DPO widens f:xf→CoTf→yf.f: x_f \to \mathrm{CoT}_f \to y_f.0 moderately, but also shifts probability mass away from even the correct reverse chain toward generic safe outputs (Deng et al., 16 Sep 2025).

These results establish the paper’s main contrast: reverse-only SFT is beneficial, mixed-direction SFT is harmful, and DPO only incompletely restores directional separation.

6. Interpretation, failure modes, and direction-aware alignment

The study’s interpretation is that mixed reasoning data introduce conflicting supervision signals. In the paper’s wording, forward and reverse examples carry conflicting gradient signals, and under an NTK view the interaction term f:xf→CoTf→yf.f: x_f \to \mathrm{CoT}_f \to y_f.1 cancels out, producing what the paper calls directional interference (Deng et al., 16 Sep 2025).

The observed failure modes are described in two parts. First, models trained on mixed data exhibit higher hallucination, specifically in-distribution f:xf→CoTf→yf.f: x_f \to \mathrm{CoT}_f \to y_f.2 spikes. Second, they assign lower correct reasoning likelihood, in that f:xf→CoTf→yf.f: x_f \to \mathrm{CoT}_f \to y_f.3 underperforms pure-direction SFT. DPO can partially restore the preference margin, but the paper cautions that it does so by shifting probability mass toward outputs that are not the desired reverse chains, characterizing these as overly guarded outputs or generic safe outputs depending on the context (Deng et al., 16 Sep 2025).

The paper therefore recommends several direction-aware alignment strategies. It proposes keeping separate SFT stages for each reasoning direction, followed by a targeted DPO that respects direction tags. It also recommends introducing explicit direction tokens or prefixes, such as “[FORWARD]” vs. “[BACKWARD]”, so the model conditions on the desired reasoning orientation. Further recommendations are to use curriculum or weighted sampling to avoid simultaneous conflicting updates, and to explore alternative preference objectives that penalize only “off-direction” hallucinations rather than lumping them together (Deng et al., 16 Sep 2025).

A plausible implication is that r1k functions as a stress test for alignment objectives in settings where semantically related examples encode mutually incompatible reasoning trajectories. In that sense, the dataset is relevant not only to reverse reasoning per se, but also to the broader question of how fine-tuning pipelines should represent and preserve latent task orientation.

7. Significance within multi-stage fine-tuning research

Within the paper’s scope, r1k serves as a compact, controlled benchmark for examining whether limited but high-quality reverse reasoning data can outperform corresponding forward data in downstream evaluation. The answer given by the experiments is affirmative for pure-direction SFT: r1k provides a compact, high-quality signal that, when used alone for SFT, yields consistent gains over s1k (Deng et al., 16 Sep 2025).

Its broader significance lies in what it reveals about the limitations of naive data mixing. The paper does not claim that more data is inherently worse; rather, it shows that mixed directionality can be worse when the optimization procedure does not explicitly encode the intended reasoning orientation. This reframes the problem from one of raw dataset size to one of supervision compatibility.

Accordingly, r1k occupies a specific position in the study of reasoning alignment: it is a reverse reasoning corpus built by inversion from s1k, a testbed for bidirectional reasoning objectives, and an empirical basis for the conclusion that robust and direction-aware alignment strategies are needed in multi-stage fine-tuning pipelines.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to r1k.