r1k: A Reverse Reasoning Dataset Study
- r1k is a reverse reasoning dataset that inverses 1,000 forward examples (from s1k) to study directional reasoning and optimization dynamics.
- The dataset is constructed via a three-step inversion process including role-swapping, chain-of-thought generation, and automatic quality filtering.
- Empirical findings reveal that SFT on r1k outperforms s1k by up to 6.8% accuracy, while mixing forward and reverse data leads to conflicting supervision signals.
Searching arXiv for the cited paper and related references. r1k is a reverse reasoning dataset introduced to study bidirectional reasoning in multi-stage fine-tuning. It is constructed by inverting 1,000 forward reasoning examples from s1k into naturally occurring reverse reasoning instances, so that a model must recover an original question from a newly posed question whose answer corresponds to that original prompt. In the underlying study, r1k is used to compare supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) under forward, reverse, and mixed-direction training regimes, with the central finding that SFT on reverse data alone outperforms SFT on the corresponding forward data, while naive forward–reverse mixing produces conflicting supervision signals and directional interference (Deng et al., 16 Sep 2025).
1. Conceptual definition and formalization
The study distinguishes forward reasoning and reverse reasoning as two related but directionally distinct objectives. Forward reasoning corresponds to s1k, a set of 1,000 high-quality “question → chain-of-thought → answer” examples from Muennighoff et al. (2025). Each forward instance is written as , where is the prompt or question and
A teacher model implements the mapping
Reverse reasoning is defined through an inversion operator applied to each forward example:
where is a new question whose answer is the original question , and
The paper characterizes this as an approximate inversion of the teacher’s forward mapping,
0
so that solving 1 requires “running the chain of thought in reverse” (Deng et al., 16 Sep 2025).
This formulation makes r1k more than a simple paraphrase set. The intended distinction is directional: forward examples and reverse examples encode different reasoning trajectories even when they remain semantically coupled. A plausible implication is that the dataset probes whether current fine-tuning methods preserve reasoning orientation or collapse distinct paths into a weaker averaged behavior.
2. Construction of the dataset
r1k is derived from all 1,000 examples in s1k. No additional human filtering was applied; DeepSeek-R1 both generated and validated each reverse instance. The construction pipeline is described as a three-step inversion methodology for each 2 (Deng et al., 16 Sep 2025).
First, Prompt template #1 is used to “swap roles” of question and answer. Its input is the original question 3 together with the final answer in 4, and its output is a new question 5 and a new answer 6 expressed as a statement. Second, Prompt template #2 generates a detailed chain-of-thought and boxed final answer for 7. Third, a quality filter applies automatic sanity checks on numeric consistency, while token counts are plotted in Figure A.3 of the paper.
The resulting dataset statistics are concise and specific. The dataset size is
8
examples, with no train/val/test split; all examples are used for SFT. Domain coverage matches s1k, including math word problems, logic puzzles, physics, …. The median token length is approximately 600 tokens, with the full distribution reported in Figure A.3. In addition, a small probe set of 50 forward + 50 reverse examples is held out for learning-dynamics analysis (Deng et al., 16 Sep 2025).
These design choices matter because they isolate the effect of directional inversion while holding data scale fixed. The absence of additional human filtering suggests that quality control is delegated to model-based generation and validation plus automatic consistency checks, rather than manual curation.
3. Relationship to s1k and the role of inversion
The relation between s1k and r1k is one of paired directional transformation rather than independent corpus construction. s1k supplies high-quality forward exemplars, and r1k is built as their inverted counterpart. In forward form, the answer terminates the reasoning process; in reverse form, the new prompt is constructed so that the original question becomes the answer target. The study therefore treats the two datasets as aligned but directionally opposed supervision sources (Deng et al., 16 Sep 2025).
This pairing is central to the paper’s analysis of bidirectional reasoning objectives. The authors do not frame reverse reasoning as merely data augmentation. Instead, reverse instances are treated as a distinct supervision regime with its own chain-of-thought pattern. The experimental results then test whether training on one direction, the other direction, or both simultaneously changes downstream behavior.
A common misconception would be to assume that forward and reverse examples are interchangeable because they are derived from the same underlying semantic content. The paper argues against this view indirectly through its empirical findings: mixed training does not preserve clean directional preferences, and performance can deteriorate markedly relative to pure-direction training. This suggests that equivalence at the level of semantic content does not guarantee equivalence at the level of optimization dynamics.
4. Fine-tuning protocols and optimization setup
The experimental setup comprises an SFT stage and, in some conditions, a subsequent DPO stage. For SFT, the models are Qwen2.5-Instruct (7B and 14B). The training objective is standard cross-entropy:
9
Parameter-efficient adaptation uses LoRA with rank 0 and 1 on q, k, v, o, gates, down/up_proj, lm_head. The hyperparameters are: initial learning rate 2 with cosine schedule + restarts; weight decay 3; AdamW with 4; clip = 1.0; warmup = 5% epochs; 10 epochs total; batch size = 1; mixed-precision BF16; max length = 20k tokens (Deng et al., 16 Sep 2025).
For DPO, preference pairs 5 are collected from the combined 6 set. For a forward prompt 7, the preferred response is 8 and the dispreferred response is 9; for a reverse prompt 0, the preference relation is reversed, so 1 and 2. The DPO objective is
3
together with an implicit KL penalty that keeps 4 close to the SFT reference. The reported DPO hyperparameters are 5 for training and 6 for probe analyses; learning rate 7 with cosine decay; gradient accumulation = 4; 200 steps (approximately 3 h). The same memory optimizations are retained: FlashAttention, DeepSpeed ZeRO-3, and BF16 (Deng et al., 16 Sep 2025).
The two-stage design enables a separation between instruction-following adaptation and preference-based directional correction. In the paper’s framing, this is necessary because mixed SFT alone does not maintain a strong distinction between forward and reverse chains.
5. Empirical performance and learning dynamics
The paper reports downstream zero-shot accuracy using lm-eval-harness on AIME24, Math500, and GPQA. The central quantitative result is that SFT on r1k outperforms SFT on s1k for both model sizes (Deng et al., 16 Sep 2025).
| Data | Model | AIME / Math / GPQA / Average |
|---|---|---|
| s1k | 7B | 16.7% / 77.0% / 34.0% / 42.6% |
| r1k | 7B | 20.0% / 77.4% / 42.4% / 46.6% (+4.0) |
| s1k | 14B | 20.0% / 83.2% / 48.4% / 50.6% |
| r1k | 14B | 33.3% / 86.0% / 53.0% / 57.4% (+6.8) |
The paper summarizes this pattern as a 1.6%–6.8% accuracy improvement over s1k across evaluated benchmarks. The reported table specifically shows average gains of +4.0 for the 7B model and +6.8 for the 14B model, with especially large movement on AIME24 for 14B (Deng et al., 16 Sep 2025).
By contrast, mixing forward and reverse data degrades performance. The paper considers a half-mix (500 forward + 500 reverse) and a full mix (2k), and reports average accuracy approximately 31–38%, substantially below pure r1k training. The stated explanation is that directional consistency collapses: the model cannot strongly favor the correct chain in either direction.
DPO offers only a partial recovery. Applied to the mixed-trained 7B model, it yields a +7.1% gain (31.5 → 38.6), but this remains inferior to SFT on r1k alone. To analyze the dynamics, the paper tracks the margin
8
where
9
Under mixed-data SFT, the margin remains very small, approximately 0.05–0.10. DPO widens 0 moderately, but also shifts probability mass away from even the correct reverse chain toward generic safe outputs (Deng et al., 16 Sep 2025).
These results establish the paper’s main contrast: reverse-only SFT is beneficial, mixed-direction SFT is harmful, and DPO only incompletely restores directional separation.
6. Interpretation, failure modes, and direction-aware alignment
The study’s interpretation is that mixed reasoning data introduce conflicting supervision signals. In the paper’s wording, forward and reverse examples carry conflicting gradient signals, and under an NTK view the interaction term 1 cancels out, producing what the paper calls directional interference (Deng et al., 16 Sep 2025).
The observed failure modes are described in two parts. First, models trained on mixed data exhibit higher hallucination, specifically in-distribution 2 spikes. Second, they assign lower correct reasoning likelihood, in that 3 underperforms pure-direction SFT. DPO can partially restore the preference margin, but the paper cautions that it does so by shifting probability mass toward outputs that are not the desired reverse chains, characterizing these as overly guarded outputs or generic safe outputs depending on the context (Deng et al., 16 Sep 2025).
The paper therefore recommends several direction-aware alignment strategies. It proposes keeping separate SFT stages for each reasoning direction, followed by a targeted DPO that respects direction tags. It also recommends introducing explicit direction tokens or prefixes, such as “[FORWARD]” vs. “[BACKWARD]”, so the model conditions on the desired reasoning orientation. Further recommendations are to use curriculum or weighted sampling to avoid simultaneous conflicting updates, and to explore alternative preference objectives that penalize only “off-direction” hallucinations rather than lumping them together (Deng et al., 16 Sep 2025).
A plausible implication is that r1k functions as a stress test for alignment objectives in settings where semantically related examples encode mutually incompatible reasoning trajectories. In that sense, the dataset is relevant not only to reverse reasoning per se, but also to the broader question of how fine-tuning pipelines should represent and preserve latent task orientation.
7. Significance within multi-stage fine-tuning research
Within the paper’s scope, r1k serves as a compact, controlled benchmark for examining whether limited but high-quality reverse reasoning data can outperform corresponding forward data in downstream evaluation. The answer given by the experiments is affirmative for pure-direction SFT: r1k provides a compact, high-quality signal that, when used alone for SFT, yields consistent gains over s1k (Deng et al., 16 Sep 2025).
Its broader significance lies in what it reveals about the limitations of naive data mixing. The paper does not claim that more data is inherently worse; rather, it shows that mixed directionality can be worse when the optimization procedure does not explicitly encode the intended reasoning orientation. This reframes the problem from one of raw dataset size to one of supervision compatibility.
Accordingly, r1k occupies a specific position in the study of reasoning alignment: it is a reverse reasoning corpus built by inversion from s1k, a testbed for bidirectional reasoning objectives, and an empirical basis for the conclusion that robust and direction-aware alignment strategies are needed in multi-stage fine-tuning pipelines.