Papers
Topics
Authors
Recent
Search
2000 character limit reached

ReasonGR: Enhanced Retrieval via Reasoning

Updated 4 July 2026
  • ReasonGR is a reasoning-enhanced generative retrieval framework that integrates structured prompting with LoRA and QLoRA for multi-step semantic reasoning in complex financial queries.
  • It employs a mix of task-specific and chain-of-thought prompts to guide token-level docid generation, thereby improving retrieval accuracy and consistency.
  • The framework uses parameter-efficient adaptation and an adaptive loss scaling mechanism to outperform baselines like BM25 and DSI on the FinQA benchmark.

ReasonGR is a reasoning-enhanced generative retrieval framework designed to improve retrieval when queries require multi-step semantic reasoning, especially in numerical and financial document retrieval settings such as FinQA. In the standard generative retrieval formulation, a model directly generates a document identifier, or docid, for a query; ReasonGR retains that autoregressive docid-generation setup but adds structured prompting and reasoning-focused parameter adaptation with LoRA and QLoRA. The paper positions the method as a response to two limitations of ordinary generative retrieval: treating retrieval as a black-box generation task and relying on maximum-likelihood estimation over query-docid pairs without explicitly teaching intermediate reasoning (Dong et al., 12 Mar 2026).

1. Problem formulation and target setting

ReasonGR is defined over a corpus

D={d1,…,dN}\mathcal{D} = \{d_1, \dots, d_N\}

and a query set

Q={q1,…,qM}.\mathcal{Q} = \{q_1, \dots, q_M\}.

Each document has a unique docid jij_i, and retrieval is cast as sequence generation: Pθ(ji∣qi)=∏p=0∣ji∣Pθ(jip∣qi,ji<p).P_\theta(j_i \mid q_i) = \prod_{p=0}^{|j_i|} P_\theta(j_i^p \mid q_i, j_i^{<p}). In this formulation, jipj_i^p is the pp-th token of the docid and ji<pj_i^{<p} are the previously generated tokens (Dong et al., 12 Mar 2026).

The framework is motivated by complex financial queries for which retrieval depends on more than semantic matching. The paper identifies cases in which a model must locate relevant sections, interpret tables, connect numbers to textual context, and perform multiple reasoning steps before deciding which document is relevant. It argues that ordinary generative retrieval often fails on such inputs because it is trained to output a docid in one pass, while the training objective does not directly teach the model how to reason through intermediate steps.

Within that framing, ReasonGR is intended to improve retrieval accuracy, token-level consistency, and robustness on reasoning-heavy queries while preserving parameter efficiency. The central claim is not that docid generation should be replaced, but that docid generation should be preceded or accompanied by explicit reasoning traces. This suggests a shift from pure identifier prediction toward reasoning-conditioned identifier generation.

2. Prompt structure and reasoning-focused adaptation

ReasonGR combines structured prompting, reasoning-guided training, and a reasoning-focused adaptation module (Dong et al., 12 Mar 2026).

The prompting strategy concatenates a task-specific instruction with a chain-of-thought-style instruction. The task-specific templates include formulations such as “Answer the query with a document ID.”, “Generate the document ID that answers the question.”, “Based on the question, predict the document ID.”, “Retrieve a document ID that fits the query.”, and “Using the question, find the document ID.” These are paired with reasoning prompts such as “Use step-by-step reasoning.”, “You need to explain your answer.”, “Think this through carefully.”, “Let’s think step-by-step.”, and “Explain your reasoning before answering.” The paper describes this as prompt mixing: one task instruction is selected, one reasoning instruction is selected, they are combined into a full prompt, and the model is trained on outputs that include reasoning traces plus the docid.

In some variants, the framework also uses few-shot prompting. The stated purpose is to diversify the training distribution, expose the model to different phrasings, and encourage stable reasoning behavior rather than overfitting to one prompt style. The intended effect is to make the model behave less like a direct classifier and more like a reasoner that identifies important entities and numbers, traces evidence in the document, and only then generates the correct docid.

The adaptation module introduces LoRA into the transformer backbone while keeping the original model weights frozen: θ′=θ+Δθ,\theta' = \theta + \Delta \theta, with the update parameterized as

Δθ=AB⊤,A∈Rd×r,  B∈Rk×r,\Delta \theta = A B^\top, \quad A \in \mathbb{R}^{d \times r}, \; B \in \mathbb{R}^{k \times r},

where

r≪min⁡(d,k).r \ll \min(d,k).

The paper also applies QLoRA, quantizing the frozen backbone weights to 4-bit precision. According to the paper’s interpretation, this freezes general language and retrieval knowledge while learning a compact task-specific reasoning adapter. A plausible implication is that the reasoning module is meant to specialize the model toward numerical and multi-step retrieval behavior without requiring full fine-tuning.

3. Training procedure and loss design

ReasonGR uses two main training tasks (Dong et al., 12 Mar 2026). The first is corpus memorization: the model learns document-to-docid mappings using MLE loss. Because documents often contain tables with numerical data, the table contents are split and augmented with row and column headers to help the model learn from structured numeric evidence.

The second task is relevance and reasoning learning. The model takes queries as input and learns to generate correct multi-step reasoning traces together with the corresponding docid. To better match real-world query distributions, the paper generates pseudo-queries per document using a query generation model, and the implementation uses 10 pseudo-queries per document.

The paper introduces a custom adaptive penalty scaling loss: Q={q1,…,qM}.\mathcal{Q} = \{q_1, \dots, q_M\}.0 where Q={q1,…,qM}.\mathcal{Q} = \{q_1, \dots, q_M\}.1 is the cross-entropy loss and Q={q1,…,qM}.\mathcal{Q} = \{q_1, \dots, q_M\}.2 is a weighted penalty based on Exact Match (EM), Part Match (PM), Set Match (SM), and Structure Score (S-Score) (Dong et al., 12 Mar 2026). The paper describes this as fine-grained supervision: incorrect or incomplete predictions are penalized more strongly, especially when they fail to match the target docid structure. The text also notes that the formula is written somewhat awkwardly, but the intended idea is that cross-entropy is reweighted by a penalty term derived from retrieval-quality metrics.

The training design therefore goes beyond plain query-docid modeling in two ways. First, the target sequence includes reasoning traces followed by the final docid. Second, the loss is modulated by retrieval-sensitive evaluation signals rather than relying only on exact sequence likelihood. The paper’s argument is that this richer supervision is necessary for multi-step semantic reasoning in financial retrieval.

4. FinQA benchmark setting and empirical results

The evaluation is conducted on FinQA, described as a financial reasoning dataset containing questions that require numerical and semantic reasoning over financial reports and tables (Dong et al., 12 Mar 2026). The dataset has 8,281 samples, split into 75% training, 10% validation, and 15% test.

The docid specification appears in two forms in the paper details. One description states that the filename of each document is used as the unique docid. The implementation description states that docids are formed by concatenating keywords extracted using KeyBERT, with company and year prepended and joined by hyphens. The article’s reported results use that stated FinQA setup.

The baselines are BM25 and DSI, together with two internal variants: ReasonGR (Zero), which removes prompt training, and ReasonGR (CoT), which uses CoT prompting in a modified form. Evaluation uses EM, PM, SM, and S-Score; BM25 does not have PM or SM reported because it is not a generative token sequence model.

On the evaluation set, the reported results are:

  • BM25: EM 0.623
  • DSI: EM 0.563, PM 0.646, SM 0.651
  • ReasonGR (Zero): EM 0.572, PM 0.732, SM 0.748
  • ReasonGR (CoT): EM 0.571, PM 0.728, SM 0.748
  • ReasonGR: EM 0.607, PM 0.751, SM 0.765

On the test set, the reported results are:

  • BM25: EM 0.625
  • DSI: EM 0.578, PM 0.654, SM 0.659
  • ReasonGR (Zero): EM 0.601, PM 0.750, SM 0.767
  • ReasonGR (CoT): EM 0.612, PM 0.755, SM 0.774
  • ReasonGR: EM 0.626, PM 0.762, SM 0.779

The paper highlights several patterns in these numbers. ReasonGR consistently outperforms DSI across all reported metrics. The full ReasonGR model gives the best PM and SM, indicating stronger token-level and unordered token matching. It also reaches test EM 0.626, slightly above BM25 and better than DSI. The gains are strongest in PM and SM, which the paper interprets as evidence that the model generates structurally correct and semantically aligned docids more reliably than standard generative retrieval.

The internal comparisons function as ablations. ReasonGR (Zero) performs worse than the full model, which the paper uses to argue that prompting matters. ReasonGR (CoT) performs better than DSI and close to the full model, suggesting that reasoning prompts help, but are strongest when combined with the broader prompting strategy and parameter-efficient adaptation.

5. Computational profile and operational characteristics

The framework is designed to remain feasible under parameter-efficient adaptation rather than full-model updating (Dong et al., 12 Mar 2026). All variants have the same model-only memory, reported as 653.875 MiB. Training memory differs substantially across prompt regimes:

  • ReasonGR (Zero): 8496.19 MiB
  • ReasonGR (CoT): 25424.19 MiB
  • ReasonGR: 18280.19 MiB

Training time also varies by variant:

  • ReasonGR (Zero): 40 epochs, approximately 3:36:38
  • ReasonGR (CoT): 60 epochs, approximately 6:49:29
  • ReasonGR: 55 epochs, approximately 5:11:09

The paper interprets these measurements as showing that CoT prompts improve reasoning but are computationally heavier, while LoRA and QLoRA make the method feasible on a single GPU. This is consistent with the model design: the reasoning enhancement is introduced primarily through prompting and low-rank adaptation rather than through a large increase in trainable parameter count.

The paper’s qualitative description of the intended behavior is that ReasonGR should extract key information from a query, identify relevant sections in financial reports, connect numbers and semantic clues, and generate the docid formed by company name and report year. In that sense, the framework is meant to behave as a reasoning path from question to document identifier rather than as a pure lexical or dense matching mechanism.

6. Position within generative retrieval and stated limitations

ReasonGR belongs to a line of work that attempts to make generative retrieval less dependent on opaque one-pass docid prediction and more responsive to structured reasoning demands (Dong et al., 12 Mar 2026). A nearby but conceptually distinct direction is QUESTER, which reframes generative retrieval as query specification generation, specifically keyword query generation for BM25, and trains the policy with GRPO (Satouf et al., 7 Nov 2025). ReasonGR generates reasoning traces and docids; QUESTER generates an actionable search specification executed by a conventional search engine. A plausible implication is that recent generative retrieval research is splitting into at least two formulations: direct identifier generation augmented with reasoning, and specification generation coupled to symbolic retrieval.

The limitations stated for ReasonGR are specific. First, the evaluation is benchmark-centric: the method is validated on FinQA, but generalization to broader retrieval settings remains uncertain. Second, the input length is limited, and the paper notes that the 512-token limit is shorter than average document length, which may limit performance. Third, the backbone size is a constraint: FLAN-T5-base may be too small for strong CoT reasoning. Fourth, the loss granularity is limited, since sample-level loss may provide less fine-grained feedback than token-level supervision.

The future directions suggested in the paper follow directly from those limitations: use larger models, explore decoder-only architectures, add a classification head to better separate reasoning from retrieval, design token-level loss functions, and test on broader reasoning-intensive retrieval benchmarks. Taken together, these suggestions indicate that ReasonGR is presented not as a closed solution to generative retrieval in numerical settings, but as an initial framework showing that reasoning-aware prompting and parameter-efficient adaptation can significantly improve retrieval when the query requires multi-step semantic reasoning.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ReasonGR.