Papers
Topics
Authors
Recent
Search
2000 character limit reached

Align to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math Reasoning

Published 25 Jan 2026 in cs.CL | (2601.17671v1)

Abstract: Despite the impressive reasoning abilities demonstrated by LLMs, empirical evidence indicates that they are not language agnostic as expected, leading to performance declines in multilingual settings, especially for low-resource languages. We attribute the decline to the model's inconsistent multilingual understanding and reasoning alignment. To address this, we present Pivot-Aligned Self-Feedback Multilingual Reasoning (PASMR), aiming to improve the alignment of multilingual math reasoning abilities in LLMs. This approach designates the model's primary language as the pivot language. During training, the model first translates questions into the pivot language to facilitate better alignment of reasoning patterns. The reasoning process in the target language is then supervised by the pivot language's reasoning answers, thereby establishing a cross-lingual self-feedback mechanism without relying on external correct answers or reward models. Extensive experimental results demonstrate that our method enhances both the model's understanding of questions and its reasoning capabilities, leading to notable task improvements.

Summary

  • The paper introduces the PASMR framework to improve multilingual mathematical reasoning by explicitly mapping target language inputs to a pivot language with self-feedback reinforcement learning, enhancing overall accuracy by up to 32.8% on MGSM for low-resource languages.
  • Stronger reasoning models like Mistral-7B-Instruct benefit significantly from PASMR, boosting performance by 25.5 points on MGSM and 32.8 points on low-resource languages.
  • Self-feedback without gold answers or external translators ensures the consistency of answers across languages, showing a direct equivalence between the pivot and target languages, and achieving up to 79.1% improvement on MSVAMP.

Motivation and problem statement

LLMs exhibit substantial performance asymmetries across languages on reasoning tasks, despite the intuition that mathematical reasoning should be language-agnostic. The authors attribute this degradation to two failure points in the internal processing pipeline that prior interpretability work has characterized—mapping input to a dominant language, reasoning there, then mapping back to the target language: (1) imperfect understanding of multilingual inputs during the initial mapping, and (2) lossy projection of reasoning results back into the target language. Their diagnostic experiment shows that non-English answers from Mistral-7B-Instruct overlap with English answers but at lower accuracy, motivating a training framework that explicitly supervises both stages.

The PASMR framework

Pivot-Aligned Self-Feedback Multilingual Reasoning (PASMR) consists of two stages. In Pivot-Aligned Mapping (PAM), the model is fine-tuned to translate a target-language question QTQ_T into English QEnQ_{\text{En}} (delimited by special tokens) and then produce the answer ATA_T in the target language, optimizing the standard SFT objective over the joint sequence (QEn,AT)(Q_{\text{En}}, A_T) conditioned on QTQ_T. This makes the pivot mapping explicit rather than implicit, reducing bias introduced when the model silently maps problems to its dominant language.

In Self-feedback Reinforcement Learning (SRL), the PAM-trained model samples a trajectory containing the English translation and target-language answer; feeding the extracted QEnQ_{\text{En}} back into the model yields an English answer AEnA_{\text{En}}. A token-wise reward assigns 1 for agreement between ATA_T and AEnA_{\text{En}}, 0.1 for disagreement, 0 for format errors, with a KL penalty (β=0.01\beta = 0.01) against the PAM policy and discounted advantages (QEnQ_{\text{En}}0). Optimization uses REINFORCE++ with clipped ratios (QEnQ_{\text{En}}1). Crucially, the reward requires no gold answers or external reward models—the pivot-language output serves as the supervision signal, exploiting the fact that accuracy in the dominant language acts as an empirical upper bound for cross-lingual consistency.

The training data derives from GSM8K: solutions generated by Qwen2.5-instruct-7B, translated into nine languages by Qwen2.5-instruct-32B, and rule-filtered, yielding 2,048 instances per language (20,480 total), split evenly between SFT and RL. An OOD set built from NumGLUE tasks tests robustness.

Main results

Evaluation covers MGSM (in-domain) and MSVAMP (out-of-domain) across Mistral-7B-Instruct, Llama-3-8B-Instruct, and Deepseek-math-7b-instruct, against non-training baselines (Pipeline, MCOT) and training baselines (MSFT, MSFT with gold-answer RL, MAPO).

Model MGSM avg. MGSM low-res. MSVAMP avg. MSVAMP low-res.
Mistral-7B-Instruct 27.0 11.6 41.7 22.5
+ PASMR 52.5 44.4 60.2 53.1
Llama-3-8B-Instruct 57.8 45.3 69.5 62.0
+ PASMR 66.9 58.8 71.3 65.8
Deepseek-math-7b 61.8 39.2 70.0 51.4
+ PASMR 72.4 59.9 79.1 71.0

The headline gains are largest where baselines are weakest: Mistral improves 25.5 points on MGSM average and 32.8 points on low-resource languages (Bengali, Thai, Swahili); DeepSeek-math gains 20.7 points on low-resource MGSM. PASMR also outperforms gold-answer RL on several settings—for example, DeepSeek-math reaches 79.1 on MSVAMP versus 75.2 for gold-answer RL—which is notable because it achieves this without any ground-truth labels. Pipeline methods show high variance, including severe degradation on DeepSeek-math (MGSM average dropping from 61.8 to 33.9).

On robustness, PASMR trained on OOD NumGLUE data maintains or improves benchmark performance (Llama: 67.0/74.0 on MGSM/MSVAMP versus 66.9/71.3 with in-domain data), whereas MSFT improves in-domain but degrades out-of-domain (MSVAMP 69.5 → 64.1). This supports the claim that RL-based self-feedback generalizes better than supervised fine-tuning, which the data-scaling analysis reinforces: SRL improves steadily with more data while MSFT saturates or declines (48.4 → 45.7 on MSVAMP).

Ablations and analysis

Removing SRL consistently hurts performance, confirming that the RL stage contributes beyond PAM alone. Replacing self-generated pivots with NLLB-200 translations slightly underperforms self-generated ones, suggesting the fine-tuned model's internal translation better serves its own alignment than an external professional translator—a somewhat counterintuitive result worth noting. Substituting gold English pivots yields large additional gains (e.g., Llama MGSM average rises from 66.9 to 75.4; DeepSeek to 81.4), establishing pivot quality as a primary bottleneck.

The paper quantifies this bottleneck directly: target–pivot answer consistency reaches 82–98% when the pivot answer is correct but only 12–35% when it is wrong. Since the self-feedback reward rewards mere agreement, an incorrect pivot can reinforce incorrect target answers; the framework's ceiling is therefore bounded by pivot correctness, which itself depends on translation fidelity and task difficulty relative to model capability. Two further patterns emerge: languages far below English at baseline (Swahili, Bengali) gain most, largely via PAM; and weaker reasoners (Mistral) benefit more overall, while stronger models gain mainly in cross-lingual consistency rather than raw reasoning ability.

Limitations and open questions

Several constraints qualify the results. First, the method presupposes a well-defined dominant "pivot" language in which the model reasons reliably; for models without such a dominant language, the upper-bound argument underlying SRL weakens. Second, because the reward is agreement-based rather than correctness-based, the mechanism cannot correct errors that originate in the pivot language itself—the gold-pivot ablation shows substantial headroom remains (up to ~9 points on MGSM for Llama). Third, training data is machine-translated by Qwen models, so translation noise propagates into both PAM supervision and evaluation conditions; the paper does not measure sensitivity to translator quality beyond the NLLB comparison. Fourth, experiments are confined to math word problems with short verifiable answers; whether answer-consistency rewards transfer to open-ended or non-verifiable reasoning tasks is untested. Finally, the evaluation covers ten languages and three 7–8B models; scaling behavior and applicability to truly low-resource languages outside this set remain open questions, as does the interaction between pivot quality and task difficulty that the authors identify but do not fully disentangle.

Conclusion

PASMR couples explicit pivot-aligned mapping with a self-feedback RL loop that uses cross-lingual answer consistency as its reward signal, eliminating dependence on gold answers, external translators, or reward models. It delivers consistent gains across three base models on MGSM and MSVAMP, with the largest improvements on low-resource languages, and generalizes better under distribution shift than supervised alternatives. The central residual limitation—that the approach inherits the correctness ceiling of the pivot language—is clearly demonstrated by the authors' own ablations and constitutes the most direct target for subsequent work.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.