Replicate the results across datasets at a higher-accuracy operating point

Replicate the matrix-CODI rank-ablation experiments on GSM8K at a higher-accuracy operating point to determine whether the observed rank-indifference generalizes beyond ProsQA and the low-learning GSM8K-Aug result.

Background

The core experiments use ProsQA and GPT-2 backbones. The only GSM8K-Aug experiment operates at approximately 6% accuracy, which the paper characterizes as barely learning the task and therefore weak evidence for the broader claim.

A higher-accuracy cross-dataset replication would test whether flat rank-k curves and rank-indifferent behavior persist on a substantially different reasoning benchmark.

References

Cross-dataset replication on GSM8K at a higher-accuracy operating point is pending.

The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA  (2609.03090 - Larson, 2 Sep 2026) in Section 7, paragraph “Single task, single architecture family”

A three-seed replication per variant ($\sim!42$ H100-hours) and re-running the four positive-control rank-$k$ evaluations on the full $500$-problem ProsQA test set are both pending; the $500$-problem eval raises power to detect $|r_s|!\geq!0.15$ from $\sim!40\%$ to $\sim!80\%$ at $\alpha!=!0.05$.

The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA  (2609.03090 - Larson, 2 Sep 2026) in Section 7, paragraph “One seed per positive control”

A reasoning task whose ground truth provably requires $k>1$ independent quantities at the answer position would disambiguate; we do not have one at this scale.

The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA  (2609.03090 - Larson, 2 Sep 2026) in Section 7, paragraph “Alternative explanation: the task is rank-1-solvable”