Close the NAS-Bench-NLP performance gap

Close the gap between the reproduced zero-cost proxy performance on the recurrent NAS-Bench-NLP search space and the approximately 0.56 correlation reported in the literature.

Background

The paper evaluates CoRA-NAS on NAS-Bench-NLP, a recurrent/LSTM language-cell search space, as an out-of-family test beyond the vision benchmarks. The authors report that the available proxies and architecture encoding are weak in this setting, producing correlations around 0.3, and that the refinement stage does not reliably improve the prior.

In reproducing prior results, the authors were unable to match the literature's reported correlation of approximately 0.56 across five implementations. This leaves an explicit unresolved reproducibility or performance-gap issue for zero-cost proxy evaluation on NAS-Bench-NLP.

References

It is an honest weak regime, and the reason is instructive. Two of the five proxies (AZ-NAS's feature-map views) are undefined on recurrent nets and dropped; the rest are weak (best, jacov, $0.31$ in our faithful reproduction; we could not close the gap to the literature's $0.56$ over five implementations, and report ours straight), so the prior itself is poor.

— CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search  (2609.11884 - Yang et al., 10 Sep 2026) in Section 4, subsection “The framework's boundary: a non-vision space (NAS-Bench-NLP)”