Performance of Smaller and Specialized Models on DeltaML-Bench
Determine the performance of smaller or specialized machine-learning models on DeltaML-Bench, extending evaluation beyond GPT-5 and Claude Sonnet 4.
References
While necessary for feasibility, this leaves the performance of smaller or specialized models on DeltaML-Bench an open question.
— DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories
(2608.19653 - Moukpe et al., 20 Aug 2026) in Section Limitations