Performance of Smaller and Specialized Models on DeltaML-Bench

Determine the performance of smaller or specialized machine-learning models on DeltaML-Bench, extending evaluation beyond GPT-5 and Claude Sonnet 4.

Background

The study evaluates only two frontier model families, GPT-5 and Claude Sonnet 4, because of computational constraints. It also evaluates only the Modular and ARG scaffolding designs. Consequently, the benchmark’s results do not characterize how smaller or specialized models perform on autonomous machine-learning experimentation tasks in real-world research repositories.

References

While necessary for feasibility, this leaves the performance of smaller or specialized models on DeltaML-Bench an open question.

DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories  (2608.19653 - Moukpe et al., 20 Aug 2026) in Section Limitations