Establish whether larger diagnosis-training sets improve or saturate performance

Determine whether increasing the full-diagnosis training set beyond 480 or 948 source tasks yields further improvement in exact-step attribution accuracy or instead produces saturation under the evaluated training procedures.

Background

The nested full-diagnosis training ladder reports gains as the number of source tasks increases, but the contrast between the intermediate training sizes has an interval that includes zero. Because training-set size also changes optimizer updates, checkpoint selection, and decoding conditions, the reported results cannot determine whether performance continues to improve, has saturated, or reflects procedural differences.

References

The latter establishes neither further improvement nor saturation.

— Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training  (2609.40111 - Zhu et al., 30 Sep 2026) in Appendix I.11, page 47