Re-evaluate specialized agents on matched backbone models

Re-run each specialized CAE simulation agent on the same backbone model used by the Direct Baseline to determine whether performance differences arise from the agent scaffolding or from stronger underlying base models.

Background

The comparison reports published results for specialized systems rather than re-running those systems under the experimental conditions used for the Direct Baseline. Although information access, repair budget, and scoring criteria were matched, the underlying backbone models were not. Consequently, any observed advantage of the Direct Baseline may partly reflect differences in base-model capability rather than the absence of simulation-specific scaffolding. A matched-backbone re-evaluation would isolate the contribution of specialized agent design.

References

Part of the gap between the Direct Baseline and the specialized systems may therefore reflect stronger base models rather than the absence of scaffolding; re-running each specialized system on the same backbone would separate the two, and we leave this to future work.

What Do CAE Simulation Agents Really Need Beyond a Generic Harness?  (2609.03718 - Shi et al., 3 Sep 2026) in Section "Limitations", paragraph "Reported rather than re-run specialized baselines"