Test whether the observed separation survives harder optimization settings

Investigate whether the apparent interchangeability of architectural knowledge and structured critique persists for heterogeneous operators, tighter physical constraints, and substantially smaller evaluation budgets, where incorrect search hypotheses cannot be corrected through extensive sampling.

Background

The reported experiments use a small FP16 GEMM workload basket, a single modeled accelerator, and relatively generous evaluation opportunities. The authors suggest that these conditions produce broad near-optimal plateaus, allowing opaque search to perform well without recovering the physical meaning of the variables.

The proposed follow-up settings include attention, collectives, sparse layers, tighter constraints, and materially smaller evaluation budgets. These settings are intended to determine whether the observed relationship between architectural knowledge and dialectical critique remains stable when the optimization problem is more difficult and sampling is less effective.

References

Making the question harder to dodge is what we are working on now: heterogeneous operators such as attention, collectives and sparse layers; tighter constraints; and materially smaller evaluation budgets, the axis we expect to bite first, since it is where a wrong hypothesis can no longer be corrected by sampling.

Do AI Agents Understand Computer Architecture?  (2609.19387 - Sharan et al., 16 Sep 2026) in Section 3, Discussion and Future Work