Explain Model- and Task-Dependent Agent Trajectory Dynamics

Explain the causes of the differing relationships between token usage, trajectory length, model scale, task complexity, and HLS design outcomes observed in passing and failing agent traces.

Background

The paper analyzes token counts and trajectory steps for the gpt-oss-20b and gpt-oss-120b models, grouping runs by whether the generated HLS design passes the testbench and synthesis stages. The smaller model tends to use more tokens and steps on passing traces, whereas the larger model tends to use more tokens and steps on failing traces. The paper reports that the causes of these contrasting dynamics, and their relationship to model scale and task complexity, have not been established.

References

Although the causes of these dynamics and their relationship to model scale and task complexity remain unclear, these findings motivate deeper investigation.

Benchmarking Agentic HLS Design Tasks With HLS-Eval  (2609.09526 - Abi-Karam et al., 8 Sep 2026) in Section 3.3, “Agent Trajectories”