Re-run the Track B campaign under the corrected tool protocol

Re-run the GateTruth Track B agentic-repair campaign using the corrected file-access allowlist, directory-discovery interface, and strict per-task budget enforcement, and determine whether the reported objective-met results persist under the repaired protocol.

Background

The reported Track B campaign was conducted before several protocol defects were fixed. Agents could read public testbenches, had limited ability to discover available files, and could exceed nominal token budgets because of flaws in the pre-call and post-call checks.

The paper states that the corrected tool has since been implemented but that the campaign was not rerun. A corrected rerun is therefore needed to assess whether the original results were affected by these protocol gaps.

References

We have since added a strict allowlist restricting read_file to the design directory only, confirmed directly against this exact scenario (repeating the census's own action now returns a rejection instead of the testbench content), but the campaign in Table~\ref{tab:trackb} predates that fix and we do not re-run it here; a re-run under the corrected tool is future work.

GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing  (2608.12635 - Bhadra, 12 Aug 2026) in Section 10, Limitations and Threats to Validity, Four Track B protocol gaps