Generalizability of the LLM-assisted hardware-design protocol

Determine whether the six-step LLM-assisted hardware-design protocol—comprising executable golden-reference validation, a versioned project constitution, experiment logging, hardware-coupling-based review, typed feedback rules, and a staged validation ladder—transfers with a similar success profile to other LLMs, toolchains, operators, task mixes, cryptographic schemes, accelerator domains, or FPGA/ASIC flows.

Background

The paper evaluates the protocol in a single six-month campaign using Claude-family models, one developer, one FPGA platform, and a particular unified ML-KEM-768/ML-DSA-65 accelerator. The reported success rates therefore describe one tightly coupled model–operator–toolchain combination rather than a controlled comparison across systems.

The authors explicitly identify transferability as unresolved because the study does not establish whether the observed hardware-coupling gradient or the approximately 71.6% task-success profile would persist for different LLMs, toolchains, operators, task distributions, cryptographic schemes, accelerator domains, or FPGA/ASIC implementations. Resolving this problem would require replication with broader experimental populations and appropriate baselines.

References

Those evaluations, however, focused on zero-shot generation and editing, leaving open how agents achieve HLS design tasks.

Benchmarking Agentic HLS Design Tasks With HLS-Eval  (2609.09526 - Abi-Karam et al., 8 Sep 2026) in Abstract

Whether it transfers with a similar success profile is untested (Section~\ref{sec:threats}, H1--H2).

AI-Assisted Design of a Post-Quantum Cryptographic Accelerator: A Deployed-Silicon Case Study  (2609.04058 - Park et al., 3 Sep 2026) in Section 2.6, subsection “A Reusable Protocol for Other Implementations” (following the six-step protocol)