Papers
Topics
Authors
Recent
Search
2000 character limit reached

VHDL-Eval: LLM Benchmark for VHDL

Updated 3 July 2026
  • VHDL-Eval is a comprehensive framework that evaluates large language models on VHDL code generation, summarization, and self-consistency through curated tasks and testbenches.
  • It leverages a detailed dataset derived from Verilog translations and tutorial problems, categorizing tasks into simple, medium, and complex for precise performance assessment.
  • The framework employs zero-shot, in-context, and fine-tuning methods to expose common LLM error modes, with structured metrics and chain-of-descriptions improving overall reliability.

VHDL-Eval is a comprehensive benchmarking and evaluation framework designed to rigorously assess the capabilities of LLMs in the tasks of VHDL code generation, summarization, and self-consistency verification. Emerging from recent developments in automated hardware design tools, VHDL-Eval systematically addresses the pronounced gap between LLM competence in popular programming languages and their often-limited performance on hardware description languages—most notably, VHDL. The framework integrates a meticulously curated dataset, self-verifying testbenches, robust evaluation metrics, and comparative baselines, forming a reference point for practitioners and researchers in electronic design automation, LLM-based HDL synthesis, and code intelligence domains (Vijayaraghavan et al., 2024, Vijayaraghavan et al., 16 Jul 2025).

1. Dataset Design and Composition

VHDL-Eval comprises 202 canonical VHDL tasks, derived through two principal channels: translation from Verilog-based evaluation suites and aggregation of publicly available VHDL tutorial problems. The translation pipeline incorporates automated and manual stages. An initial pass employs ICARUS Iverilog to port Verilog RTL artifacts to VHDL, followed by manual post-processing to correct translation anomalies, canonicalize problem descriptions, and ensure functional fidelity (Vijayaraghavan et al., 2024).

The dataset encompasses three tiers of circuit complexity, categorized by lines of code (LOC) and signal architecture:

  • Simple: Gates (8–16 LOC, ~35%)
  • Medium: Adders, counters (17–40 LOC, ~50%)
  • Complex: FSMs, datapaths (>40 LOC, ~15%)

Each task includes the following fields:

  • Problem statement (English/natural language)
  • VHDL entity declaration (port/interface specification)
  • Canonical VHDL solution (as ground truth)
  • Self-verifying VHDL testbench (VUnit + GHDL based)
  • Meta-information: task ID, origin, task class, human-written summaries (Vijayaraghavan et al., 16 Jul 2025).

2. Testbench Architecture and Correctness Validation

Every VHDL-Eval task is paired with a self-checking testbench harness, employing VUnit and GHDL for simulation and assertion-driven output checking. The testbench structure is fully automated and adheres to the following canonical pattern:

1
2
3
4
5
6
7
8
9
10
for each test_vector in test_vectors loop
    apply_inputs(UUT, test_vector.inputs);
    wait for test_vector.delay;
    if UUT.outputs /= test_vector.expected then
        fail("Mismatch: got " & to_string(UUT.outputs)
             & " expected " & to_string(test_vector.expected));
    else
        pass("OK");
    end if;
end loop;

A task is deemed functionally correct if all test vectors pass, ensuring precise semantic validation. Coverage metrics are calculated per-task (fraction of passing vectors), with aggregate correctness defined as the fraction of tasks with full pass coverage (Vijayaraghavan et al., 2024, Vijayaraghavan et al., 16 Jul 2025).

3. Evaluation Methodologies and Metrics

Multiple evaluation modalities are supported:

The primary metrics include:

Functional Correctness: Pass@k

Defined for k sampled generations per task: Pass@k=1−(n−ck)(nk)\text{Pass@}k = 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}} where nn is the number of generations and cc is the number of correct solutions (Vijayaraghavan et al., 2024, Shen et al., 11 Jun 2026).

Self-Consistency (SC1\mathrm{SC}_1)

Measures whether the round-trip process (code → summary → regenerated code) yields functionally equivalent results, using equivalence checking against the canonical solution (Vijayaraghavan et al., 16 Jul 2025).

ROUGE-L and LLM Preference Rate (PR)

ROUGE-L measures the longest common subsequence between generated and reference summaries, while PR quantifies human or LLM preference for generated outputs over references in summarization tasks (Vijayaraghavan et al., 16 Jul 2025).

Structural Metrics

Tree-sitter–based weighted sequence and Jaccard similarities quantify the structural preservation between generated and reference code (Shen et al., 11 Jun 2026).

4. Baseline Performance and Error Profiles

Empirical results demonstrate that even state-of-the-art LLMs exhibit limited proficiency in VHDL code generation:

The self-consistency metric (SC1\mathrm{SC}_1) further reveals that bidirectional transformations (code-summarization-regeneration) are a principal failure point, with leading models achieving only ≈23.5% in baseline settings (Vijayaraghavan et al., 16 Jul 2025).

5. Methodological Advances: Chain-of-Descriptions and Validation Pipelines

The Chain-of-Descriptions (CoDes) strategy introduces structured intermediate planning steps—such as explicit design decompositions and synthesized high-level explanations—prior to code emission. This multi-step planning notably increases Pass@1 and SC1\mathrm{SC}_1 (e.g., Pass@1 (TB) from 19.2% to 25.4% for Granite-Code-34B), confirming that stepwise prompt construction and design intent clarification are critical for LLM VHDL competence (Vijayaraghavan et al., 16 Jul 2025).

Related frameworks, including VHDLSuite, present automated pipelines that enable large-scale Verilog-to-VHDL task synthesis, iterative simulation-driven repair, and robust error categorization. These pipelines apply LLMs in a generate→simulate→repair loop, leveraging feedback from VUnit/GHDL to refine outputs and filter only those that are functionally validated (Shen et al., 11 Jun 2026).

6. Diagnostic Analysis and Error Taxonomy

Sophisticated failure taxonomy reveals VHDL-specific challenges: uninitialized signals, strict type resolution requirements, declaration order constraints, and initialization semantics enforce a discipline absent from most programming language tasks. Automated classifiers segment failures into six categories: declaration errors, type/data object mismatches, instantiation/interface errors, combinational and sequential logic bugs, and miscellaneous/initialization problems (Shen et al., 11 Jun 2026).

Notably, successful generations show significantly higher structural similarity to reference code, whereas compile-time and runtime failures are associated with low syntactic correspondence.

7. Comparative Perspectives and Future Directions

Comparison with alternative frameworks, such as LLHD—a multi-level SSA-based IR—clarifies that operational semantics, strict event-queue scheduling, and delta-phase resolution are foundational for cycle- and delta-accurate simulation in VHDL evaluation (Schuiki et al., 2020). VHDL-Eval and VHDLSuite each highlight that language-specific rigor and semantics—such as nine-valued logic and explicit process-wait constructs—compound the design complexity relative to Verilog-centric benchmarks. Iterative validation/repair and feedback-driven fine-tuning are repeatedly shown to be essential for progress in LLM-driven hardware code synthesis (Vijayaraghavan et al., 2024, Shen et al., 11 Jun 2026).

Emergent trends suggest expanding task corpora, advancing in-context learning paradigms, automating testbench and golden reference generation, and integrating formal equivalence checking as promising routes for closing the performance gap. Open-sourcing of benchmark suites and evaluation harnesses is anticipated to accelerate the systematic development and assessment of VHDL-aware LLMs (Shen et al., 11 Jun 2026).


Key References:

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to VHDL-Eval.