---
title: 'VHDL-Eval: LLM Benchmark for VHDL'
url: https://www.emergentmind.com/topics/vhdl-eval
type: topic
---

# VHDL-Eval: LLM Benchmark for VHDL

VHDL-Eval is a comprehensive benchmarking and evaluation framework designed to rigorously assess the capabilities of Large Language Models (LLMs) in the tasks of VHDL code generation, summarization, and self-consistency verification. Emerging from recent developments in automated hardware design tools, VHDL-Eval systematically addresses the pronounced gap between LLM competence in popular programming languages and their often-limited performance on hardware description languages—most notably, VHDL. The framework integrates a meticulously curated dataset, self-verifying testbenches, robust evaluation metrics, and comparative baselines, forming a reference point for practitioners and researchers in electronic design automation, LLM-based HDL synthesis, and code intelligence domains [2406.04379, 2507.12308].

## 1. Dataset Design and Composition

VHDL-Eval comprises 202 canonical VHDL tasks, derived through two principal channels: translation from Verilog-based evaluation suites and aggregation of publicly available VHDL tutorial problems. The translation pipeline incorporates automated and manual stages. An initial pass employs ICARUS Iverilog to port Verilog RTL artifacts to VHDL, followed by manual post-processing to correct translation anomalies, canonicalize problem descriptions, and ensure functional fidelity [2406.04379].

The dataset encompasses three tiers of circuit complexity, categorized by lines of code (LOC) and signal architecture:
- **Simple**: Gates (8–16 LOC, ~35%)
- **Medium**: Adders, counters (17–40 LOC, ~50%)
- **Complex**: FSMs, datapaths (>40 LOC, ~15%)

Each task includes the following fields:
- Problem statement (English/natural language)
- VHDL entity declaration (port/interface specification)
- Canonical VHDL solution (as ground truth)
- Self-verifying VHDL testbench (VUnit + GHDL based)
- Meta-information: task ID, origin, task class, human-written summaries [2507.12308].

## 2. Testbench Architecture and Correctness Validation

Every VHDL-Eval task is paired with a self-checking testbench harness, employing VUnit and GHDL for simulation and assertion-driven output checking. The testbench structure is fully automated and adheres to the following canonical pattern:

```vhdl
for each test_vector in test_vectors loop
    apply_inputs(UUT, test_vector.inputs);
    wait for test_vector.delay;
    if UUT.outputs /= test_vector.expected then
        fail("Mismatch: got " & to_string(UUT.outputs)
             & " expected " & to_string(test_vector.expected));
    else
        pass("OK");
    end if;
end loop;
```

A task is deemed functionally correct if all test vectors pass, ensuring precise semantic validation. Coverage metrics are calculated per-task (fraction of passing vectors), with aggregate correctness defined as the fraction of tasks with full pass coverage [2406.04379, 2507.12308].

## 3. Evaluation Methodologies and Metrics

Multiple evaluation modalities are supported:
- **Zero-shot prompting**: LLMs generate VHDL code from a problem statement without prior examples.
- **In-Context Learning (ICL)**: K-shot demonstrations (e.g., k=5) are prefixed to each prompt.
- **Parameter-Efficient Fine-Tuning (PEFT)**: Using LoRA adapters and small-scale supervised tuning on VHDL-specific corpora.

The primary metrics include:

### Functional Correctness: Pass@k
Defined for k sampled generations per task:
\[
\text{Pass@}k = 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}}
\]
where $n$ is the number of generations and $c$ is the number of correct solutions [2406.04379, 2606.13735].

### Self-Consistency ($\mathrm{SC}_1$)
Measures whether the round-trip process (code → summary → regenerated code) yields functionally equivalent results, using equivalence checking against the canonical solution [2507.12308].

### ROUGE-L and LLM Preference Rate (PR)
ROUGE-L measures the longest common subsequence between generated and reference summaries, while PR quantifies human or LLM preference for generated outputs over references in summarization tasks [2507.12308].

### Structural Metrics
Tree-sitter–based weighted sequence and Jaccard similarities quantify the structural preservation between generated and reference code [2606.13735].

## 4. Baseline Performance and Error Profiles

Empirical results demonstrate that even state-of-the-art LLMs exhibit limited proficiency in VHDL code generation:
- **Zero-shot Pass@1** values rarely exceed 23% for leading models (e.g., CodeLlama-70B-instruct, Granite-Code-34B) [2406.04379, 2507.12308].
- **Pass@5** remains below 40% for most models; Gemini 3 Pro Preview achieves ≈48% Pass@1 and ≈78% Pass@5 [2606.13735].
- **Common error modes**: syntax errors (missing or mismatched library/use clauses), port and type mismatches, misconfigured sensitivity lists, off-by-one-indexing in vectors, incomplete FSM logic, missing signal initializations [2406.04379, 2606.13735].

The self-consistency metric ($\mathrm{SC}_1$) further reveals that bidirectional transformations (code-summarization-regeneration) are a principal failure point, with leading models achieving only ≈23.5% in baseline settings [2507.12308].

## 5. Methodological Advances: Chain-of-Descriptions and Validation Pipelines

The Chain-of-Descriptions (CoDes) strategy introduces structured intermediate planning steps—such as explicit design decompositions and synthesized high-level explanations—prior to code emission. This multi-step planning notably increases Pass@1 and $\mathrm{SC}_1$ (e.g., Pass@1 (TB) from 19.2% to 25.4% for Granite-Code-34B), confirming that stepwise prompt construction and design intent clarification are critical for LLM VHDL competence [2507.12308].

Related frameworks, including VHDLSuite, present automated pipelines that enable large-scale Verilog-to-VHDL task synthesis, iterative simulation-driven repair, and robust error categorization. These pipelines apply LLMs in a generate→simulate→repair loop, leveraging feedback from VUnit/GHDL to refine outputs and filter only those that are functionally validated [2606.13735].

## 6. Diagnostic Analysis and Error Taxonomy

Sophisticated failure taxonomy reveals VHDL-specific challenges: uninitialized signals, strict type resolution requirements, declaration order constraints, and initialization semantics enforce a discipline absent from most programming language tasks. Automated classifiers segment failures into six categories: declaration errors, type/data object mismatches, instantiation/interface errors, combinational and sequential logic bugs, and miscellaneous/initialization problems [2606.13735]. 

Notably, successful generations show significantly higher structural similarity to reference code, whereas compile-time and runtime failures are associated with low syntactic correspondence.

## 7. Comparative Perspectives and Future Directions

Comparison with alternative frameworks, such as LLHD—a multi-level SSA-based IR—clarifies that operational semantics, strict event-queue scheduling, and delta-phase resolution are foundational for cycle- and delta-accurate simulation in VHDL evaluation [2004.03494]. VHDL-Eval and VHDLSuite each highlight that language-specific rigor and semantics—such as nine-valued logic and explicit process-wait constructs—compound the design complexity relative to Verilog-centric benchmarks. Iterative validation/repair and feedback-driven fine-tuning are repeatedly shown to be essential for progress in LLM-driven hardware code synthesis [2406.04379, 2606.13735].

Emergent trends suggest expanding task corpora, advancing in-context learning paradigms, automating testbench and golden reference generation, and integrating formal equivalence checking as promising routes for closing the performance gap. Open-sourcing of benchmark suites and evaluation harnesses is anticipated to accelerate the systematic development and assessment of VHDL-aware LLMs [2606.13735].

---

**Key References:**  
- [2406.04379] VHDL-Eval: A Framework for Evaluating Large Language Models in VHDL Code Generation  
- [2507.12308] Chain-of-Descriptions: Improving Code LLMs for VHDL Code Generation and Summarization  
- [2606.13735] VHDLSuite: Unified Pipeline for L

Source: https://www.emergentmind.com/topics/vhdl-eval