VHDL-Xform: Assessing Functional Equivalence
- VHDL-Xform is a benchmark dataset designed to evaluate language models' ability to recognize functionally equivalent VHDL code despite lexical differences.
- It employs systematic transformations—including identifier renaming, inert statement insertions, and cross-language back-translation—to ensure semantic equivalence.
- Evaluation using metrics like ROUGE-L and LLM Preference Rate, along with Chain-of-Descriptions improvements, reveals both challenges and performance gains in VHDL code summarization.
VHDL-Xform is a benchmark dataset and methodology developed to rigorously evaluate LLMs' ability to recognize and reason about functional equivalence in VHDL (VHSIC Hardware Description Language) code. It targets the specific challenge of distinguishing functionally identical yet lexically dissimilar VHDL snippets, addressing limitations of existing models on tasks related to code generation and summarization for hardware description languages (Vijayaraghavan et al., 16 Jul 2025).
1. Formal Definition and Equivalence Semantics
VHDL-Xform is formally defined over the set of all valid VHDL code snippets. Three classes of transformation functions, , are considered:
- : Type-2 transformations (identifier renaming)
- : Type-3 transformations (statement-level inert variations)
- : Type-4 transformations (semantic rewriting via cross-language back-translation)
For any and , a transformed maintains semantic equivalence: , where represents the mapping from input vectors to output vectors over all valid VHDL test sequences. The dataset is defined as: 0 This construction enforces that all (original, transformed) pairs are functionally indistinguishable at the behavior level, a key feature for evaluating the robustness of code understanding in models (Vijayaraghavan et al., 16 Jul 2025).
2. Construction Methodology and Transformation Categories
VHDL-Xform is built from permissively licensed public VHDL code repositories. Three transformation pipelines are applied:
- Type-2 (Identifier Renaming): Each code entity is parsed to extract names, then identifiers are systematically renamed using patterns (e.g., variable-to-single-character, snake to camel case), typically with LLM-assisted suggestions.
- Type-3 (Statement-Level Variations): Functionality-preserving changes are made by inserting inert statements (e.g., comments, unused signals) and reordering declarations that do not affect design semantics or compilation.
- Type-4 (Semantic Rewriting via Back-Translation): The snippet is compiled into Verilog (using GHDL and Icarus Verilog), then re-synthesized back to VHDL. This process introduces deeper structural and naming variations, often creating new temporary state variables and minimizing direct lexical overlap compared to the original.
Functional equivalence for all pairs is verified either through sequential equivalence checking or via internal testbenches, ensuring no behavioral divergence is introduced (Vijayaraghavan et al., 16 Jul 2025).
3. Representative Transformation Examples
Three transformation classes are realized concretely within the VHDL-Xform corpus:
- Type-2 Example:
- Original:
- 9
- Transformed:
- 0
- Type-3 Example:
- Original:
- 1
- Transformed:
- 2
- Type-4 Example:
- Original:
- 3
- Transformed: (after VHDL→Verilog→VHDL)
- 4
- These exemplify systematic code drift while preserving semantics (Vijayaraghavan et al., 16 Jul 2025).
4. Evaluation Protocols and Metrics
Two principal metrics are used to assess VHDL code summarization consistency over VHDL-Xform:
- ROUGE-L (1): Measures the longest common subsequence (LCS) between generated (2) and reference (3) summaries, using the 4 score:
5
Where 6 and 7.
- LLM Preference Rate (PR): For each (summary, reference) pair, a judge LLM (Llama-3-70B-Instruct) is queried: "Which summary better captures the functionality?" PR is the fraction of times the generated summary is preferred.
Performance for various models under zero-shot conditions and with Chain-of-Descriptions (CoDes) improvements is summarized below:
| Model | ROUGE-L (Zero-Shot) | PR (%) (Zero-Shot) |
|---|---|---|
| CodeLlama-34B | 36.63 | 36.6 |
| Granite-Code-34B | 38.60 | 35.7 |
| Deepseek-33B | 35.10 | 28.6 |
| Granite-Code-20B | 31.70 | 28.1 |
| Mixtral-8x7B | 19.60 | 14.8 |
| Mistral-7B | 18.20 | 6.6 |
| Granite-Chat-13B | 28.10 | 19.5 |
| Llama-2-Chat-13B | 26.60 | 18.5 |
The ROUGE-L range is typically 20–40%, showing significant challenges for current models (Vijayaraghavan et al., 16 Jul 2025).
5. Chain-of-Descriptions (CoDes) and Performance Gains
CoDes is a multi-step prompting paradigm designed to inject a series of intermediate, descriptive steps into LLM workflows. For code summarization on VHDL-Xform, this involves decomposing the task into a series of explanations or plans—which are then concatenated with the input and supplied to the model.
Empirical results demonstrate notable, though model-dependent, improvements:
| Model | ROUGE-L (ZS→CoDes) | Δ8 | PR (ZS→CoDes) | ΔPR |
|---|---|---|---|---|
| CodeLlama-34B | 36.63→39.20 | +2.57 |