VeriPPA: LLM-Driven Verilog PPA Optimization
- VeriPPA is a two-stage, feedback-driven framework that generates Verilog by first ensuring syntactic and functional correctness through iterative simulation-based repair.
- It refines functionally correct RTL with synthesis-driven power, performance, and area (PPA) optimization using detailed reports and in-context learning.
- Experimental results demonstrate significant improvements in syntax and functionality metrics over baseline LLM methods, highlighting its potential in automated hardware design.
VeriPPA is a two-stage, feedback-driven framework for large-language-model-based Verilog generation that couples correctness refinement with synthesis-informed Power, Performance, and Area optimization. In the hardware-design literature, the term denotes a system that takes a textual hardware specification , generates Verilog , first improves syntactic and functional correctness through iterative simulation-driven repair, and then refines the functionally correct RTL under explicit PPA constraints using synthesis reports and in-context learning. The framework is presented explicitly as VeriPPA in "LLM-VeriPPA: Power, Performance, and Area Optimization aware Verilog Code Generation with LLMs" (Thorat et al., 10 Sep 2025) and in the earlier open-source formulation "Advanced LLM-Driven Verilog Development: Enhancing Power, Performance, and Area Optimization in Code Synthesis" (Thorat et al., 2023). A later closely related system, "VeriOpt: PPA-Aware High-Quality Verilog Generation via Multi-Role LLMs" (Tasnia et al., 20 Jul 2025), does not use the exact name VeriPPA, but is explicitly described as a PPA-aware Verilog-generation system of the same general type.
1. Problem setting and conceptual scope
VeriPPA is motivated by a gap in prior LLM-for-Verilog work: many earlier methods emphasize prompt engineering, fine-tuning for syntax and functionality, or benchmark gains, but do not address PPA optimization, even though a design that passes syntax and functional checks may still be poor in clock speed, power, or area (Thorat et al., 10 Sep 2025). The framework therefore adopts the principle that before optimizing PPA, one must first produce correct Verilog, and only then apply constraint-aware optimization.
In its canonical form, VeriPPA comprises four architectural elements: a natural-language hardware description , an LLM generator, a verification loop based on ICARUS Verilog and design-specific testbenches, and a PPA loop based on Synopsys Design Compiler on the ASAP7 7nm Predictive PDK (Thorat et al., 2023). The overall control structure is multi-round conversation with error feedback: generate RTL, inspect simulator or synthesis output, revise the RTL, and repeat until either constraints are satisfied or an iteration limit is reached.
This formulation places VeriPPA between benchmark-oriented code synthesis and EDA-aware design automation. Its central claim is not merely that LLMs can emit parsable HDL, but that they can participate in a closed-loop workflow spanning simulation, synthesis, and iterative architectural refinement.
2. Stage 1: VeriRectify and iterative correctness repair
The first stage of VeriPPA is VeriRectify, which targets syntactic correctness and functional correctness through recursive repair. The framework runs the generated Verilog against ICARUS Verilog and a design-specific testbench , then feeds simulator diagnostics back into the LLM for another correction round (Thorat et al., 10 Sep 2025).
The refinement loop is formalized as
where is the code at iteration , is the set of detected errors, detects errors, and is the LLM-based repair function. The loop terminates when 0 or the maximum iteration budget 1 is reached. The earlier paper also gives the symbolic relation
2
where the subtraction denotes removal or correction of the identified errors rather than literal arithmetic subtraction (Thorat et al., 2023).
The distinctive feature of VeriRectify is the granularity of feedback. Rather than re-prompting blindly, the framework supplies exact simulator error messages and precise failure locations from the testbench. The papers describe this as analogous to human debugging: diagnose, fix, re-run, and repeat. In the reported experiments, the setup uses up to four correction rounds, with 3, temperature 4, and context length 5 (Thorat et al., 10 Sep 2025).
The correctness metrics follow the benchmark conventions used in the cited work. Syntax correctness, or “linguistic accuracy,” means that the generated Verilog compiles or parses successfully. Functional correctness, or “operational efficacy,” is counted when a generated design passes the functionality test; the framework notes that this follows RTLLM’s evaluation style (Thorat et al., 2023).
3. Stage 2: synthesis-guided PPA optimization
Once Stage 1 yields functionally correct RTL, VeriPPA enters a second stage devoted to PPA optimization. The synthesized design is compiled with Synopsys Design Compiler using the command compile_ultra, targeting the ASAP 7nm Predictive PDK, and the framework extracts clock period (ps), power 6, and area 7 (Thorat et al., 10 Sep 2025).
The stage is formalized as
8
The optimization criterion is therefore constraint satisfaction and feedback-based refinement rather than a scalar learned objective. The papers explicitly note that VeriPPA does not define a formal optimization objective such as 9; it keeps a design if the PPA goal is met and otherwise revises it (Thorat et al., 2023).
The PPA-aware prompt is built through in-context learning (ICL) using example pairs of Verilog code, corresponding PPA reports, and optimization strategies. The cited optimization ideas include pipelining, clock gating, parallel operation, and hierarchical design (Thorat et al., 10 Sep 2025). In the earlier formulation, ICL is motivated by scarce labeled Verilog data and the risk that fine-tuning may overfit. That paper also gives the token-level ICL expression
0
with demonstrations chosen to cover diverse design types such as addition, multiplication, single-stage designs, and pipelined designs (Thorat et al., 2023).
A representative example is adder_32bit: the paper states that a version synthesizing at 500 ps can be reworked under an application requirement of under 300 ps, with the optimized RTL reaching 180 ps (Thorat et al., 10 Sep 2025). The mechanism is not framed as autonomous physical reasoning from raw RTL; rather, it is a prompt-based loop in which synthesis outcomes are exposed to the model and used to drive architectural revision.
4. Benchmarks, models, and reported results
VeriPPA is evaluated on RTLLM and VerilogEval. The RTLLM setting uses 29 designs. The VerilogEval setting is split into VerilogEval-human: 156 designs and VerilogEval-machine: 108 designs (Thorat et al., 10 Sep 2025). The earlier paper notes that RTLLM originally contained 30 designs, but risc_cpu is unavailable, leaving 29 designs in the reported evaluation (Thorat et al., 2023).
The experimental model set includes GPT-3.5, GPT-4, GPT-4o, Llama-2-7B, Llama-3-8B, CodeLlama-7B, Llama3.1-405B, RTL-Coder, and DeepSeek Coder (Thorat et al., 10 Sep 2025). The hardware platform for the later evaluation is a Linux system with AMD EPYC 7543 and NVIDIA A100 80GB.
On RTLLM, the headline result reported for VeriPPA is 81.37% syntactic correctness and 62.06% functional correctness, stated as outperforming prior SOTA by +8.3% syntax and +16% functionality (Thorat et al., 10 Sep 2025). The detailed GPT-4 comparisons on RTLLM include 66.20% syntax / 37.93% functionality without VeriPPA versus 81.37% syntax / 48.27% functionality with VeriPPA; for GPT-4 (4-shot) with RTLLM-style prompting, the comparison is 66.89% syntax / 44.82% functionality versus 81.37% syntax / 62.06% functionality (Thorat et al., 10 Sep 2025). The earlier paper reports closely aligned RTLLM improvements, including GPT-4 syntax correctness rising from 66.2% to 81.37% and functionality from 37.93% to 48.27%, with the combined method reaching 62.06% after four attempts (Thorat et al., 2023).
On VerilogEval-Machine, VeriPPA reports 99.56% syntactic correctness and 43.79% functional correctness for GPT-4, compared with the cited baseline values 92.11% syntax and 33.57% functionality. On VerilogEval-Human, GPT-4 improves from 91.28% syntax / 29.48% functionality to 97.17% syntax / 39.74% functionality (Thorat et al., 10 Sep 2025).
The PPA results are presented as optimized outcomes for selected designs. Reported optimized values include adder_32bit at 180 ps, 587.31 1, 1005.67 2; multi_booth at 123.2 ps, 42.39 3, 42.92 4; pe at 325 ps, 1206.0 5, 4863.88 6; and asyn_fifo at 114.8 ps, 988.92 7, 1344.86 8 (Thorat et al., 10 Sep 2025).
A separate efficiency comparison against self-planning in a one-iteration GPT-4o setting reports the following: self-planning uses 647589.84 MACs and 317446 tokens, with 77.93% syntax and 41.37% functionality; VeriPPA uses 552317.76 MACs and 270744 tokens, with 80.68% syntax and 41.37% functionality. The paper summarizes this as 46,702 fewer tokens and 95,272.08 trillion MACs less computation while improving syntax and matching functionality (Thorat et al., 10 Sep 2025).
5. Relation to VeriOpt and later PPA-aware RTL generation
A later paper, "VeriOpt: PPA-Aware High-Quality Verilog Generation via Multi-Role LLMs" (Tasnia et al., 20 Jul 2025), is not named VeriPPA, but the authors explicitly state that if one asks about “VeriPPA,” the closest accurate interpretation is that VeriOpt is the full framework and its PPA-aware prompting and synthesis-feedback loop is the component most naturally associated with that label. This establishes a direct line of continuity between the two-stage VeriPPA conception and a later multi-role variant.
VeriOpt preserves the correctness-first, PPA-aware-refinement structure but decomposes generation into four explicit roles: Planner, Programmer, Reviewer, and Evaluator. The Planner breaks a natural-language RTL specification into granular steps such as internal registers and parameters, FSM states, number of always blocks, and blocking versus non-blocking assignments. The Programmer writes Verilog while mapping code regions to planning steps. The Reviewer checks whether each planned step appears in the RTL and whether any step was skipped or changed. The Evaluator runs syntax and functional checks against a testbench and feeds error logs back into the loop (Tasnia et al., 20 Jul 2025).
The PPA stage in VeriOpt also broadens the optimization vocabulary. It injects domain knowledge for power optimization—including clock gating, power gating, operand isolation, register update suppression, and conditional accumulation—performance optimization—including pipelining, retiming, loop unrolling, logical effort, and path restructuring—and area optimization—including resource sharing, FSM state encoding, logic consolidation, register optimization, and technology mapping (Tasnia et al., 20 Jul 2025). It further supplies granular synthesis-report context such as dynamic power, leakage power, internal or cell switching power, cell area, design area, critical path length, critical path slack, total negative slack, and levels of logic.
Quantitatively, VeriOpt reports 25/29 successful designs and 100% syntactic correctness, outperforming several correctness baselines, and it reports up to 88% power reduction, 76% area reduction, and 73% timing improvement across RTLLM benchmarks (Tasnia et al., 20 Jul 2025). For a 16-bit adder case study, the paper lists 58.70% reduction in cell area, 59.11% reduction in design area, 57.03% reduction in dynamic power, 61.21% reduction in leakage power, and 58.40% reduction in critical path length. This suggests that the literature around VeriPPA has evolved toward richer role decomposition and more explicit use of synthesis evidence, while retaining the same central thesis: correctness alone is insufficient for industrial RTL.
6. Limitations, caveats, and significance
The VeriPPA papers identify several limitations. PPA optimization depends on the availability of an overview flow and design-specific testbenches. Results are tied to Synopsys Design Compiler and ASAP7, so portability may vary. The correction loop is bounded by a maximum number of rounds 9, and repeated outputs can reduce correction efficiency. Smaller open-source LLMs may show syntax gains while still achieving 0% functional correctness on RTLLM in the reported settings (Thorat et al., 10 Sep 2025).
The earlier and later papers also converge on a broader methodological caveat: optimization is report-driven and heuristic rather than a formal closed-loop mathematical objective (Thorat et al., 2023, Tasnia et al., 20 Jul 2025). VeriOpt adds that some designs are already close to optimal, some optimizations improve one metric while slightly worsening another, leakage power can increase in some cases, and some metrics are N/A for purely combinational designs (Tasnia et al., 20 Jul 2025). The earlier VeriPPA paper similarly notes that some designs, such as radix2_div, could not be optimized successfully (Thorat et al., 2023).
The significance of VeriPPA lies in its reframing of LLM-generated RTL as a mutable artifact inside an EDA loop. The framework turns Verilog generation into a closed-loop design-refinement system: generate RTL, detect syntax and functional errors, correct them with simulator feedback, synthesize, evaluate clock, power, and area, and optimize again under explicit constraints (Thorat et al., 10 Sep 2025). In that sense, VeriPPA marks a shift from one-shot HDL code completion toward synthesis-informed, constraint-aware hardware design automation.