---
title: 'SPICE Bench: Benchmark Suites in EDA & SWE'
url: https://www.emergentmind.com/topics/spice-bench
type: topic
---

# SPICE Bench: Benchmark Suites in EDA & SWE

Searching arXiv for recent papers that use or define “SPICE Bench” and closely related benchmark frameworks.
Tool call: arxiv_search(query="SPICE Bench benchmark SPICEAssistant Masala-CHAI MemSPICE", max_results=10, sort_by="relevance")
Searching arXiv for exact and near-exact matches to “SPICE Bench”.
Tool call: arxiv_search(query="\"SPICE Bench\"", max_results=10, sort_by="relevance")
In recent arXiv usage, **SPICE Bench** does not denote a single standardized artifact. The name is applied to several distinct benchmark suites and SPICE-centered evaluation workflows across electronic design automation and software engineering. In circuit-design work, it refers to a SPICE-based benchmarking flow for memristor-based logic-in-memory, a 256-question LTSpice benchmark for adapting switched-mode power-supply netlists, and an open schematic-to-netlist suite for analog-circuit generation and verification. In software engineering, it denotes a corpus of 6 802 labeled SWE-bench-like repair tasks curated from 291 open-source projects in SWE-Gym [2309.04868] [2507.10639] [2411.14299] [2507.09108].

## 1. Nomenclature and domain scope

The shared label reflects a common emphasis on **benchmarked interaction with SPICE artifacts or SPICE-style evaluation loops**, but the concrete object being benchmarked varies substantially by paper. In some cases the benchmark target is a circuit or workload executed in a simulator; in others it is an LLM’s ability to generate or modify netlists; in software engineering it is a labeled dataset whose name is unrelated to electronic simulation.

| Usage in the literature | Benchmark object | Reported scale |
|---|---|---|
| MemSPICE-style “SPICE Bench” | memristor-based LiM evaluation flow | six ISCAS’85 circuits in a single 512-device row |
| SPICEAssistant “SPICE Bench” | LTSpice-based SMPS adaptation tasks | 256 questions |
| Masala-CHAI “SPICE Bench” | analog schematic-to-netlist generation and verification | ≈2,100 schematics automatically converted into SPICE netlists |
| SPICE Bench in software engineering | SWE-bench-like repair tasks with automated labels | 6 802 instances from 291 projects |

A recurring source of confusion is that the same label spans at least two unrelated research lineages: **SPICE as circuit simulation infrastructure** and **SPICE as a software-engineering labeling pipeline**. This suggests that the phrase functions primarily as a project-specific benchmark name rather than as a field-wide standard.

## 2. Test-bench antecedents in circuit simulation

Before the appearance of benchmark suites explicitly named SPICE Bench, SPICE-centered research often used the term **test bench** to denote a validated simulation environment. A representative example is the LTspice and SimScape power-cycling test bench for accelerated life testing of the 1.2 kV SiC-MOSFET Cree CAS300M12BM2 in harsh offshore environments [2001.05580].

That bench comprises a **DC voltage source** with \(V_{\mathrm{DC}} \approx 50\ \mathrm{V}\), a **device-under-test** subcircuit for the CAS300M12BM2, a **series load resistor**, **gate-drive logic**, a **thermal network**, and **measurement and data capture** for \(V_{\mathrm{DS(on)}}\), \(I_D\), and \(T_j(t)\). The subcircuit uses gradual-channel MOSFET equations with temperature dependence, voltage- and temperature-dependent capacitances \(C_{gd}\), \(C_{gs}\), and \(C_{ds}\), and a thermal network with \(R_{\theta JC}=0.15\ \mathrm{K/W}\), \(R_{\theta CA}=2.5\ \mathrm{K/W}\), and \(C_{\mathrm{th,jc}} \approx 1\ \mathrm{J/K}\). Power-cycling profiles use \(t_{\mathrm{on}}=1\ \mathrm{s}\), \(t_{\mathrm{off}}=1\ \mathrm{s}\), ambient temperatures from \(25^\circ\mathrm{C}\) to \(150^\circ\mathrm{C}\), and failure thresholds such as a 20% increase in \(R_{\mathrm{DS(on)}}\) [2001.05580].

The same work formalizes degradation and lifetime extraction through a power-law drift model,
\[
\Delta R(n)=B\cdot n^\gamma,
\]
and a Coffin–Manson type relation,
\[
N_f = A\cdot (\Delta T_j)^{-\alpha},
\]
with remaining useful life estimated as \(RUL(n_0)=n_{\mathrm{fail}}-n_0\). Validation is reported against the datasheet: static \(I\)-\(V\) curves show maximum error below 5%, \(R_{\mathrm{DS(on)}}\) versus temperature matches within \(\pm 3\%\), and thermal impedance matches Cree’s \(R_{\theta JA}\) specification within 10%; early hardware data show \(R_{\mathrm{DS(on)}}\) drift trends within 15% over \(1\ \mathrm{k}\) cycles [2001.05580].

This antecedent is significant because it locates the later notion of SPICE Bench within a broader tradition of **simulation-grounded, measurement-oriented, and validation-heavy** circuit research.

## 3. MemSPICE and SPICE-level benchmarking for logic-in-memory

In "MemSPICE: Automated Simulation and Energy Estimation Framework for MAGIC-Based Logic-in-Memory" [2309.04868], the term **“SPICE Bench”** is used for a stepwise benchmarking flow that starts from HDL and ends with per-benchmark energy numbers for memristor-based logic-in-memory on a crossbar.

The flow begins with an **automated HDL→SPICE-netlist pipeline**. Verilog or VHDL RTL is synthesized with **ABC** into 2-input NOR and NOT primitives, then converted by **SIMPLER** into a JSON mapping that records row size, inputs, outputs, and an ordered execution sequence of INIT, NOR, and NOT steps. A JSON→SPICE translator emits a **single-row crossbar subcircuit** with memristor instances \(I_0 \ldots I_{N_{\mathrm{col}}-1}\), relay-controlled row and column access, PWL sources for operation pulses, and a transient analysis with `.save` of device currents [2309.04868].

Automatic testbench generation derives one PWL waveform file per net. INIT drives the target column to \(V_{\mathrm{INIT}}\), exemplified as \(2.0\ \mathrm{V}\), for \(\tau_{\mathrm{INIT}}\); NOR and NOT operations apply \(V_{\mathrm{OPCODE}}\) on input nets and \(0\ \mathrm{V}\) on the output net, while selected relays are asserted at \(V_{\mathrm{SW}}=2.0\ \mathrm{V}\). Simulations use the **VTEAM memristor model**, relay elements with \(r_{\mathrm{open}}=1\ \mathrm{T}\Omega\) and \(r_{\mathrm{closed}}=1\ \Omega\), `.tran` for transient execution, and optional `.dc` sweeps for read-level profiling [2309.04868].

Energy estimation is defined directly at the SPICE waveform level. For device \(i\),
\[
E_i = \int_{t_0}^{t_f} V_i(t)\, I_i(t)\, dt,
\]
and the total energy is
\[
E_{\rm total} = \sum_{i=1}^{N_{\rm devices}} \int_{t_0}^{t_f} V_i(t)\,I_i(t)\,dt.
\]
MemSPICE supports numerical integration through Python or Matlab and also auto-generates a Spectre-compatible `.ocn` file so that Spectre can print per-device integrals [2309.04868].

The benchmark results reported for six ISCAS’85 circuits, all mapped into a single 512-device row, are: **c17** at \(1.16\ \mathrm{pJ}\), **c432** at \(1157\ \mathrm{pJ}\), **c499** at \(1821\ \mathrm{pJ}\), **c880** at \(1842\ \mathrm{pJ}\), **c1908** at \(1876\ \mathrm{pJ}\), and **c3540** at \(2714\ \mathrm{pJ}\), with total SPICE wall-clock time per benchmark of 2–20 minutes on a modern workstation. The paper further notes that average energy per logic operation is typically 2–3 pJ/op for medium-size designs, although re-initialization overhead can dominate, and that input-pattern variation is \(\pm 5\)–10% [2309.04868].

Within the EDA literature, this version of SPICE Bench is best understood as a **reproducible SPICE workflow** rather than merely a dataset.

## 4. SPICEAssistant and the 256-question LTSpice benchmark

"SPICEAssistant: LLM using SPICE Simulation Tools for Schematic Design of Switched-Mode Power Supplies" defines **SPICE Bench** as a benchmark of **256 questions** testing an LLM’s ability to adapt circuit netlists to fulfill switched-mode power-supply design tasks [2507.10639].

The benchmark spans three topology tiers: **Easy** with 72 cases based on an idealized general buck converter; **Medium** with 72 cases using a buck converter with the **LTC3419 dual-step-down IC**; and **Hard** with 112 cases using a **2-phase synchronous buck with LTC7802 controller**. Task categories comprise **32 topology-adaptation questions** and **224 parameter-tuning questions**. The rationale is to mirror a common SMPS workflow in which one starts from a typical application circuit in a datasheet and then either changes component values or slightly rewires pins to achieve new specifications [2507.10639].

All simulations are performed in **LTSpice** with a transient directive
```text
.TRAN 0 5 ms 0 1 µs startup
```
and, where needed, `.AC DEC 100 10 Hz 10 MHz` and `.DC Iload 0 2 A 0.01 A`. `.OPTIONS POST=2 NOMOD` controls data export. Performance metrics are extracted through `.MEAS` commands such as average output voltage, peak-to-peak ripple, switching frequency, and settling time to 90% of final value. The framework’s Python reading tools invoke these measurements so that the LLM does not parse raw waveforms or plotted images [2507.10639].

Evaluation is organized as an **LLM↔Simulator loop**. The LLM proposes a netlist change; the framework writes the modified netlist, calls LTSpice, extracts requested metrics, and returns numeric feedback; the process repeats for up to \(N=5\) iterations, after which solve rate plateaus. Scoring uses **Solve Rate (SR)**, the **Absolute Percentage Error**
\[
APE_i=\left|\frac{A_i-F_i}{A_i}\right|\times 100\%,
\]
and \(\mathrm{median}(APE)\) over non-topology tasks [2507.10639].

The reported results show a large separation between a standalone LLM and simulator-assisted operation. On the full 256-question benchmark, **GPT-4o** attains **14.8% SR** and **64.3% APE**, whereas **SPICEAssistant** attains **53.0% SR** and **4.2% APE**. By difficulty level, the Easy suite improves from **29.0% ± 1.0%** to **64.9% ± 1.8%** SR; the Medium suite from **12.5%** to **50.0%**; and the Hard suite from **7.1%** to **47.3%**. The paper reports an **absolute solve-rate gain of 38.2 percentage points**, a drop of median APE from **64% to 4%**, and only marginal improvement from adding RAG alone [2507.10639].

This SPICE Bench instantiation is therefore a **closed-loop LLM evaluation benchmark** defined by repeated simulation, automatic measurement extraction, and tolerance-based success criteria.

## 5. Analog-circuit netlist generation: Masala-CHAI and related SPICEPilot metrics

In "Masala-CHAI: A Large-Scale SPICE Netlist Dataset for Analog Circuits by Harnessing AI," SPICE Bench denotes a benchmark suite for **analog-circuit design and verification** built from automatically generated SPICE netlists and their associated schematics [2411.14299].

The benchmark currently comprises **≈2,100 analog-circuit schematics automatically converted into SPICE netlists**, with an open-source framework that users can grow to **≈7,500 or more**. Covered topologies range from **single-stage RC filters and basic common-source amplifiers** with **2–5 devices** and **≤10 nets** to **two-stage op-amps, differential amplifiers, LC oscillators, and bandgap references** with **20–40+ devices** and **15–30 nets**. Device naming conventions follow standard SPICE forms such as **R#**, **C#**, **L#**, **M#**, **Q#**, **V#**, and **I#**; node numbering uses integer nets with ground at node 0; values use unit-suffixed syntax such as `10k`, `1p`, and `100n`; and MOSFET instance lines specify `W=` and `L=` parameters [2411.14299].

Its three-step workflow consists of **labeling analog circuits**, **prompt tuning for LLMs**, and **SPICE netlist verification**. Labeling uses a **YOLOv8 object detector** trained on **∼4,300 textbook images** across **12 classes** and a **Deep Hough Transform** plus heuristic clustering with endpoint radius \(r=40\) px for net detection. Prompt tuning addresses component differentiation and net-annotated SPICE generation. Verification checks for floating nets, missing multi-terminal connections, and basic simulation sanity in **NGSpice/XYCE**, re-prompting the LLM until all checks pass [2411.14299].

Benchmark correctness is formalized through a **graph-based netlist metric**. For netlist graphs \(G_1\) and \(G_2\), similarity is
\[
S=\left(1-\frac{\mathrm{GED}(G_1,G_2)}{\mathrm{GED}_{\max}}\right)\times 100\%,
\]
where
\[
\mathrm{GED}_{\max}=|V_1|+|E_1|+|V_2|+|E_2|.
\]
On 20 benchmark designs with 10 samples each, fine-tuning improves correctness from **8/10 to 10/10** on a common-source amplifier for both GPT-3.5-turbo and GPT-4o-mini, from **6/10 to 10/10** for GPT-4o-mini on a single-stage RC low-pass filter, and from **0/10 to 7/10** for GPT-4o-mini on a 2-stage op-amp with compensation; the bandgap reference remains at **0/10** [2411.14299].

A closely related but separately named benchmark framework is **SPICEPilot**, which standardizes evaluation for LLM-based SPICE code generation through the metrics **Syntax Correctness Rate (SCR)**, **Netlist Functional Coverage (NFC)**, **Simulation Fidelity (SF)**, and **Pass@k** [2410.20553]. SPICEPilot’s dataset contains **60 unique circuits** across digital and analog categories, with **51 transistor-based “easy” to “extreme” tasks**, transistor counts from **4 up to 88**, supply voltage fixed at **1.2 V**, and temperature at **27 °C**. It reports, for example, **Pass@1 = 79.8%** and **Pass@5 = 85.5%** for **SPICEPilot+Claude** on the 24-circuit AnalogCoder benchmark, and average **Pass@1 = 80%** and **Pass@3 = 88%** for **SPICEPilot+GPT-4o** on a 25-circuit mixed benchmark [2410.20553].

Taken together, Masala-CHAI and SPICEPilot establish a methodological axis in which SPICE Bench denotes either a **curated benchmark corpus** or a **metric-standardized benchmarking framework** for schematic interpretation and netlist synthesis.

## 6. SPICE Bench in software engineering

A separate usage appears in "SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation," where **SPICE Bench** is a dataset of **6 802 labeled “SWE-bench-like” repair tasks** drawn from **291 real-world open-source projects in SWE-Gym** [2507.09108].

Each instance includes the **GitHub issue title and body**, the **golden patch**, the **test patch**, and three automated labels with rationales: **Issue Clarity**, **Test Coverage**, and **Effort Estimation**. The pipeline comprises **Context-Aware Code Navigation (via Aider)**, **Rationale-Driven Prompting**, and **Multi-Pass Consensus (Self-Consistency)**. Context is restricted to files touched by the golden patch and test patch, with **RepoMap** using ctags, Tree-sitter, and graph ranking; prompting is derived from **20 fully-agreed SWE-V instances**; and consensus fuses three stochastic runs by majority vote or, if no majority exists, by the median ordinal label [2507.09108].

The paper formalizes binary label discretization from human 0–3 scores, reports inter-rater agreement with **Krippendorff’s \(\alpha\)**,
\[
\alpha = 1 - \frac{\sum_{i,j}o_{ij}\,\delta^2(i,j)}{\sum_{i,j}e_{ij}\,\delta^2(i,j)},
\]
and gives a cost model
\[
C = n\,m\,(c+n_c)\,\frac{t+n_t}{60}.
\]
Using \(m=4\), \(c=\$75/\mathrm{hr}\), and \(t=20\) minutes per instance yields a manual baseline of approximately **\$100 000 per 1 000 instances** [2507.09108].

Relative to **SWE-bench Verified**, which contains **500 verified instances** after **1 699 initially labeled** examples, SPICE Bench is approximately **13.6× larger**. On a 110-instance SWE-V test set, reported accuracies include **87.3% ICA** for **GPT-4o-mini**, **68.2% TCA** for **GPT-4o**, and **68.5% TCA** for **DeepSeek-Reasoner**. On a 48-instance manual agreement-based evaluation, **ICA label correctness** is **93.5%** and **TCA label correctness** is **60%**. Median per-1 000-instance API cost is reported as **\$3.29** for GPT-4o-mini, **\$4.32** for DeepSeek-Reasoner, **\$43.92** for GPT-4.1, and **\$54.91** for GPT-4o, with the default combined configuration at approximately **\$5.10 per 1 000 instances**; total time for 1 000 instances is **≈25 h** versus **1 333 h** manual, a **×53 speedup** [2507.09108].

This use of the term is conceptually distinct from circuit-simulation SPICE Bench. Here, the benchmark object is a **large labeled corpus for SE-focused foundation models**, not a simulator-driven electronic-design task.

## 7. Methodological commonalities and terminological boundaries

Across these otherwise heterogeneous uses, several common patterns recur. First, SPICE Bench almost always involves **automation of a translation layer**: HDL to SPICE netlists in MemSPICE, natural-language design requests to netlist edits in SPICEAssistant, schematic images to SPICE netlists in Masala-CHAI, and repository state to task labels in the software-engineering SPICE pipeline. Second, the evaluation loop is typically **programmatic and tool-mediated**, whether through `.tran` runs and current integration, `.MEAS` extraction, graph-edit comparison, or consensus fusion over repeated model calls. Third, success is operationalized through **explicit quantitative metrics** rather than informal inspection [2309.04868] [2507.10639] [2411.14299] [2507.09108].

A plausible implication is that the phrase **SPICE Bench** now denotes a family of benchmark designs centered on **closed-loop validation against executable artifacts**. In circuit work, the executable artifact is the SPICE simulation itself; in software engineering, it is a structured instance enriched by automated labels and rationales. This convergence explains why the same name can recur across otherwise unrelated subfields.

The term should also be distinguished from **benchmark studies of the Solar Orbiter SPICE instrument**. Brooks et al. benchmarked Solar Orbiter/**SPICE** against **Hinode/EIS** for plasma-composition measurements in active region AR 12781, using spectral-line diagnostics, DEM inversion with **PintOfAle**, density-sensitive ratios such as **Fe XIII 202.044/203.826** and **Mg VIII 772.31/782.34**, and Mg/Ne abundance assessments in which the coronal case reproduced **85–95%** of lines within **35%** in one region and **85%** in another [2210.08899]. That study is a benchmark **of** SPICE as an EUV spectrometer, not a benchmark suite **named** SPICE Bench.

The current literature therefore supports a precise but non-singular definition: **SPICE Bench** is a reused benchmark label spanning multiple research programs, most prominently simulator-grounded EDA evaluation and automated dataset construction for software-engineering foundation models.

Source: https://www.emergentmind.com/topics/spice-bench