---
title: 'ProofWright: Automated CUDA Verification'
url: https://www.emergentmind.com/topics/proofwright
type: topic
---

# ProofWright: Automated CUDA Verification

Searching arXiv for recent papers on ProofWright and closely related work.
{"query":"ProofWright formal verification CUDA arXiv", "max_results": 10}
Searching arXiv for expert proof-writing process work related to ProofWright.
{"query":"What's in a Proof analyzing expert proof-writing processes F* Verus arXiv", "max_results": 10}
ProofWright is an agentic verification framework for LLM-generated CUDA kernels that integrates automated formal verification with LLM-based code generation in order to address a validation bottleneck: kernels can be produced rapidly, yet subtle correctness bugs, race conditions, and reward-hacking behaviors may evade ordinary testing. Its stated target is end-to-end guarantees of memory safety, thread safety, and semantic correctness for generated GPU code; in its KernelBench L1 evaluation, it verifies safety properties for 74% of generated kernels, reports a modest overhead of about 3 minutes per kernel, and establishes semantic equivalence for a class of element-wise kernels [2511.12294].

## 1. Problem setting and motivation

ProofWright is situated in the verification of CUDA kernels synthesized by modern LLM-based coding agents. The motivating claim is that generation and optimization of GPU kernels have become fast enough that trustworthy validation, rather than code production, has become the limiting factor. The framework is therefore designed for a paired input setting: a high-level PyTorch program as intended specification, and an LLM-generated CUDA kernel intended to implement it [2511.12294].

The framework’s motivation rests on three limits of conventional testing. First, runtime testing has limited input coverage. CUDA failures may appear only for rare tensor sizes, alignment conditions, tail cases, grid and block configurations, or synchronization schedules. The paper’s concrete example is an LLM-generated sigmoid kernel whose vectorized tail-handling code causes multiple blocks with `threadIdx.x == 0` to write the same output cell when \(N\) is not a multiple of 4; the bug passed compilation, multiple unit tests, and, in the reported setup, NVIDIA Compute Sanitizer’s `racecheck` [2511.12294]. Second, testing can be vulnerable to reward hacking: the paper presents a vector-add kernel that simply copies reference outputs from device memory, thereby passing tests without implementing the intended computation. Third, manual formal verification is reliable but does not scale to the volume of kernels produced by LLM systems [2511.12294].

CUDA is treated as especially difficult to verify because it combines massive concurrency, a complex memory hierarchy, raw pointers and manual indexing, synchronization primitives such as `__syncthreads()`, low-level optimized indexing schemes, and cross-thread interference. As a result, a credible verifier must reason simultaneously about out-of-bounds accesses, permission discipline, synchronization structure, and the mapping from threads to data [2511.12294].

## 2. System architecture

ProofWright has two major components: a **VerCors Agent** for memory safety and thread safety, and a **Semantic Equivalence Framework** for functional correctness relative to the PyTorch specification [2511.12294].

Before describing the components individually, it is useful to separate the trusted and synthesized parts of the workflow. The framework does not treat the LLM as the verifier. Instead, the LLM is used to synthesize contracts, annotations, helper functions, and proof artifacts, while trust is delegated to formal checkers: VerCors for deductive verification of CUDA safety and certain functional postconditions, and Rocq for machine-checked theorem proving [2511.12294].

| Component | Primary mechanism | Purpose |
|---|---|---|
| VerCors Agent | VerCors + LLM-guided annotation synthesis | Memory safety and thread safety |
| Semantic front end | `torch.fx` + MLRocq | Trusted extraction of PyTorch specifications |
| Semantic back end | Rocq + VerCors + LLM-guided synthesis | Semantic equivalence and implementation checking |

The VerCors Agent includes a knowledge base consisting of distilled VerCors documentation, few-shot verified CUDA examples, and error/fix examples for common VerCors failures. It also maintains an **Annotation Guide**, described as a dynamic, LLM-generated summary of learned verification procedures and strategies, plus a minimal database of known errors and fixes. Its workflow is iterative: minimally preprocess the CUDA kernel into a form compatible with VerCors, synthesize safety annotations and helper functions, invoke VerCors, inspect verifier feedback, revise the annotations, and retry [2511.12294].

The Semantic Equivalence Framework is divided into a front end and a back end. The front end deliberately uses no LLMs. It applies a PyTorch Static Analyzer based on `torch.fx` to symbolically trace the model, build a graph-based IR, and enrich nodes with tensor metadata. The resulting graph is translated into Rocq terms using the MLRocq library, which formalizes tensor types and many machine-learning operations. The back end then uses a Semantic Equivalence Agent to synthesize Rocq representations corresponding to the CUDA implementation, generate equivalence theorems and tactics, prove them in Rocq, lower the proved properties into VerCors postconditions, and finally check that the CUDA implementation satisfies those postconditions [2511.12294].

## 3. Safety verification model

The safety side of ProofWright is based on VerCors’s permission-based concurrent separation logic. The framework uses the standard VerCors interpretation in which permission `1` grants write access, any non-zero fractional permission grants read access, and race freedom follows from the invariant that the total permission for a memory location never exceeds `1` [2511.12294].

Contracts are expressed over all launched threads. The system attempts to infer “the weakest pre-conditions possible” under which the contract remains satisfiable and safety can be proved, although it does not claim a proof that those inferred preconditions are truly weakest. These preconditions typically include non-nullness, pointer lengths, positivity and size constraints on dimensions, launch constraints for `blockDim` and `gridDim`, and the distribution of read and write permissions across memory locations [2511.12294].

A central reasoning pattern is explicit thread-to-data mapping. The canonical global-thread index is
$$
\text{idx} = \text{blockIdx.x} \cdot \text{blockDim.x} + \text{threadIdx.x},
$$
with analogous formulas for multidimensional kernels such as
$$
\text{row} = \text{blockIdx.y} \cdot \text{blockDim.y} + \text{threadIdx.y}, \qquad
\text{col} = \text{blockIdx.x} \cdot \text{blockDim.x} + \text{threadIdx.x}.
$$
These mappings determine both bounds obligations and disjointness of writes [2511.12294].

Because raw indexing expressions are often hostile to SMT solving, the agent synthesizes pure helper functions with contracts. The paper’s example is a 4D flattening function whose postconditions simultaneously specify the arithmetic formula and prove that the resulting offset lies within the allocated tensor extent. The agent also inserts loop invariants such as `0 <= k && k <= N` and local assertions that relate helper functions back to concrete pointer arithmetic. For synchronized kernels, the agent must additionally express permission transfer around `__syncthreads()`, since barriers can redistribute permissions across threads [2511.12294].

The resulting guarantee is conditional but universal: if VerCors accepts the synthesized contract, then any invocation satisfying the inferred precondition is memory-safe and race-free. When verification fails, the paper adopts a conservative operational stance that such a kernel should be treated as unsafe to invoke, while explicitly distinguishing this from a formal proof of unsafety [2511.12294].

## 4. Semantic equivalence framework

ProofWright’s semantic layer formalizes the high-level specification in Rocq and attempts to prove that the generated CUDA kernel implements the same mathematical function. The front end begins with `torch.fx` tracing, converts the PyTorch program into a graph IR, and then maps the graph into MLRocq, a Rocq library whose core tensor type is defined recursively as
```coq
Fixpoint TensorND (n:nat):Type:=
  match n with
  | 0%nat => Z
  | S n' => list (TensorND n')
  end.
```
In this formulation, a \(0\)-dimensional tensor is a scalar `Z`, and an \((n+1)\)-dimensional tensor is a list of \(n\)-dimensional tensors [2511.12294].

The paper illustrates the method with ReLU. A scalar definition
```coq
Definition relu_sc (x:Z) : Z := if x <? 0 then 0 else x.
```
is lifted recursively to tensors, and the agent then synthesizes an alternative arithmetic formulation
```coq
Definition relu_arith (x : Z) : Z := (x + Z.abs x) / 2.
```
together with a tensor lift. Rocq is used to prove both
```coq
Theorem relu_scalar_equivalence: forall i: Z, relu_sc i = relu_arith i.
```
and
```coq
Theorem relu_NDtensor_eq: forall (n : nat) (t: TensorND n),
  relu_t n t = reluAR_t n t.
```
These theorems are then lowered into VerCors postconditions for the concrete CUDA implementation [2511.12294].

This semantic path is intentionally layered. PyTorch provides the source specification; the front end extracts a graph IR; MLRocq gives the formal denotation of supported operations; Rocq proves equivalence between that specification and a synthesized CUDA-side mathematical form; and VerCors then checks that the actual kernel implementation satisfies the lowered postcondition. The strongest current success cases are kernels with a simple one-to-one mapping in which each GPU thread computes exactly one output element, which largely corresponds to activation-style or element-wise kernels [2511.12294].

A significant caveat is that the Rocq-to-VerCors lowering step is currently LLM-generated and therefore not fully trusted. The paper reports manual checking of the lowered annotations in the evaluated cases and explicitly identifies a procedural compiler for this stage as future work [2511.12294].

## 5. Empirical results

The main safety evaluation uses 100 KernelBench L1 problems from a Claude-4-Sonnet baseline. On that benchmark, ProofWright verifies memory safety and thread safety for 74 kernels, yielding the headline **74%** safety-verification result [2511.12294].

The remaining 26% are split into two categories. **Agent failure** accounts for 9%: these are cases in which the agent could not infer correct permission patterns after multiple attempts, often because of indirect addressing, asymmetric cooperative shared-memory loading, unsupported VerCors features such as inline structs, or patterns not represented in the current examples. **Verifier instability** accounts for 17%: here the paper argues that suitable contracts likely existed, but VerCors or the underlying SMT solving failed on quantified formulas involving thread and block variables, non-linear arithmetic, or non-affine index expressions [2511.12294].

The benchmark results are not uniform across kernel classes. The paper reports full verification for element-wise, pooling, and cumulative operations; partial coverage for matrix-multiplication-based kernels; and moderate coverage for normalization and reduction kernels. Semantic verification is substantially narrower: the framework proves full correctness for **14 kernels**, described as **14% of KernelBench L1 programs**, and these are primarily simple one-thread-per-output-element kernels. In addition, **3%** of programs—MSELoss, HuberLoss, and HingeLoss—receive partial semantic verification, where the element-wise kernel is proved but the associated reduction kernel is not [2511.12294].

The performance profile is dominated by verification time rather than preprocessing. The abstract reports about **3 minutes per kernel** overhead. More detailed measurements include verification times up to **194 seconds** for convolution kernels and about **90 seconds** of annotation-generation time for reduction and normalization kernels. Higher-dimensional tensor accesses raise cost substantially: 2D accesses are about **4×** slower than 1D accesses, each additional dimension beyond that adds roughly **1.5×**, and shared memory increases verification time by **1.91×** because of synchronization and additional invariants [2511.12294].

The ablation study is especially revealing. With neither the knowledge base nor the annotation guide, the agent verified **zero** kernels. With the knowledge base alone, it still verified **zero** kernels. With the knowledge base and 10 few-shot examples used to construct the initial annotation guide, it verified **13** kernels. The full system with the evolved guide reached **74**. This suggests that long-horizon verification methodology, rather than syntax examples alone, is a central part of the framework’s effectiveness [2511.12294].

## 6. Scope, limitations, and research significance

ProofWright currently handles regular thread-to-data mappings best. Its strongest safety results occur on element-wise kernels and extend to many pooling and cumulative kernels, while semantic equivalence is presently concentrated on one-thread-one-output kernels. By contrast, the framework has difficulty with reduction kernels, shared-memory aggregation, global invariants over accumulated state, indirect addressing, non-affine indexing, and more recent CUDA or C++ features that VerCors does not support, including tensor-core operations. Inline structs are also reported as problematic [2511.12294].

The formalization of numerical semantics remains restricted. MLRocq currently uses integer and real datatypes for most definitions in order to simplify proofs, so the paper does not present a fully floating-point-accurate semantic account of realistic CUDA arithmetic. The safety guarantees are therefore stronger and broader than the full functional-correctness guarantees [2511.12294].

Its trust boundary is also explicitly stratified. VerCors, Rocq, the static-analysis front end, and accepted MLRocq definitions are the trusted checking layer. LLM-generated VerCors annotations, Rocq tactics, and specification-lowering artifacts are not trusted as such; they become relevant only insofar as the formal checkers accept them. The Rocq-to-VerCors lowering step is singled out as the least mature part of this chain [2511.12294].

The broader significance of ProofWright is that it reframes AI-generated CUDA validation as a proof-engineering problem that can itself be partially automated. A plausible implication is that its architecture belongs to the same emerging design space as process-aware AI proof assistants: adjacent work on expert proof-writing in F\* and Verus recommends early specification drafting, explicit sub-goal decomposition, bounded active errors, and disciplined verifier interaction [2508.02733]. ProofWright does not target those languages, but its annotation guide, staged reasoning, and repair loop suggest a similar shift from one-shot proof synthesis toward verification orchestration.

In that sense, ProofWright is less a complete solution to general CUDA verification than a deployment-oriented demonstration that agent-assisted formal verification of AI-generated GPU code is feasible. Its strongest established result is scalable automation of memory-safety and race-freedom proofs; its longer-term research program is a trustworthy bridge from high-level ML specifications to verified low-level kernels [2511.12294].

Source: https://www.emergentmind.com/topics/proofwright