PMRobust: Robust Persistent Memory Compiler
- PMRobust is a compiler-based system that enforces strict persistency by automatically inserting cache-line flushes and fences to eliminate missing flush and fence bugs.
- It uses whole-program, flow-sensitive analysis with CFL-reachability and field sensitivity to track persistent memory states and escape conditions for accurate optimization.
- Empirical evaluations show that its optimized variants, especially PMRobust_flit, achieve near-original performance while maintaining robust crash consistency across PM and CXL systems.
PMRobust is a compiler-based system for persistent-memory programs that automatically inserts cache-line flush and fence instructions so that code using persistent memory is free from missing flush and fence bugs, while keeping performance overhead extremely low. It is designed for both traditional persistent memory, such as Intel Optane, and CXL shared memory, treating them uniformly as “PM.” The system is formulated around robustness to weak persistency models: every execution under the weak model should correspond to some strict-persistency execution, thereby eliminating missing flush and fence bugs by construction (Guo et al., 23 Sep 2025).
1. Persistent-memory semantics and the target bug class
In PM and CXL shared memory systems, main memory contents outlive crashes, but CPU caches remain volatile, and writes to PM go through caches and store buffers. Under Intel’s Px86 architectural persistency semantics, stores may reach other cores before they reach persistent memory, cache lines may be written back to PM in arbitrary order, and cache contents may be lost on crashes. Correct crash-consistent data structures therefore require cache-line flushes such as clwb, clflushopt, or clflush, together with fences such as sfence, to force durability and ordering (Guo et al., 23 Sep 2025).
The central bug class addressed by PMRobust is the missing flush or missing fence bug. These bugs induce lost durability, crash-consistency violations, and robustness violations. The paper characterizes the last notion using the terminology of PSan: execution under a weak persistency model is not equivalent to any execution under strict persistency, so post-crash states can exhibit reorderings or omissions of stores that would never arise in a strictly persistent execution. A canonical example is the case where post-crash execution observes y == 1 while x == 1 was not persisted, although strict persistency would require the persistence order to match the volatile order (Guo et al., 23 Sep 2025).
The correctness criterion is therefore not merely “eventual durability,” but robustness to weak persistency models. Strict persistency is defined as the condition that persistency memory order is identical to volatile memory order. PMRobust does not attempt to implement strict persistency literally after every store, which would be costly; rather, it enforces enough flush and fence structure that every weak-persistency execution is equivalent to some strict-persistency execution. This distinction is essential because the ordering visible to threads under volatile consistency is decoupled from the order in which data reach PM (Guo et al., 23 Sep 2025).
A common misconception is that the problem is only about persistent media such as NVDIMMs. The same reasoning extends to CXL shared memory: compute nodes cache updates to shared memory, and node failures can lose dirty cache lines, so software must use flush and fence operations to make writes durable or at least crash-consistent. PMRobust therefore treats PM and CXL uniformly under the same persistency discipline (Guo et al., 23 Sep 2025).
2. Hardware model, ISA assumptions, and robustness formulation
PMRobust is built on Px86, Intel’s formal persistency semantics for x86. The underlying model includes per-core store buffers, a conceptual global persistent buffer, cache-line granularity persistence, and crashes that lose volatile state while leaving PM contents as of some prefix of persistent-buffer writes. For CXL, the model includes multiple processing nodes sharing memory over a CXL network, with each node caching shared memory; optional Global Persistent Flush can handle certain failures, but many failure modes remain, including cable or NIC failures (Guo et al., 23 Sep 2025).
At the instruction level, PMRobust assumes the standard x86 persistency operations. clflush flushes and evicts a cache line immediately; clflushopt flushes and evicts lazily and requires a fence for completion; clwb writes back without guaranteed eviction and also requires a fence; sfence orders and ensures completion of prior flush operations. PMRobust chooses clwb as its default flush primitive because it has weaker ordering properties than clflush, often retains cache lines, and pairs naturally with sfence to finalize persistence (Guo et al., 23 Sep 2025).
The semantic reasoning is expressed with happens-before and sequenced-before relations. Strict persistency requires that persist order respect happens-before: if , then must be persisted before . Robustness to weak persistency models relaxes this only when reordering would be unobservable, for example for objects that are not reachable in post-crash executions. This observation leads directly to PMRobust’s analysis of whether an object is reachable from persistent roots and whether its writes have progressed from dirty to flushed to durable (Guo et al., 23 Sep 2025).
This suggests that PMRobust is best understood as a static enforcement mechanism for observability-constrained persistence ordering. It is not merely a flush-placement heuristic. Its formal target is the equivalence class of executions admitted by strict persistency, and its transformation is constructed so that the compiled program remains within that class under Px86 (Guo et al., 23 Sep 2025).
3. Whole-program static analysis
PMRobust is a whole-program, flow-sensitive, intraprocedural plus interprocedural dataflow analysis that identifies which pointers refer to PM, tracks whether PM locations have escaped, tracks a persistency state for each PM location, detects robustness violations and durability bugs, and guides insertion of flushes and fences (Guo et al., 23 Sep 2025).
To distinguish PM from volatile memory, PMRobust adapts the CFL-reachability alias analysis of Zheng and Rugina. Developers list PM allocators such as pmalloc or libpmemobj allocators; pointer values returned by these allocators are treated as initial PM roots; aliases of PM pointers are computed via CFL-reachability; and the analysis is demand-driven so that only pointers reachable from PM roots are analyzed. PMRobust also adds type-based field sensitivity: for each struct type, it tracks which fields are PM pointers, and all objects of the same type are treated uniformly (Guo et al., 23 Sep 2025).
The analysis separates captured objects from escaped objects. Captured objects are not reachable from any persistent root, so stores to them are invisible after a crash and flushes can be delayed. Escaped objects are reachable from a persistent root, so their persistency ordering may be observable after a crash. This distinction is encoded by an escape lattice
and a map
The analysis is conservative: if a location may have escaped along any path, it is treated as escaped (Guo et al., 23 Sep 2025).
Persistency is tracked by a second finite-state abstraction. The persistency lattice is
with ordering
A PM location is modeled as , where is an object reference and is a field or byte offset. At each program point, PMRobust maintains
0
A store marks a location dirty; clwb or clflushopt transition dirty to clwb; sfence transitions clwb to clean; and clflush directly yields clean. Atomic loads are treated conservatively by marking the loaded location dirty, because another thread may have written a value but not flushed it yet (Guo et al., 23 Sep 2025).
The dataflow update for both escape and persistency maps is the standard transfer
1
Using these state abstractions, PMRobust checks two main categories of violations. First, at function exit, if any PM location that is not a parameter or return value is both escaped and non-clean, a violation is reported. In the top-level main, this captures durability bugs. Second, inside a function, if some escaped PM location is non-clean and the program stores to another escaped PM location, then a robustness violation is reported, because two distinct escaped, non-clean cache lines may enable a post-crash state inconsistent with strict persistency. The analysis also treats a release to a non-PM location as a potential violation if escaped non-clean PM state remains outstanding, ensuring that critical-section updates are persisted before unlock or release (Guo et al., 23 Sep 2025).
The key theorem states that if a program has a robustness violation, PMRobust’s analysis will detect it. The supporting argument proceeds by enumerating cases where two stores to distinct escaped PM locations have a happens-before relation and showing that the finite-state abstraction must discover either two escaped, non-clean locations or a release with outstanding non-clean state (Guo et al., 23 Sep 2025).
4. Automatic insertion of flushes and fences
PMRobust’s repair strategy is staged. First, when the analysis finds that a PM location becomes dirty in a way that contributes to a robustness violation, it inserts clwb immediately after the store. For variable-size writes such as memcpy over PM ranges, it inserts a loop or helper that issues clwb over the entire range. For atomic loads, PMRobust uses the FliT scheme to avoid always flushing. Second, after flushes are inserted, PMRobust re-runs the analysis and inserts sfence either at function exits where some PM location remains in clwb state or just before a second object would become escaped and non-clean (Guo et al., 23 Sep 2025).
This transformation is governed by three correctness results. The detection-completeness theorem states that if a program has a robustness violation, it will be detected by PMRobust. The transformation-soundness theorem states that after PMRobust transforms a program by inserting flushes and fences, PMRobust will not report any robustness violations in the program. The strict-persistency equivalence theorem states that if PMRobust does not report a robustness violation, then any execution of the program is one under strict persistency. Taken together, these establish the claimed absence of missing flush and fence bugs under the system’s assumptions (Guo et al., 23 Sep 2025).
A central optimization concerns newly allocated objects. At allocation time, a PM object is captured, not yet reachable from any persistent root. Stores to its fields are therefore not immediately observable after a crash, so PMRobust can delay flushes until the object escapes. The paper’s persistent-stack example illustrates the pattern. A new node n is allocated by pmalloc, fields n->data and n->next are written while n is still captured, and only then is a commit store such as s->top = n executed. PMRobust does not flush after each field write; instead, it inserts flushes for the node just before the commit store, followed by sfence, and then flushes the root update. This ensures that when the node becomes reachable, its fields are already clean, and at no program point do two distinct escaped, non-clean cache lines exist (Guo et al., 23 Sep 2025).
The system explicitly avoids over-serialization. Naively enforcing strict persistency would require a flush after every PM store and a fence between every store to distinct escaped locations. PMRobust reduces this cost through the captured-versus-escaped distinction, flow-sensitive per-location state tracking, and FliT optimization for atomic loads. For cases that cannot be analyzed precisely, such as pointer arithmetic on PM addresses, PMRobust conservatively inserts fences to ensure correctness (Guo et al., 23 Sep 2025).
5. Implementation and empirical evaluation
PMRobust is implemented as a transformation pass in LLVM, building on LLVM 8 and extended as whole-program analysis via WLLVM. It targets C and C++ and was tested with clang and clang++. The evaluated PM libraries include PMDK, libpmem, libvmmalloc, and custom allocators in the benchmark suite. Developers compile code with the PMRobust pass and provide a small configuration identifying allocation functions that return PM pointers, together with optional annotations for regions where strict persistency can be safely relaxed (Guo et al., 23 Sep 2025).
The evaluation covers RECIPE concurrent persistent indexes, a PM-enabled Memcached, and a PMDK data store. The RECIPE set includes P-ART, P-BwTree, P-CLHT, P-Masstree, FAST&FAIR, and CCEH; the workloads use YCSB A, B, C, and E with integer and string keys, and persistent memory is provided via PMDK’s libvmmalloc. The Memcached evaluation uses a PM-enabled version with libpmem and a memaslap workload of 100K operations, 90% get and 10% set, 16-byte keys, 1024-byte values, and 16 threads. The PMDK data store evaluation uses seven map backends—btree, rbtree, ctree, hashmap_atomic, hashmap_tx, hashmap_rp, and skiplist—with a benchmark that inserts 100K random 8-byte keys and then removes them. The system used for evaluation is Ubuntu 22.04.4 on a 16-core 2.4 GHz Intel Xeon Silver 4314 with 256 GB DRAM and 256 GB Intel Optane PM (Guo et al., 23 Sep 2025).
Three PMRobust variants are reported.
| Variant | Description | Geometric mean overhead |
|---|---|---|
| PMRobust_base | Flush + fence immediately after every store and atomic load to PM | 11.21% |
| PMRobust_opt | Adds escape + persistency state analysis, delaying flushes and fences | 6.41% |
| PMRobust_flit | Adds FliT’s counter-based optimization for atomic loads | 0.26% |
Across all benchmarks, the geometric mean overhead over original, hand-optimized programs is 11.21% for PMRobust_base, 6.41% for PMRobust_opt, and 0.26% for PMRobust_flit (Guo et al., 23 Sep 2025). Many benchmarks, including P-CLHT, P-Masstree, FAST&FAIR, CCEH, and Memcached, show PMRobust_base already close to original performance, suggesting that extra flushes and fences are often off critical paths. PMRobust_opt further reduces the gap where overhead is noticeable, especially for structures with complex escape patterns. PMRobust_flit substantially improves atomic-load-heavy workloads such as P-ART and P-BwTree, often matching or exceeding original performance. In Memcached, PMRobust_flit outperforms the original by about 20%, because PMRobust replaces flush function calls with inlined clwb instructions (Guo et al., 23 Sep 2025).
The paper also compares analysis runtime with PMBugAssist on PMDK tests and other benchmarks. PMRobust’s analysis times scale roughly with code size, from tens of seconds to a few minutes for large code bases such as Memcached with PMDK libraries. The reported table includes 199 seconds for PMRobust on Memcached, versus 1861 seconds for PMBugAssist (Guo et al., 23 Sep 2025).
6. Related systems, scope, and limitations
PMRobust is positioned against dynamic checkers, repair systems, and other static PM analyses. Dynamic tools such as PSan, XFDetector, PMTest, pmemcheck, Agamotto, and Yashme focus on bug finding through instrumentation, checkpoints, symbolic execution, or trace analysis; they require bug-revealing test cases, heavy annotations, or workload generation, and they do not guarantee the absence of bugs. Repair tools such as Hippocrates and PMBugAssist repair given bug traces but do not ensure that the repaired program is globally free of PM bugs. StaticPersist and AutoPersist determine which objects should reside in PM and preserve reachability invariants, while Corundum statically enforces persistent memory safety; these are complementary rather than substitutes. Within this landscape, PMRobust is described as the first static tool that detects missing flush and fence bugs using robustness as the correctness criterion, automatically inserts flushes and fences, and guarantees absence of missing flush and fence bugs (Guo et al., 23 Sep 2025).
This positioning clarifies another misconception: PMRobust is not a dynamic verifier that happens to insert repairs. Its methodology is purely static, whole-program, and trace-independent. A plausible implication is that PMRobust is especially attractive where test coverage is weak or crash schedules are hard to enumerate, although its guarantees remain contingent on the assumptions built into the analysis (Guo et al., 23 Sep 2025).
Those assumptions are explicit. PMRobust assumes data race freedom for non-atomic accesses, consistent with C and C++ semantics. It relies on x86 TSO plus Px86 semantics for atomics and flush and fence instructions. The current implementation requires whole-program IR after linking via WLLVM. Function pointers and complex polymorphic dispatch are not fully supported; calls through function pointers may trigger conservative flushing of all escaped, non-clean objects. Pointer arithmetic on PM addresses is treated conservatively, with fences inserted after such stores and warnings reported. The system also exposes “escape hatch” annotations for patterns that intentionally allow robustness violations but remain logically safe, such as checksums, certain link-and-persist designs, or pointer tagging; misuse of these annotations weakens guarantees (Guo et al., 23 Sep 2025).
Within that scope, PMRobust’s contribution is twofold. First, it formalizes persistent-memory correctness through robustness to weak persistency models rather than ad hoc flush placement. Second, it shows that a compiler can enforce that criterion with very low measured overhead—0.26% geometric mean for the strongest optimized variant relative to hand-placed flush and fence operations—while targeting both persistent memory and CXL shared memory under a unified abstraction (Guo et al., 23 Sep 2025).