CausalMesh: Causal Cache for Serverless Computing
- CausalMesh is a distributed cache system for stateful serverless computing that ensures causal+ consistency through a dual-cache design and dependency integration.
- It achieves coordination-free, abort-free read/write operations with asynchronous propagation, resulting in lower latency and higher throughput compared to existing solutions.
- The protocol is formally verified with Dafny, guaranteeing strict causal cuts across caches and robust performance even when clients roam among different servers.
CausalMesh is a cache system for stateful serverless computing that provides causally consistent caching when a computation may migrate from one machine to another. It is designed for workflows composed of multiple serverless functions that access shared state on a remote database through per-node cache servers, and it targets the anomaly-prone setting in which successive functions in the same workflow may execute on different machines with different caches. The system is presented as the first cache system that supports coordination-free and abort-free read/write operations and read transactions when clients roam among multiple servers; it also offers a transactional variant, CausalMesh-TCC, for read-write transactional causal consistency in the presence of client roaming, but at the cost of abort-freedom. Its protocol is formally verified in Dafny, and its evaluation reports lower latency and higher throughput than existing proposals (Zhang et al., 21 Aug 2025).
1. Problem setting and motivation
Stateful serverless workflows are built from multiple serverless functions that access shared data on a remote database. In practice, developers often place a cache layer between the serverless runtime and the database to reduce I/O latency. The motivating problem is that, in a serverless environment, functions in the same workflow may be scheduled to different nodes with different caches, so weakly consistent caches can expose non-intuitive anomalies. The examples given include a workflow that does not see its own writes, recommendation behavior that does not reflect a user’s latest preferences, and access-control changes that take effect out-of-order (Zhang et al., 21 Aug 2025).
The paper situates this problem against the cost of direct database access and against the limitations of existing cache designs. It cites remote database access such as DynamoDB at 10–20ms, and it notes that eventually consistent caches such as AWS DAX can exhibit anomalies under function migration. A specific result reported for DAX is a 14.2% anomaly rate in a simple two-function workflow when using 8 nodes. Existing remedies, including HydroCache and FaaSTCC, are described as providing caches with transactional causal consistency, but only by introducing expensive coordination or aborts, thereby weakening the primary performance rationale for caching (Zhang et al., 21 Aug 2025).
Within this setting, CausalMesh’s goal is to provide causal+ consistency, denoted CC+, across caches so that data dependencies within a workflow are respected regardless of migration between nodes. The central design constraint is to preserve low-latency, high-throughput cache operation without coordination in the critical path and without aborts on ordinary read/write operations (Zhang et al., 21 Aug 2025).
2. System architecture and protocol mechanisms
CausalMesh is deployed between the serverless platform and the backend database. The architecture comprises cache servers on worker nodes, a client library used by workflows and functions, and RPC communication among cache servers. The intended deployment is ideally 1:1 between cache servers and serverless function workers for locality, though other configurations are possible. The client library exposes key-value read and write operations together with a multi-key read transaction API. Cache servers exchange updates using FIFO-ordered RPC that may be delayed, and the protocol is explicitly decentralized rather than centrally coordinated (Zhang et al., 21 Aug 2025).
A defining mechanism is the “dual cache.” Each cache server maintains a C-cache and an I-cache:
1 2 |
C-cache := { Key → (Value, VC) }
I-cache := { Key → [ (Value, VC, deps) ] } |
The C-cache contains versions known to be visible on all servers and is intended to represent a causally consistent cut across the system. The I-cache is a multi-version store for items that are not yet globally visible, including local writes that have not completed propagation. Versioning is handled with vector clocks, and dependencies are recorded as mappings of the form {Key → VC}. The dependency metadata is restricted to “nearest” dependencies in order to minimize metadata size (Zhang et al., 21 Aug 2025).
The key local operation is dependency integration. When a cache server serves a read, or when a write becomes globally visible, it integrates the dependencies needed to maintain what the paper calls a “strict causal cut.” The operation is local and does not require network communication. The protocol description gives the following pseudo-code fragment:
1 2 3 4 5 6 |
def integrate(self, deps): all_deps = {k: v for k, v in transitive_predecessors_of(deps)} for k, vcs in all_deps.items(): consistent_versions = self.Inconsistent[k].remove(filter(vc in vcs)) self.vc.merge_all(vcs) self.Consistent[k].merge_all(consistent_versions) |
The propagation path for writes uses a double-round chain inspired by chain replication. A cache server receiving a write assigns a vector clock, places the write in the I-cache, and returns to the client immediately. The write is then sent asynchronously to the server’s successor, continuing for a full round. On the second round, once the write reaches the tail, it is integrated into the C-cache at the tail and optionally at other servers for faster visibility. Reads consult the local C-cache, but before returning a result the server integrates any dependencies carried by the client so that the returned state covers those dependencies (Zhang et al., 21 Aug 2025).
3. Consistency model, operational semantics, and transactional variant
The semantics of CausalMesh are organized around causal+ consistency. In the base system, reads and writes are coordination-free and abort-free. Reads are served from the local C-cache after dependency integration, while writes update the local I-cache and propagate asynchronously, so propagation is not on the critical path. CausalMesh also supports multi-key read transactions within a single function, with the caveat that the transaction must not overlap with the client’s own previous writes to those keys; otherwise it may abort (Zhang et al., 21 Aug 2025).
The paper formalizes the C-cache invariant using the notion of a strict causal cut:
This definition captures the requirement that all dependencies of every write in the C-cache are also in the C-cache, or have been overwritten by newer causally-after writes. In operational terms, the client library tracks dependencies and local buffered state; the read path merges returned values with local state when needed, and the write path merges local dependencies before assigning a version and initiating propagation (Zhang et al., 21 Aug 2025).
CausalMesh-TCC extends the design to transactional causal consistency for read-write workflows. The trade-off is explicit: full read-write transactional semantics require the loss of abort-freedom. In this mode, writes are buffered on the client until the end of the workflow and sent in a batch for atomic integration into the C-cache. Each read in a transaction must observe a consistent system-wide cut in which either all or none of a transaction’s writes are visible. If the desired cut does not exist, such as under parallel or concurrent operations, the transaction or workflow is aborted and retried (Zhang et al., 21 Aug 2025).
A recurrent misconception is to equate causal consistency with immediate global visibility or with universally abort-free transactions. The reported design is narrower and more precise. In CausalMesh, ordinary reads and writes are abort-free and coordination-free, but visibility to other caches occurs only after two rounds of propagation, and CausalMesh-TCC preserves stronger transactional semantics by allowing aborts. The paper states that new writes become visible to other caches after 2N RTTs, where is the number of cache servers (Zhang et al., 21 Aug 2025).
4. Formal verification and proof obligations
A prominent feature of CausalMesh is its formal verification in Dafny. The paper reports that an earlier version of the protocol had passed TLA+ model checking for small configurations, yet a subtle causal consistency violation was missed. That experience motivated a full deductive proof rather than bounded model exploration (Zhang et al., 21 Aug 2025).
The verification methodology proceeds by encoding the entire protocol, including server state, client state, and message state, in Dafny with Z3 as the SMT backend. Server state includes the C-cache and I-cache; client state includes dependencies and local buffered values. Node-level actions are expressed as two-state predicates. The proof then identifies invariants, proves base cases, and proves inductive preservation under protocol actions. The stated scope is arbitrary system size rather than a finite bounded configuration (Zhang et al., 21 Aug 2025).
The paper identifies several key proven properties. First, C-caches are always strict causal cuts. Second, all client and server dependencies referenced in a read are globally available. Third, clients always observe monotonic reads and writes, specifically read your writes, monotonic read, writes follow reads, and monotonic writes. The stated conclusion is that CausalMesh guarantees causal+ consistency for all executions and all workloads (Zhang et al., 21 Aug 2025).
The formal-verification aspect is significant because the protocol combines client roaming, asynchronous propagation, multi-version state, and dependency integration. These are precisely the sorts of interacting mechanisms for which testing and small-state exploration can fail to expose corner-case violations. The paper’s emphasis is therefore not merely on proof-assisted implementation, but on machine-checked preservation of semantic invariants under the full protocol transition system (Zhang et al., 21 Aug 2025).
5. Experimental evaluation and reported performance
The experimental evaluation uses CloudLab m510 nodes with 3–16 cache servers, Nightcore as the serverless runtime, and Redis as the underlying database with 5ms added RTT. The evaluation covers a microbenchmark and a real application. The microbenchmark is a three-function workflow with reads and writes on random keys under a Zipfian distribution. The real application is the Movie Review Service from DeathStarBench, described as a realistic workload with 13 serverless functions (Zhang et al., 21 Aug 2025).
On the microbenchmark, the reported results for CausalMesh are 1.3–2.2× higher throughput than HydroCache or FaaSTCC, 7–59% lower median latency, and up to 97% lower tail (P99) latency. The paper further reports that latency remains consistently low as the number of servers increases. For CausalMesh-TCC, latency matches transactional causal-consistency baselines while throughput is reported as up to 1.6× higher (Zhang et al., 21 Aug 2025).
Resource-efficiency results are also given. CPU usage is reported as 23–57% lower for CausalMesh than for HydroCache or FaaSTCC, and memory or metadata footprint is 17–20% lower. The memory result is accompanied by the statement that footprint remains bounded over time, unlike HydroCache, which can grow indefinitely without garbage collection. Scalability results report that CausalMesh scales linearly with the number of cache servers and that normalized throughput remains constant as nodes increase, whereas HydroCache and FaaSTCC do not scale as well; HydroCache is stated to be unable to handle more than 8 nodes because of coordination costs (Zhang et al., 21 Aug 2025).
For the Movie Review Service workload, the paper reports that CausalMesh achieves 2× higher throughput and up to 64% reduction in tail latency compared to competitors before their performance degrades. CausalMesh-TCC is reported as providing similar or up to 1.35× higher throughput than other TCC systems. The write-visibility evaluation states that the inconsistency window remains small—for example, approximately 3ms for 8 nodes—whereas HydroCache and FaaSTCC expose visibility only after periodic refresh intervals of 50–100ms (Zhang et al., 21 Aug 2025).
6. Terminological scope and relation to other causal-literature usages
The term CausalMesh in this context denotes a distributed-systems protocol for causally consistent caching, not a framework for statistical causal inference, geometric morphometry, or categorical semantics. Here, “causal” refers to causal consistency in the distributed-systems sense, implemented through vector clocks, dependency tracking, and happens-before relations. This is distinct from work on causal diagrams for physical models, which defines causality through temporal derivatives in systems of differential equations (Kinsler, 2015).
The distinction matters because adjacent arXiv literatures use the vocabulary of meshes and causal models in different ways. In medical imaging, “Deep Structural Causal Shape Models” study counterfactual reasoning over 3D surface meshes using geometric deep learning and deep structural causal models, enabling subject-specific counterfactual mesh generation rather than cache coherence (Rasal et al., 2022). In categorical work, the summary of “Topos Causal Models” states that “CausalMesh and similar frameworks build on the universal, compositional, and functorial approach championed in TCMs,” and UCLA describes a layered architecture in which simplicial, graph, data, and homotopy layers yield mesh-like causal representations (Mahadevan, 5 Aug 2025, Mahadevan, 2022).
These neighboring usages suggest a broader semantic field around “causal” and “mesh,” but they do not redefine the serverless system. Within the distributed-systems literature, CausalMesh is specifically a formally verified causal cache for stateful serverless computing. Its principal contributions are the dual-cache design, local dependency integration, double-round propagation, support for client roaming, a coordination-free and abort-free CC+ path for reads and writes, and a separate TCC mode when stronger transactional semantics are required (Zhang et al., 21 Aug 2025).