Size-Minimal Checkpointing
- Size-minimal checkpointing is a strategy that minimizes the stored state needed for recovery by selecting optimal checkpoint placements under performance constraints.
- It is applied across diverse domains such as ML, HPC, AD, and distributed databases, utilizing methods like differential, incremental, and layer-wise checkpointing.
- The approach balances trade-offs among checkpoint size, recomputation time, and system correctness while adapting to constraints like memory, bandwidth, and processing resources.
Searching arXiv for recent and foundational papers relevant to size-minimal checkpointing across ML, HPC, AD, and distributed systems. Size-minimal checkpointing denotes a family of checkpointing strategies that minimize the amount of state that must be stored, transmitted, or kept live for recovery, rollback, or recomputation, subject to correctness and performance constraints. The phrase is not uniform across the literature. In online checkpoint placement, it refers to the best worst-case rewind quality achievable with a fixed budget of checkpoints (Bringmann et al., 2013). In weakly consistent fully replicated databases, it is formalized as Strict Concision, meaning that the stored checkpoint has only one copy of each object in the database and does not store any channel states (Ravishankar et al., 7 Oct 2025). In large-scale machine learning, it appears as sparse, incremental, differential, input-aware, or layer-wise checkpointing, each designed to reduce checkpoint size or memory footprint without violating recovery or training requirements (Gandhi et al., 2024, Eisenman et al., 2020, Liao et al., 2022, Sun et al., 25 Feb 2026).
1. Definitions and scope
The literature uses “size-minimal” to denote different constrained optimization problems rather than a single universal formalism. In mobile and ad hoc distributed systems, the objective is often minimum-process coordinated checkpointing: only the processes directly or transitively needed for a consistent global checkpoint should take checkpoints, because wireless bandwidth, stable storage, and battery power are limited (Kumar et al., 2010, Tuli et al., 2011). In online checkpointing, the budget is the number of maintained checkpoints, and the central quantity is the discrepancy between the longest interval induced by online placement and the ideal spacing at time (Bringmann et al., 2013). In reverse-mode automatic differentiation, the problem is to reduce tape growth by using checkpoint trees that trade recomputation for lower storage growth (Siskind et al., 2017).
In storage systems and databases, the notion is closer to physical concision. Application-level differential checkpointing reduces bytes written by storing only changed blocks and adds a file format that supports fragmentation of protected datasets in order to support dynamic sizes (Keller et al., 2019). In fully replicated weakly consistent databases, size-minimality is stricter: one stored copy per logical object, no redundant per-replica state, and no channel state (Ravishankar et al., 7 Oct 2025).
In model training, the budget may be GPU memory, host-memory bandwidth, network bandwidth, or remote storage capacity. Mimose formulates the problem as choosing a tensor checkpointing plan for each input size such that while minimizing recomputation overhead (Liao et al., 2022). Check-N-Run reduces checkpoint write bandwidth and capacity for recommendation models through incremental checkpointing and quantization (Eisenman et al., 2020). MoEtion reduces the amount of MoE state captured per iteration by checkpointing only subsets of experts while always checkpointing non-expert parameters (Gandhi et al., 2024). LLMTailor turns checkpointing into a layer-wise assembly problem in which layers and their optimizer state can be merged from different checkpoints to form a composite, resumable state (Sun et al., 25 Feb 2026).
| Setting | Size-minimal interpretation | Representative mechanism |
|---|---|---|
| Online rewinding | Fixed checkpoints, minimum worst-case discrepancy | Online placement and replacement (Bringmann et al., 2013) |
| Prediction-aware fault tolerance | Minimum platform waste under prediction windows | Two-mode regular/proactive checkpointing (Aupy et al., 2013) |
| Mobile or MANET systems | Minimum number of checkpointing processes | Dependency-vector-based coordinated checkpointing (Kumar et al., 2010, Tuli et al., 2011) |
| HPC dynamic datasets | Minimum changed data written | Differential checkpointing with hashing and fragmentation (Keller et al., 2019) |
| GPU training memory budget | Minimum recomputation under memory cap | Input-aware tensor checkpoint planner (Liao et al., 2022) |
| Large-model training I/O budget | Minimum persisted model state per interval | Incremental, sparse, or layer-wise checkpointing (Eisenman et al., 2020, Gandhi et al., 2024, Sun et al., 25 Feb 2026) |
| Fully replicated databases | One copy per object, no channel state | Strict Concision via DTCS (Ravishankar et al., 7 Oct 2025) |
2. Objective functions and analytical metrics
Across domains, size-minimal checkpointing is typically expressed as a trade-off between checkpoint cost and recovery cost. In the fault-prediction literature, waste is the fraction of time the platform is not doing useful work. Without prediction, regular checkpointing alone contributes , where is checkpoint cost and is the period; with prediction windows, the analysis introduces a regular mode and a proactive mode, with distinct costs 0 and 1, periods 2 and 3, and a total waste that also depends on downtime 4, recovery time 5, platform MTBF 6, prediction precision 7, recall 8, and window size 9 (Aupy et al., 2013). The resulting optimal periods are generalized Young/Daly-type formulas, including
0
for proactive mode and a corresponding closed form for 1 under prediction windows (Aupy et al., 2013).
In application-level differential checkpointing, the key question is when hashing overhead is outweighed by reduced I/O. If 2 is time to write a block, 3 is time to hash it, and 4 is the dirty-block fraction, the normalized net benefit is
5
The break-even dirty fraction is
6
Differential checkpointing is beneficial when the dirty fraction is below 7 (Keller et al., 2019).
For memory-budgeted activation checkpointing, the optimization is closer to knapsack scheduling. Mimose models per-layer activation memory as
8
defines excess memory as the predicted activation requirement above a budget 9, and then selects a set of layers whose dropped activations cover the excess while attempting to minimize recomputation (Liao et al., 2022). The paper does not provide a formal optimality theorem, but its scheduler is explicitly built around the constraint 0 and a greedy cover of excess memory.
These formulations show that “size” is rarely optimized in isolation. A plausible implication is that size-minimal checkpointing is best understood as constrained minimization of persisted or live state under a fixed correctness criterion and an explicit latency, recomputation, or availability budget.
3. Scheduling and decomposition algorithms
A foundational line of work studies checkpoint placement itself, independent of storage format. In online checkpointing, the active checkpoints partition 1 into 2 intervals, and the worst rewind cost is the longest interval 3. The discrepancy of an algorithm 4 at time 5 is
6
with 7 and 8 the infimum over all online algorithms (Bringmann et al., 2013). The paper improves known upper bounds to 9 for all 0 and 1 when 2 is a power of two, and proves the lower bound 3 (Bringmann et al., 2013). Here size-minimality means that a fixed checkpoint budget is used as efficiently as possible.
Reverse-mode AD presents a different scheduling problem: classical taping has storage that can grow linearly in runtime, and divide-and-conquer checkpointing reduces this by splitting execution intervals and replaying subintervals as needed. The central innovation in arbitrary-program checkpointing is to move the mechanism into the language implementation, using interruption and resumption primitives that allow checkpoints to span any execution interval rather than only syntactic loops (Siskind et al., 2017). Binary bisection produces 4 increases in both snapshot space and recomputation time, while treeverse and binomial schedules retain the fixed-space, fixed-time, or logarithmic-overhead regimes known from Griewank’s framework (Siskind et al., 2017).
The connection between these literatures is structural. Both treat checkpoint placement as a scheduling problem over an execution interval, and both optimize the geometry of stored states rather than only their byte representation. This suggests that size-minimal checkpointing has a placement dimension as important as its serialization dimension.
4. Storage-reducing checkpoint representations
Several systems minimize checkpoint size by changing what is persisted. Differential checkpointing stores only changes relative to prior state. In the FTI-based approach for dynamic datasets, a protected dataset is split into blocks, hashes are maintained per block, dirty blocks are identified by hash comparison, and a new file format, FTI-FF, uses virtual containers whose positions remain immutable as datasets grow (Keller et al., 2019). This avoids the failure mode of page-based dCP when dynamic-size datasets relocate or resize.
Incremental checkpointing uses update sparsity rather than block equality. In Check-N-Run, recommendation models are dominated by embedding tables, and only a fraction of embedding rows is modified in a checkpoint interval. Each GPU maintains a bit-vector with one bit per embedding vector, marking rows accessed during the interval; the bit-vectors are typically less than 5 of model size (Eisenman et al., 2020). The system compares one-shot incremental checkpoints, consecutive incrementals, and an intermittent policy that alternates incrementals with occasional full checkpoints according to a simple cost rule. Quantization is applied only for storage, not for training computation, with asymmetric and adaptive asymmetric methods used to reduce checkpoint size while respecting restart-accuracy constraints (Eisenman et al., 2020).
Layer-wise selective checkpointing moves the representation boundary from data blocks or parameter rows to semantically meaningful model modules. LLMTailor restructures AdamW parameter groups so that each transformer layer is represented by two optimizer groups, one for non-decay and one for decay parameters, and the total number of groups becomes 6, where 7 is the number of transformer layers and 8 the number of auxiliary layers (Sun et al., 25 Feb 2026). A merge recipe then specifies which layer should be copied from which checkpoint, and the corresponding model weights and optimizer shards are assembled into a composite checkpoint. The paper evaluates parity checkpointing and filtered checkpointing rather than a formal “significant update” norm, so the selectivity policy is external to the merging mechanism (Sun et al., 25 Feb 2026).
The strongest formalization of storage concision appears in weakly consistent fully replicated databases. There the checkpoint itself must contain one copy of each logical object and no channel state, even though all replicas hold copies. MuFASA’s DTCS construction is explicitly designed to satisfy this “Strict Concision” requirement (Ravishankar et al., 7 Oct 2025).
5. Model-training checkpointing under memory, bandwidth, and failure budgets
Large-model training has produced some of the clearest system-level demonstrations of size-minimal checkpointing. Mimose addresses GPU-memory-constrained training with dynamic input sizes. It collects per-layer memory and time online through shuttling forwarding, fits quadratic memory predictors, and uses a greedy scheduler with buckets of similarly sized layers. The estimator achieves approximately 9 error on TC-Bert and about 0 across four NLP tasks; estimator and scheduler overhead is about 1 ms per invocation, and the total Mimose overhead over an epoch is equivalent to only 2 iterations (Liao et al., 2022).
Check-N-Run targets terabyte-scale recommendation models in which embeddings account for more than 3 of model size. It combines per-row incremental checkpointing with checkpoint-only quantization and reports 4 reductions in required write bandwidth and 5 reductions in required capacity on real-world models (Eisenman et al., 2020). The snapshot stall is measured as less than 7 seconds per checkpoint for 128 GPUs, and tracking overhead is about 6 of iteration time (Eisenman et al., 2020).
MoEtion is specialized for Mixture-of-Experts training, where expert parameters can dominate checkpoint size. Its selective checkpointing stores all non-expert parameters every iteration but only 7 of 8 experts per MoE layer, giving
9
Expert selection is popularity-aware and budget-aware, driven by local and remote bandwidth budgets 0 and 1 (Gandhi et al., 2024). The paper reports checkpoint size reductions up to 2, checkpointing-overhead reduction up to 3, recovery-overhead reduction up to 4, and ETTR up to 5, even with MTBF as low as 20 minutes (Gandhi et al., 2024).
LLMTailor addresses full-model LLM checkpoints from a different angle. Instead of every checkpoint storing the entire model and optimizer, the system permits layer-wise partial checkpoints and reconstructs a resumable state by merging layers and optimizer groups across checkpoints. The evaluation reports that filtered checkpointing can reduce total checkpoint size from 1799.52 GB to 420 GB for Llama3.1-8B and can reduce checkpoint-time share for Qwen2.5-7B from 6 to 7, which the abstract summarizes as 8 times smaller and 9 times faster, while parity checkpointing preserves final train and evaluation loss in the reported settings (Sun et al., 25 Feb 2026).
6. Distributed consistency, minimum-process protocols, and limitations
In distributed systems, minimizing checkpoint size often means minimizing the number of participants rather than the number of bytes per participant. In mobile distributed systems, a coordinated checkpoint should involve only the transitive closure of the initiator’s current dependencies. The proposed algorithm in (Kumar et al., 2010) uses dependency vectors, message-sent vectors, piggybacked metadata, and a weight-based termination scheme so that exactly the necessary processes take tentative checkpoints, no blocking occurs, and no useless checkpoints are taken. Its message overhead is summarized as 0, and the number of checkpoints is 1, the minimum-process number (Kumar et al., 2010).
A closely related MANET scheme on top of cluster-based routing likewise seeks a consistent set of checkpoints while ensuring that only the minimum number of nodes in the cluster are required to take checkpoints and that the protocol uses very few control messages (Tuli et al., 2011). Here the checkpointing objective is constrained by limited storage capacity, limited power, and wireless communication overhead rather than by disk throughput.
MuFASA generalizes the consistency side of the problem to weakly consistent databases by defining Distributed Transaction Consistent Snapshot (DTCS). A checkpoint 2 is characterized by a virtual event 3 such that committed transactions are either entirely before or entirely after that event: 4 The algorithm uses a three-color protocol and a single counter per replica, requires only 5 new messages, and stores the checkpoint only at the initiator replica, thereby satisfying Strict Concision (Ravishankar et al., 7 Oct 2025).
A recurring misconception is that size-minimal checkpointing always means the smallest possible serialized file. The literature shows otherwise. In online checkpointing, the budget is the number of checkpoints, not bytes (Bringmann et al., 2013). In mobile and MANET checkpointing, minimality is the minimum number of processes (Kumar et al., 2010, Tuli et al., 2011). In MoE training, the paper explicitly states that the method is not mathematically minimal in an information-theoretic sense, but is close to overhead-minimal under its assumptions (Gandhi et al., 2024). A second misconception is that minimizing checkpoint size automatically preserves recovery semantics. The opposite is often true: sparsity introduces staleness, differential formats require stable logical-to-physical mapping, and minimum-process protocols require precise dependency tracking. The main research trajectory therefore balances three quantities rather than one: checkpoint size, steady-state overhead, and recovery correctness.