---
title: Size-Minimal Checkpointing
url: https://www.emergentmind.com/topics/size-minimal-checkpointing
type: topic
---

# Size-Minimal Checkpointing

Searching arXiv for recent and foundational papers relevant to size-minimal checkpointing across ML, HPC, AD, and distributed systems.
Size-minimal checkpointing denotes a family of checkpointing strategies that minimize the amount of state that must be stored, transmitted, or kept live for recovery, rollback, or recomputation, subject to correctness and performance constraints. The phrase is not uniform across the literature. In online checkpoint placement, it refers to the best worst-case rewind quality achievable with a fixed budget of \(k\) checkpoints [1302.4216]. In weakly consistent fully replicated databases, it is formalized as **Strict Concision**, meaning that the stored checkpoint has only one copy of each object in the database and does not store any channel states [2510.06404]. In large-scale machine learning, it appears as sparse, incremental, differential, input-aware, or layer-wise checkpointing, each designed to reduce checkpoint size or memory footprint without violating recovery or training requirements [2412.15411, 2010.08679, 2209.02478, 2602.22158].

## 1. Definitions and scope

The literature uses “size-minimal” to denote different constrained optimization problems rather than a single universal formalism. In mobile and ad hoc distributed systems, the objective is often **minimum-process coordinated checkpointing**: only the processes directly or transitively needed for a consistent global checkpoint should take checkpoints, because wireless bandwidth, stable storage, and battery power are limited [1005.5440, 1111.2208]. In online checkpointing, the budget is the number of maintained checkpoints, and the central quantity is the discrepancy between the longest interval induced by online placement and the ideal spacing \(T/(k+1)\) at time \(T\) [1302.4216]. In reverse-mode automatic differentiation, the problem is to reduce tape growth by using checkpoint trees that trade recomputation for lower storage growth [1708.06799].

In storage systems and databases, the notion is closer to physical concision. Application-level differential checkpointing reduces bytes written by storing only changed blocks and adds a file format that supports fragmentation of protected datasets in order to support dynamic sizes [1906.05038]. In fully replicated weakly consistent databases, size-minimality is stricter: one stored copy per logical object, no redundant per-replica state, and no channel state [2510.06404].

In model training, the budget may be GPU memory, host-memory bandwidth, network bandwidth, or remote storage capacity. Mimose formulates the problem as choosing a tensor checkpointing plan \(C(x)\) for each input size \(x\) such that \(\text{peak\_mem}(C(x),x)\le M\) while minimizing recomputation overhead [2209.02478]. Check-N-Run reduces checkpoint write bandwidth and capacity for recommendation models through incremental checkpointing and quantization [2010.08679]. MoEtion reduces the amount of MoE state captured per iteration by checkpointing only subsets of experts while always checkpointing non-expert parameters [2412.15411]. LLMTailor turns checkpointing into a layer-wise assembly problem in which layers and their optimizer state can be merged from different checkpoints to form a composite, resumable state [2602.22158].

| Setting | Size-minimal interpretation | Representative mechanism |
|---|---|---|
| Online rewinding | Fixed \(k\) checkpoints, minimum worst-case discrepancy | Online placement and replacement [1302.4216] |
| Prediction-aware fault tolerance | Minimum platform waste under prediction windows | Two-mode regular/proactive checkpointing [1302.4558] |
| Mobile or MANET systems | Minimum number of checkpointing processes | Dependency-vector-based coordinated checkpointing [1005.5440, 1111.2208] |
| HPC dynamic datasets | Minimum changed data written | Differential checkpointing with hashing and fragmentation [1906.05038] |
| GPU training memory budget | Minimum recomputation under memory cap | Input-aware tensor checkpoint planner [2209.02478] |
| Large-model training I/O budget | Minimum persisted model state per interval | Incremental, sparse, or layer-wise checkpointing [2010.08679, 2412.15411, 2602.22158] |
| Fully replicated databases | One copy per object, no channel state | Strict Concision via DTCS [2510.06404] |

## 2. Objective functions and analytical metrics

Across domains, size-minimal checkpointing is typically expressed as a trade-off between checkpoint cost and recovery cost. In the fault-prediction literature, waste is the fraction of time the platform is not doing useful work. Without prediction, regular checkpointing alone contributes \(C/T_r\), where \(C\) is checkpoint cost and \(T_r\) is the period; with prediction windows, the analysis introduces a regular mode and a proactive mode, with distinct costs \(C\) and \(C_p\), periods \(T_r\) and \(T_p\), and a total waste that also depends on downtime \(D\), recovery time \(R\), platform MTBF \(\mu\), prediction precision \(p\), recall \(r\), and window size \(I\) [1302.4558]. The resulting optimal periods are generalized Young/Daly-type formulas, including
\[
T_p^{\star} = \sqrt{\frac{((1-p)I + p) C_p}{p}}
\]
for proactive mode and a corresponding closed form for \(T_r^{\star}\) under prediction windows [1302.4558].

In application-level differential checkpointing, the key question is when hashing overhead is outweighed by reduced I/O. If \(t_w\) is time to write a block, \(t_h\) is time to hash it, and \(n_d=N_d/N_t\) is the dirty-block fraction, the normalized net benefit is
\[
\tau = (t_h - t_w) + n_d (t_w + t_h).
\]
The break-even dirty fraction is
\[
\eta = \frac{t_w - t_h}{t_w + t_h} \approx \frac{1-\rho}{1+\rho}, \qquad \rho=\frac{t_h}{t_w}.
\]
Differential checkpointing is beneficial when the dirty fraction is below \(\eta\) [1906.05038].

For memory-budgeted activation checkpointing, the optimization is closer to knapsack scheduling. Mimose models per-layer activation memory as
\[
m_i(x) \approx a_i x^2 + b_i x + c_i,
\]
defines excess memory as the predicted activation requirement above a budget \(M\), and then selects a set of layers whose dropped activations cover the excess while attempting to minimize recomputation [2209.02478]. The paper does not provide a formal optimality theorem, but its scheduler is explicitly built around the constraint \(\text{peak\_mem}(C(x),x)\le M\) and a greedy cover of excess memory.

These formulations show that “size” is rarely optimized in isolation. A plausible implication is that size-minimal checkpointing is best understood as constrained minimization of persisted or live state under a fixed correctness criterion and an explicit latency, recomputation, or availability budget.

## 3. Scheduling and decomposition algorithms

A foundational line of work studies checkpoint placement itself, independent of storage format. In online checkpointing, the active checkpoints partition \([0,T]\) into \(k+1\) intervals, and the worst rewind cost is the longest interval \(\bar{\ell}_T\). The discrepancy of an algorithm \(A\) at time \(T\) is
\[
q(A,T) := (k+1)\frac{\bar{\ell}_T}{T},
\]
with \(\Perf(A)=\sup_{T\ge t_k} q(A,T)\) and \(q^*(k)\) the infimum over all online algorithms [1302.4216]. The paper improves known upper bounds to \(q_k \le 1.59 + o(1)\) for all \(k\) and \(q_k \le \ln(4)+o(1)\le 1.39+o(1)\) when \(k\) is a power of two, and proves the lower bound \(q_k \ge 1.30 - o(1)\) [1302.4216]. Here size-minimality means that a fixed checkpoint budget is used as efficiently as possible.

Reverse-mode AD presents a different scheduling problem: classical taping has storage that can grow linearly in runtime, and divide-and-conquer checkpointing reduces this by splitting execution intervals and replaying subintervals as needed. The central innovation in arbitrary-program checkpointing is to move the mechanism into the language implementation, using interruption and resumption primitives that allow checkpoints to span any execution interval rather than only syntactic loops [1708.06799]. Binary bisection produces \(O(\log t)\) increases in both snapshot space and recomputation time, while treeverse and binomial schedules retain the fixed-space, fixed-time, or logarithmic-overhead regimes known from Griewank’s framework [1708.06799].

The connection between these literatures is structural. Both treat checkpoint placement as a scheduling problem over an execution interval, and both optimize the geometry of stored states rather than only their byte representation. This suggests that size-minimal checkpointing has a placement dimension as important as its serialization dimension.

## 4. Storage-reducing checkpoint representations

Several systems minimize checkpoint size by changing what is persisted. Differential checkpointing stores only changes relative to prior state. In the FTI-based approach for dynamic datasets, a protected dataset is split into blocks, hashes are maintained per block, dirty blocks are identified by hash comparison, and a new file format, FTI-FF, uses virtual containers whose positions remain immutable as datasets grow [1906.05038]. This avoids the failure mode of page-based dCP when dynamic-size datasets relocate or resize.

Incremental checkpointing uses update sparsity rather than block equality. In Check-N-Run, recommendation models are dominated by embedding tables, and only a fraction of embedding rows is modified in a checkpoint interval. Each GPU maintains a bit-vector with one bit per embedding vector, marking rows accessed during the interval; the bit-vectors are typically less than \(0.05\%\) of model size [2010.08679]. The system compares one-shot incremental checkpoints, consecutive incrementals, and an intermittent policy that alternates incrementals with occasional full checkpoints according to a simple cost rule. Quantization is applied only for storage, not for training computation, with asymmetric and adaptive asymmetric methods used to reduce checkpoint size while respecting restart-accuracy constraints [2010.08679].

Layer-wise selective checkpointing moves the representation boundary from data blocks or parameter rows to semantically meaningful model modules. LLMTailor restructures AdamW parameter groups so that each transformer layer is represented by two optimizer groups, one for non-decay and one for decay parameters, and the total number of groups becomes \(2L+x\), where \(L\) is the number of transformer layers and \(x\) the number of auxiliary layers [2602.22158]. A merge recipe then specifies which layer should be copied from which checkpoint, and the corresponding model weights and optimizer shards are assembled into a composite checkpoint. The paper evaluates parity checkpointing and filtered checkpointing rather than a formal “significant update” norm, so the selectivity policy is external to the merging mechanism [2602.22158].

The strongest formalization of storage concision appears in weakly consistent fully replicated databases. There the checkpoint itself must contain one copy of each logical object and no channel state, even though all replicas hold copies. MuFASA’s DTCS construction is explicitly designed to satisfy this “Strict Concision” requirement [2510.06404].

## 5. Model-training checkpointing under memory, bandwidth, and failure budgets

Large-model training has produced some of the clearest system-level demonstrations of size-minimal checkpointing. Mimose addresses GPU-memory-constrained training with dynamic input sizes. It collects per-layer memory and time online through shuttling forwarding, fits quadratic memory predictors, and uses a greedy scheduler with buckets of similarly sized layers. The estimator achieves approximately \(0.32\%\) error on TC-Bert and about \(0.32\text{–}0.46\%\) across four NLP tasks; estimator and scheduler overhead is about \(0.27\text{–}0.39\) ms per invocation, and the total Mimose overhead over an epoch is equivalent to only \(2.6\text{–}6.4\) iterations [2209.02478].

Check-N-Run targets terabyte-scale recommendation models in which embeddings account for more than \(99\%\) of model size. It combines per-row incremental checkpointing with checkpoint-only quantization and reports \(6\text{–}17\times\) reductions in required write bandwidth and \(2.5\text{–}8\times\) reductions in required capacity on real-world models [2010.08679]. The snapshot stall is measured as less than 7 seconds per checkpoint for 128 GPUs, and tracking overhead is about \(1\%\) of iteration time [2010.08679].

MoEtion is specialized for Mixture-of-Experts training, where expert parameters can dominate checkpoint size. Its selective checkpointing stores all non-expert parameters every iteration but only \(S\) of \(E\) experts per MoE layer, giving
\[
C_{\text{selective}} \approx \left(P_{\text{non\_expert}} + \frac{S}{E}P_{\text{expert}}\right)(B_{weight}+B_{optimizer}).
\]
Expert selection is popularity-aware and budget-aware, driven by local and remote bandwidth budgets \(B_{local}=BW_{host}\times T_{iter}\) and \(B_{remote}=BW_{net}\times T_{iter}\) [2412.15411]. The paper reports checkpoint size reductions up to \(9\times\), checkpointing-overhead reduction up to \(4\times\), recovery-overhead reduction up to \(31\times\), and ETTR up to \(0.98\), even with MTBF as low as 20 minutes [2412.15411].

LLMTailor addresses full-model LLM checkpoints from a different angle. Instead of every checkpoint storing the entire model and optimizer, the system permits layer-wise partial checkpoints and reconstructs a resumable state by merging layers and optimizer groups across checkpoints. The evaluation reports that filtered checkpointing can reduce total checkpoint size from 1799.52 GB to 420 GB for Llama3.1-8B and can reduce checkpoint-time share for Qwen2.5-7B from \(20.63\%\) to \(7.26\%\), which the abstract summarizes as \(4.3\) times smaller and \(2.8\) times faster, while parity checkpointing preserves final train and evaluation loss in the reported settings [2602.22158].

## 6. Distributed consistency, minimum-process protocols, and limitations

In distributed systems, minimizing checkpoint size often means minimizing the number of participants rather than the number of bytes per participant. In mobile distributed systems, a coordinated checkpoint should involve only the transitive closure of the initiator’s current dependencies. The proposed algorithm in [1005.5440] uses dependency vectors, message-sent vectors, piggybacked metadata, and a weight-based termination scheme so that exactly the necessary processes take tentative checkpoints, no blocking occurs, and no useless checkpoints are taken. Its message overhead is summarized as \(3\cdot N_{\min}\cdot C_{air}\), and the number of checkpoints is \(N_{\min}\), the minimum-process number [1005.5440].

A closely related MANET scheme on top of cluster-based routing likewise seeks a consistent set of checkpoints while ensuring that only the minimum number of nodes in the cluster are required to take checkpoints and that the protocol uses very few control messages [1111.2208]. Here the checkpointing objective is constrained by limited storage capacity, limited power, and wireless communication overhead rather than by disk throughput.

MuFASA generalizes the consistency side of the problem to weakly consistent databases by defining Distributed Transaction Consistent Snapshot (DTCS). A checkpoint \(CP\) is characterized by a virtual event \(CP_{\text{event}}\) such that committed transactions are either entirely before or entirely after that event:
\[
T \in CP_{\text{set}} \leftrightarrow (\text{\(T\) committed} \land T \rightarrow CP_{\text{event}}).
\]
The algorithm uses a three-color protocol and a single counter per replica, requires only \(O(n)\) new messages, and stores the checkpoint only at the initiator replica, thereby satisfying Strict Concision [2510.06404].

A recurring misconception is that size-minimal checkpointing always means the smallest possible serialized file. The literature shows otherwise. In online checkpointing, the budget is the number of checkpoints, not bytes [1302.4216]. In mobile and MANET checkpointing, minimality is the minimum number of processes [1005.5440, 1111.2208]. In MoE training, the paper explicitly states that the method is not mathematically minimal in an information-theoretic sense, but is close to overhead-minimal under its assumptions [2412.15411]. A second misconception is that minimizing checkpoint size automatically preserves recovery semantics. The opposite is often true: sparsity introduces staleness, differential formats require stable logical-to-physical mapping, and minimum-process protocols require precise dependency tracking. The main research trajectory therefore balances three quantities rather than one: checkpoint size, steady-state overhead, and recovery correctness.

Source: https://www.emergentmind.com/topics/size-minimal-checkpointing