---
title: Document-Packing Strategies
url: https://www.emergentmind.com/topics/document-packing-strategies
type: topic
---

# Document-Packing Strategies

Document-packing strategies are a set of algorithmic and workflow techniques with the goal of arranging, segmenting, or combining texts—usually documents or variable-length token sequences—into training or inference batches that maximize the utilization of hardware, preserve contextual integrity, or optimize learning objectives. These strategies have become essential for large-scale language model pre-training and fine-tuning, document layout analysis, and applications requiring efficient processing of variable-size or multi-modal inputs. Document-packing interacts deeply with core issues such as context coherence, compute throughput, dataset diversity, supervised fine-tuning efficiency, and downstream reasoning or composition ability.

## 1. Fundamental Principles and Formalisms

Document-packing is grounded in bin-packing and sequence-alignment problems:

- **Atom size ($a$) and maximum sequence length ($L$, MSL):** At the core, document-packing optimizes allocation of data atoms—token sequences of variable length but not exceeding $L$—into bins or training windows of fixed length $L$ [2408.09621]. If $a<L$, atoms must be merged or padded; if $a>L$, atoms are split across bins.
- **Packing objective:** Minimize the number of bins (for maximal throughput), maximize bin fill rate, preserve boundaries (for context), and minimize compute/memory waste or truncation.
- **Key metrics:** Perplexity for language modeling, packing ratio $\rho=\frac{\sum_{i=1}^B \ell_i}{B L_\text{max}}$, total tokens processed per GPU-hour, and empty area or density for 2D layout [2408.09621, 1210.4502, 2410.08081].
- **Packing integrity:** Preservation of contextual coherence demands that atom or document boundaries respect semantic units; misaligned packing can cause artificial context fragmentation or corruption [2408.09621, 2404.10830].

Formalizations typically specify decision variables $x_{i,j}\in\{0,1\}$ for assigning piece $i$ to bin $j$ under capacity and non-overlap constraints (1D/2D), with objectives such as:
\[
\min_{x,y} \sum_j y_j \quad \text{s.t.} \quad \sum_{i} \ell(c_i)x_{i,j}\leq L y_j, \quad \sum_{j} x_{i,j}=1
\]
[2404.10830], or their geometric and 2D variants [2410.12628, 1210.4502].

## 2. Packing Algorithms: Heuristics and Optimizations

### 2.1 Heuristic Bin Packing

- **Best-Fit-Decreasing (BFD):** Documents/chunks are sorted in descending size, packed into bins whose remaining space is minimized but sufficient (classic 1D bin packing) [2404.10830].
- **First-Fit-Decreasing (FFD):** Chunks are sorted and greedily placed into the first available bin with capacity, fast but can sometimes create more waste [2505.22018].
- **Greedy Packing:** For SFT, sequences are sorted descendingly and allocated to bins maximizing fill below $L_\text{max}$, preserving conversation or document boundaries whenever possible [2410.08081]. Complexity typically $O(N\log L)$ or $O(N^2)$; segment-tree acceleration is used in large-scale implementations [2404.10830].

### 2.2 Concatenation and Padding

- **Concatenation (“concat”):** Documents are streamed with boundary tokens and cut into bins, often resulting in context seams but perfect fill [2408.09621].
- **Padding:** Atoms (documents or sequences) are ended and right-padded to exactly $L$, preserving one document per chunk but at the cost of extra padding tokens and more steps [2408.09621].

### 2.3 Fine-Grained and Asymmetric Packing

- **SlimPack (slice-level and asymmetric):** Decomposes input into small “slices,” balancing forward and backward computational loads via MILP-based partitioning, attuned to asymmetric cost profiles (backward attention ~2.5$\times$ forward) [2509.26246]. The pipeline consists of DP-balance, MicroPack MILP, and critical path simulation.

### 2.4 Packing for Document Layout

- **Mesh-candidate BestFit (2D):** Used for image-like document layout, maintains a dynamic set of empty rectangular meshes and greedily fills with elements maximizing local area utilization, subject to containment, non-overlap, and implicit aesthetic regularity [2410.12628]. This approach favors “well-aligned” and dense layouts.

### 2.5 Packing for Retrieval/Sliding-Window Attention

- **Window-level packing:** For transformers with local attention, documents are cut into overlapping windows, batching only windows with real tokens, which substantially reduces padding overhead for variable-length documents [2005.04908]. In document ranking, this enables near 50% reduction in wasted computation.

## 3. Empirical Performance, Trade-Offs, and Metrics

Empirical benchmarking has quantified the trade-offs associated with each method:

| Method         | Perplexity (↓) | Throughput (tokens/GPU-h ↑) | Contextual Integrity | Padding Overhead | Truncation/Fragmentation | Downstream Task Gains |
|----------------|-----------------|-------------------------------|---------------------|-------------------|-------------------------|----------------------|
| Concat (a=L)   | Higher          | Highest                       | Mixed contexts      | None              | High                    | Baseline             |
| Padding (a=L)  | Lowest          | Lower (~15–45% slower)        | Full document       | Some              | None                    | + PPL, + downstream  |
| BFD/FFD BinPack| Lower           | Near Concat                   | Document intact     | ≤ 0.01% extra     | Minimal                 | +4–20% on tasks      |
| SlimPack       | N/A             | 1.15–2.8× over baselines      | Sample/slice-order  | N/A               | Flexible                | Up to 2.8× speedup   |

* Source: [2408.09621], [2404.10830], [2509.26246], [2410.08081].

- Setting atom size $a=L$ (MSL) yields minimal perplexity and best trade-off, aligning context window to atom and eliminating spurious concatenation [2408.09621].
- Padding always achieves lower perplexity at the expense of steps/$E$, while concatenation favors speed.
- In Best-fit Packing, unnecessary truncations are reduced, and sequence utilization compared with concat is essentially identical (≤+0.003% in large-scale runs) [2404.10830].
- Empirical downstream gains from optimal packing are significant: +4.7% (reading comprehension), +16.8% (context following), +9.2% (program synthesis), and hallucination reductions of up to 58.3% [2404.10830].
- SlimPack achieves up to $2.8\times$ throughput improvement by balancing slice assignments, crucial for extreme long-context or heavy-tailed input size distributions [2509.26246].
- Packing ratio $\rho$ and speedup $S$ are best supported by greedy or slice-level approaches for large SFT; typical wall-clock savings of 60–85% are reported [2410.08081].

## 4. Domain-Specific and Advanced Packing Strategies

### 4.1 Continual Pre-training

- **Seamless Packing:** Combines a sliding-window overlap mechanism for long documents ($r_\text{max} \approx 0.3$) with FFD-packing of short remainders; achieves up to 2 pp downstream task gains and reduces context discontinuity from 40% (concat) to <5% [2505.22018].

### 4.2 Multi-hop Reasoning

- **Packings for Cross-document Reasoning:** Enables latent multi-hop capability by assembling sequences containing 4–6 documents, always with cross-document attention. Packing beyond this “sweet spot” degrades precision and increases hallucination [2512.14427]. Epoch-wise repacking is necessary to avoid overfitting to static document groupings.

### 4.3 Supervised Fine-Tuning (SFT)

- **Random vs. Greedy Packing in SFT:** Greedy packing preserves multi-turn context integrity, yielding up to 4.5 pt improvement on GPT-evaluated metrics for large LLaMA-3-70B models on 1M+ datasets [2410.08081]. Gains are muted for small models or datasets. The effective batch size, batch size × learning rate scaling, and ratio of multi-turn to single-turn examples critically mediate SFT efficiency and representation learning.

### 4.4 Document Layout Analysis

- **Mesh-based Bin Packing for Layout Synthesis:** DocLayout-YOLO uses mesh-candidate best-fit to maximize local fill rate and maintain global alignment and density, achieving best-in-class alignment (0.0009) and density (0.645) in synthetic document generation [2410.12628].

## 5. Practical Implementation and Engineering Considerations

- **Implementation Pipelines:**
  - Pre-tokenize documents, append boundary (EOS/SEP), chunk to targets, and pack with chosen strategy.
  - Apply randomized shuffling or epoch-wise repacking to maximize data coverage and avoid memorized context groupings [2408.09621, 2512.14427].
  - Cross-document attention masking is essential for logical document independence in packing; disabling it converts packed batches to equivalently distinct samples [2404.10830].

- **Example Pseudocode (Padding, $a=L$) [2408.09621]:**
  ```python
  class PaddingDataset(Dataset):
      def __init__(self, tokenized_docs, L):
          self.chunks = []
          for doc in tokenized_docs:
              for i in range(0, len(doc), L-1):
                  slice_ = doc[i:i+L-1] + [EOS_ID]
                  slice_ += [EOS_ID] * (L - len(slice_))
                  self.chunks.append(torch.tensor(slice_))
          random.shuffle(self.chunks)
      def __len__(self): return len(self.chunks)
      def __getitem__(self, idx): 
          x = self.chunks[idx]
          return x[:-1], x[1:]
  ```

- **Hyperparameter Checklist:**
  - Atom size $a=L$ for coherence and randomness
  - Batch size tuned to maximize hardware use
  - Position encoding that enables efficient variable context lengths (ALiBi/rotary)
  - Rigorous monitoring of packing ratio $\rho$ and sequence fill rates

- **Limitations:** Excessively large $L$ relative to document size increases padding waste; fine-grained methods may incur MILP solver costs (SlimPack), and SFT requires bespoke hyperparameter tuning for effective LR/batch size scaling.

## 6. Implications, Best Practices, and Future Directions

- **Best Practices:**
  - For autoregressive LM training, set atom size $a$ equal to MSL; prefer padding for maximal accuracy, concat for throughput, and packing for balanced trade-off [2408.09621].
  - For SFT, greedy packing is preferred for dialog/multi-turn; random packing is adequate for single-turn [2410.08081].
  - In multi-document tasks, pack 4–6 documents per sequence and enable cross-document attention [2512.14427].
  - For layout synthesis, mesh-candidate best-fit yields high alignment and density [2410.12628].

- **Outlook:** Scaling and hybridization of fine-grained techniques (e.g., SlimPack, mesh-candidate methods), tighter integration of pipeline simulation and hardware-aware scheduling, and adaptive tuning (via auto-tuned solvers) are active research directions. For tasks requiring inter-document relations, careful balancing of pack size and attention scope is critical to avoid hallucination and maximize emergent reasoning ability.

- **Key conceptual finding:** Packing methods must balance resource utilization, representational coherence, and downstream generalization. There is no one-size-fits-all solution—the workload, data distribution, and objective all dictate the optimal packing strategy.

## 7. References

- "Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models" [2408.09621]
- "Fewer Truncations Improve Language Modeling" [2404.10830]
- "SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training" [2509.26246]
- "Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning" [2410.08081]
- "Improving Continual Pre-training Through Seamless Data Packing" [2505.22018]
- "Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models" [2512.14427]
- "DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception" [2410.12628]
- "Comparing several heuristics for a packing problem" [1210.4502]
- "Local Self-Attention over Long Text for Efficient Document Retrieval" [2005.04908]

Source: https://www.emergentmind.com/topics/document-packing-strategies