Papers
Topics
Authors
Recent
Search
2000 character limit reached

Multi-Deletion Correction Codes

Updated 14 December 2025
  • Multi-deletion correction codes are error-correcting schemes that recover sequences with multiple arbitrary deletions while addressing challenging synchronization errors.
  • They employ algebraic models, VT codes, multiplicity-free constructions, and permutation-based methods to achieve near-optimal redundancy and efficient decoding.
  • Recent advances include explicit constructions, refined combinatorial bounds, and quantum analogs with applications in genomics, network synchronization, and file alignment.

A multi-deletion correction code is an error-correcting code designed to uniquely recover codewords from sequences that have undergone multiple arbitrary deletions. Whereas classical codes address substitution errors, multi-deletion codes confront one of the most challenging types of synchronization errors where positions and values of deleted symbols are unknown. This problem is fundamental in information theory, with connections to genomic data, network synchronization, and file alignment protocols. Multi-deletion correction codes span binary, non-binary, and permutation-based constructions, and recent progress involves nearly optimal explicit constructions, tight combinatorial bounds, advanced decoding algorithms, and quantum analogs. This article surveys the main families of codes, their underlying frameworks, decoding methods, theoretical bounds, and practical advances.

1. Foundational Principles and Main Models

Multi-deletion correction considers words xx over an alphabet Σ\Sigma where up to tt symbols can be deleted, producing an (unknown) subsequence yy. The classical objective is to construct codes C⊂ΣnC \subset \Sigma^n such that for any x≠x′∈Cx \ne x' \in C, no sequence yy can be a subsequence of both after up to tt deletions. By Levenshtein’s equivalence, codes that correct tt deletions can also correct any mixture of up to tt insertions and deletions.

Algebraic Models

  • Single Deletion (VT Codes): Varshamov–Tenengolts–Levenshtein codes use weighted sum congruences; redundancy Σ\Sigma0 is optimal for Σ\Sigma1.
  • Multiple Deletion (General Codes): For Σ\Sigma2, the problem is far more complex. Existentially, random greedy selection shows that redundancy Σ\Sigma3 suffices, but until recently, explicit codes did not match this bound.

Non-Binary, Multiplicity-Free Codes

Recent non-binary constructions focus on codes over alphabets Σ\Sigma4 with Σ\Sigma5, especially multiplicity-free words—each symbol appears exactly once. The design involves splitting each codeword into its unordered set and its induced permutation, enabling modular correction of deletions on sets and permutations (Schaller et al., 23 Jan 2025).

Permutation Codes

In permutation codes, the codewords are permutations Σ\Sigma6. The Ulam metric, defined by Σ\Sigma7, governs deletion-correctability. Codes correcting Σ\Sigma8 deletions must have minimum Ulam distance Σ\Sigma9 (Wang et al., 2024).

2. Explicit Constructions and Decoding Algorithms

Multiplicity-Free Set-Permutation Construction

For tt0, every multiplicity-free word is mapped bijectively to:

  1. Induced Set tt1;
  2. Induced Permutation tt2.

Codes are built by combining a constant-weight set code tt3 (corrects up to tt4 asymmetric deletions as 1→0 errors) and a permutation-deletion code tt5 (corrects tt6 stable deletions). The code tt7 is defined as the inverse image of tt8 under the decomposition bijection (Schaller et al., 23 Jan 2025).

Decoding (Pseudocode)

  1. Recover tt9 via indicator decoding over yy0.
  2. Sort yy1 lexicographically and trace observed positions to symbol indices to derive the shortened permutation.
  3. Decode the permutation using a stable deletion-correcting decoder for yy2.
  4. Reassemble using inverse bijection.

The complexity is yy3 for set decoding, polynomial in yy4 for permutation decoding via known subroutines.

Permutation Deletion Codes via Hamming Mapping

An injective mapping yy5 converts permutation errors into Hamming errors. For yy6, augment to yy7 and apply yy8 such that Hamming errors correspond to up to yy9 translocations for C⊂ΣnC \subset \Sigma^n0 deletions. By intersecting this image with Hamming-metric codes of distance C⊂ΣnC \subset \Sigma^n1, one gets codes of size C⊂ΣnC \subset \Sigma^n2 that correct C⊂ΣnC \subset \Sigma^n3 deletions. Decoding leverages erasure and error-correcting codes in the Hamming metric (Wang et al., 2024).

Burst Deletion and Generalizations

Burst-deletion codes correct runs of consecutive deletions. Constructions interleave array-based codes and shifted VT (SVT) codes, achieving C⊂ΣnC \subset \Sigma^n4 redundancy for correcting bursts of length C⊂ΣnC \subset \Sigma^n5—provably near optimal (Schoeny et al., 2016). Extensions address non-consecutive bursts and mixed insertions/deletions.

3. Combinatorial and Asymptotic Bounds

Sphere-Packing and Hypergraph Methods

  • General Codes: The maximum code size for correcting C⊂ΣnC \subset \Sigma^n6 deletions is bounded by fractional matching in deletion-derived hypergraphs (Kulkarni et al., 2012).
  • Levenshtein Bound and Improvements: For general C⊂ΣnC \subset \Sigma^n7-ary alphabets, classical Levenshtein bounds are C⊂ΣnC \subset \Sigma^n8. A refinement via mixed packing (allowing insertions and deletions) yields strictly better bounds when C⊂ΣnC \subset \Sigma^n9 (Cullina et al., 2013):

x≠x′∈Cx \ne x' \in C0

Singleton Bound and Rate Analysis

A code correcting x≠x′∈Cx \ne x' \in C1 deletions must have x≠x′∈Cx \ne x' \in C2. The multiplicity-free set-permutation construction achieves redundancy x≠x′∈Cx \ne x' \in C3, which is asymptotically optimal as x≠x′∈Cx \ne x' \in C4 (Schaller et al., 23 Jan 2025).

Existential and Explicit Constructions in Binary Case

  • Greedy existential codes achieve x≠x′∈Cx \ne x' \in C5 code size for binary codes correcting x≠x′∈Cx \ne x' \in C6 deletions.
  • Explicit constructions, particularly the Guruswami–HÃ¥stad augmented VT codes, match redundancy x≠x′∈Cx \ne x' \in C7 for x≠x′∈Cx \ne x' \in C8 (Guruswami et al., 2020, Sun et al., 2024).

4. Quantum Multi-Deletion Correction

Quantum analogs for multi-deletion codes address deletion of x≠x′∈Cx \ne x' \in C9 qubits (tracing out unknown positions). Two systematic methods are established:

Reed–Solomon-Based Quantum Deletion Codes

  • Alternating Sandwich Mapping: Interleave RS code blocks with marker qubits (blocks of yy0 and yy1), transforming deletion correction into erasure correction.
  • Error Locator Algorithm: Precise block alignment and marker measurement allows detection of erased blocks; standard quantum RS erasure decoding completes recovery.
  • Achieves rates arbitrarily close to the RS code’s rate for any fixed yy2, does not require prior knowledge of deletion count (Hagiwara, 2023).

Marker-Periodicity and Stabilizer Conversion

Any yy3-erasure-correcting quantum code can be lifted, via periodic marker-state prefixing, to a yy4-deletion-correcting code over alphabet of size yy5. The achievable rate scales by yy6 compared to the base erasure code (Matsumoto et al., 2021).

5. Array-Based and Structured Multi-Deletion Codes

Criss-Cross Codes for Arrays

For yy7 arrays, correcting up to yy8 deletions spread arbitrarily across rows and columns (criss-cross model) requires intersection of systematic multi-deletion codes and rank-metric Gabidulin codes. Redundancy is lower bounded by yy9, and explicit codes attain tt0 (Welter et al., 2021).

Helberg-Type and Non-Binary Constructions

Number-theoretic codes generalize Helberg’s binary construction to non-binary alphabets: using moments based on an exponentially growing weight sequence and a congruence condition modulo a carefully chosen modulus, up to tt1 deletions are correctable. Decoding algorithms work in tt2 for fixed tt3; however, redundancy scales linearly with tt4 (Le et al., 2015, Segrest et al., 26 Aug 2025).

6. List-Decoding, Edit Channels, and Limitations

Recent advances include list-decodable multi-deletion codes (outputting short lists containing the true codeword), and codes for general edit channels (insertions, deletions, substitutions). For example, a code correcting two edits achieves redundancy tt5 by intersecting two-deletion, two-substitution, and one-deletion-one-substitution constraints, surpassing previous constructions (Sun et al., 2024).

7. Comparative Summary and Open Problems

Construction/Bound Redundancy Alphabet/Structure Complexity Reference
VT code (tt6 deletion) tt7 Binary/Arbitrary Linear (Kulkarni et al., 2012)
Guruswami–Håstad explicit (tt8) tt9 Binary Poly(n) (Guruswami et al., 2020)
Multiplicity-free perm+set code tt0 Non-binary (tt1) tt2 (Schaller et al., 23 Jan 2025)
Permutation-t-deletion codes tt3 Permutations Polynomial (Wang et al., 2024)
Guess & Check (zero-error, high-prob) tt4 Binary Polynomial (Hanna et al., 2017)
Quantum Reed-Solomon construction Flexible, tt5 Qudits Efficient, no prior tt6 (Hagiwara, 2023)

Despite major advances, capacity-achieving explicit codes for arbitrary tt7 remain an open challenge in both binary and non-binary settings. For small tt8, best-known explicit codes match existential bounds up to tt9 terms; for large weight or array codes, redundancy is still linear in tt0. Permutation codes currently define the best explicit tradeoffs, and quantum deletion correction is an active area leveraging erasure conversion. Future directions include reducing the constant factors in redundancy, developing faster encoding algorithms, and extending combinatorial bounds to complex constrained sources.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Multi-Deletion Correction Codes.