Papers
Topics
Authors
Recent
Search
2000 character limit reached

HapCompass: Graph-Based Haplotype Assembly

Updated 14 July 2026
  • HapCompass is a graph-based haplotype assembly method that reconstructs diploid SNP sequences by resolving conflicts among fragmented reads.
  • It employs a combinatorial optimization framework to assemble local haplotype blocks from partial and noisy NGS data.
  • The method contrasts with low-rank matrix completion approaches by focusing on discrete conflict resolution and local phasing accuracy.

Searching arXiv for papers on HapCompass and closely related haplotype assembly work. Search query: HapCompass haplotype assembly arXiv HapCompass is a haplotype-assembly method situated in the NGS-based single-individual phasing literature. In comparative descriptions of that literature, it is identified as belonging to the class of approaches that build a graph or formulate an optimization over fragment conflicts, rather than recasting haplotype assembly as a low-rank matrix-completion problem (Majidian et al., 2018). In the underlying problem setting, the aim is to reconstruct the maternal and paternal SNP sequences of a diploid individual from partially overlapping sequencing reads; because reads do not necessarily overlap continuously, the output is typically a collection of haplotype blocks rather than a single uninterrupted chromosome-scale phase.

1. Problem Setting

HapCompass belongs to the broader class of methods for assembling haplotypes directly from sequencing reads. For a diploid individual, the two haplotypes are the maternal and paternal sequences of SNP alleles. After alignment to a reference genome, non-SNP sites are removed, leaving a fragmentary representation in which each read covers only a subset of variant positions. The central inference task is to combine these local observations into two globally consistent allele sequences.

This setting is inherently combinatorial. Read fragments provide only partial and noisy information, and the phase relation between distant SNPs is often mediated through chains of overlapping reads rather than direct co-observation. As a result, any method in the HapCompass family must reconcile local consistency with genome-wide assembly constraints. The block structure of the output is not an implementation detail but a direct consequence of finite read length and incomplete overlap.

2. Methodological Identity of HapCompass

The most explicit characterization available in the comparative literature is that HapCompass is conceptually aligned with methods that use a graph or an optimization over fragment conflicts, whereas alternative formulations model the read matrix as an incomplete low-rank object and recover haplotypes by matrix completion (Majidian et al., 2018). This places HapCompass within the graph-based lineage of haplotype assembly.

That distinction is substantive. A graph-or-conflict formulation treats the observed phase relations induced by reads as discrete compatibility constraints. A plausible implication is that HapCompass organizes evidence at the level of fragment agreement and disagreement, and then seeks a globally coherent phasing by resolving contradictions in that induced structure. In this sense, HapCompass is best understood not as a statistical completion model over missing matrix entries, but as a combinatorial phasing framework whose primitive objects are conflicts among read-supported allele configurations.

3. Data Representation and Assembly Output

Although the comparative source does not restate the internal representation used by HapCompass itself, it makes clear what any competing haplotype-assembly method must consume and produce. Reads are aligned to the reference genome, reduced to SNP observations, and interpreted as partial evidence over variant sites. Because the observations are incomplete, assembly procedures operate on sparse support patterns rather than full-length haplotype vectors.

The output of such methods is ordinarily a set of haplotype blocks. This is especially clear in real-data settings where reads do not overlap continuously. For HapCompass, this implies that performance must be interpreted at two levels simultaneously: local phasing accuracy within blocks, and structural continuity across the assembled phase blocks. A common misconception is that haplotype assembly necessarily yields chromosome-wide phase; in practice, block fragmentation is a routine and expected feature of read-based assembly.

4. Accuracy and Completeness Criteria

Comparative studies in the same methodological area evaluate haplotype assembly with a mixture of accuracy and continuity metrics (Majidian et al., 2018). One standard measure is the reconstruction rate,

rr=1−1lmin⁡{HD(h^m,hm),HD(h^p,hp)},\text{rr}=1-\frac{1}{l}\min\Big\{ \mathcal{HD}(\hat{\boldsymbol h}_m,\boldsymbol h_m), \mathcal{HD}(\hat{\boldsymbol h}_p,\boldsymbol h_p) \Big\},

where HD\mathcal{HD} is Hamming distance. Another is the switch error rate (SWER), defined as the number of switches divided by haplotype length; a switch occurs when the inferred parental origin flips relative to the previous SNP. For real fosmid blocks, studies also report the SNP missing rate (SMR), mean block length, and AN50, the median block length weighted by correct allele proportion.

These criteria are important for understanding HapCompass because they expose a recurring trade-off in haplotype assembly. High local phasing accuracy does not by itself guarantee long, complete blocks, and improved continuity can come with computational or statistical costs. This suggests that HapCompass, like other graph-based baselines, is most usefully assessed in a multidimensional performance space rather than by a single scalar score.

5. Contrast with Matrix-Completion Approaches

A particularly clear foil to HapCompass is the matrix-completion paradigm developed for NGS-based haplotype assembly (Majidian et al., 2018). In that formulation, haplotypes and reads are encoded in a ±1\pm 1 alphabet, reads are arranged into an incomplete matrix R\boldsymbol R, and the latent haplotype matrix H\boldsymbol H is estimated by solving a low-rank recovery problem such as

min⁡H∑(i,j)∈Ω(Hij−Rij)2subject torank⁡(H)=2.\min_{\boldsymbol H}\sum_{(i,j)\in\Omega}(H_{ij}-R_{ij})^2 \quad \text{subject to} \quad \operatorname{rank}(\boldsymbol H)=2.

Variants of this strategy include singular value thresholding, nuclear-norm minimization, and OPTSPACE-based completion.

The contrast is methodological rather than merely notational. Matrix-completion methods assume that the hidden haplotype structure can be recovered through low-rank estimation and subsequent quantization, whereas HapCompass is described as relying on graph structure or fragment-conflict optimization. This suggests two different abstractions of the same biological problem: one continuous and algebraic, the other discrete and combinatorial. The literature presents the low-rank approach as appealing because it provides a unified framework that is not restricted to all-heterozygous sites and can accommodate both heterozygous and homozygous variants (Majidian et al., 2018). By implication, HapCompass serves as a representative of the older but still influential graph-based view against which such newer formulations are defined.

6. Position in the Literature

The continued use of HapCompass as a named comparison point indicates its status as an established reference method in haplotype assembly research (Majidian et al., 2018). Even when later work adopts a different mathematical lens, HapCompass remains part of the conceptual baseline: it marks the graph-and-conflict tradition against which matrix factorization, convex relaxation, and related methods articulate their novelty.

Its significance therefore extends beyond any single implementation detail. HapCompass represents a canonical answer to a core modeling question in haplotype assembly: whether the structure of the problem should be encoded primarily as a conflict graph or as a recoverable latent matrix. Subsequent work emphasizing more continuous haplotype blocks, lower SNP missing rate, and alternative optimization strategies does not eliminate that distinction; rather, it sharpens it. A plausible implication is that HapCompass remains relevant not only as a historical algorithm, but as a methodological category within the broader theory of read-based phasing.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HapCompass.