Pep2Prob: Probabilistic Peptide Models
- Pep2Prob is a family of peptide-centric probabilistic models that assign calibrated probabilities to peptide events in MS-based proteomics.
- Its likelihood-based scoring pipeline uses theoretical and observed spectra with supervised maximum likelihood to achieve high calibration and correct identification rates.
- Pep2Prob integrates machine learning and large-scale benchmarks to predict proteotypic peptide detectability and fragment-ion occurrence, enhancing protein quantitation.
Searching arXiv for Pep2Prob and related proteomics papers to ground the article in current records. Pep2Prob denotes a peptide-centric probabilistic orientation in MS-based proteomics in which probabilities are assigned to peptide-spectrum matches, peptide detectability, or peptide-specific fragment-ion occurrence, and are then used for identification, calibration, or downstream protein-level analysis. In the literature summarized here, the name is attached to three closely related but distinct formulations: a likelihood-based scoring pipeline for calibrated peptide-spectrum match probabilities, a machine-learning framework for predicting proteotypic peptide detection and absolute protein abundance, and a large-scale benchmark for predicting fragment-ion probability for a precursor in tandem mass spectrometry (Li et al., 2013, He et al., 2013, Xu et al., 12 Aug 2025).
1. Conceptual scope
The unifying object across Pep2Prob formulations is a probability attached to a peptide-level event. In the likelihood-based peptide-identification setting, the target quantity is a posterior probability that a candidate peptide generated an observed spectrum. In the proteotypic-peptide framework, the target is , where denotes “proteotypic” and is a physicochemical feature vector. In the fragment-ion benchmark, the target is the peptide-specific fragment-ion probability
with fragment ion , where is the ion type, its charge, and the cleavage position along the peptide backbone (He et al., 2013, Xu et al., 12 Aug 2025).
These formulations occupy different levels of the proteomics pipeline. One level concerns peptide-spectrum matching; another concerns peptide observability across experiments; a third concerns fragmentation behavior conditional on a precursor. This suggests that Pep2Prob is best understood not as a single software artifact, but as a family of probabilistic peptide models whose outputs can support identification, quantitation, and protein inference.
2. Likelihood-based Pep2Prob for peptide-spectrum match probabilities
In the likelihood-based formulation associated with Li, Eng, and Stephens, one begins with a theoretical spectrum 0 for a candidate peptide and an observed spectrum 1. Each theoretical peak is 2 and each observed peak is 3, with distinct 4 and 5 after peak clustering and preprocessing. The central quantity is the likelihood
6
where 7 is an unobserved emission configuration assigning each theoretical peak either to a matching observed peak or to “no emission” (Li et al., 2013).
The model factorizes as
8
For an emission configuration with 9 emitting theoretical peaks,
0
Matched observed peaks are modeled by a truncated normal in m/z around the corresponding theoretical peak and a matched-peak intensity density 1, whereas noise peaks are distributed uniformly over the m/z range of width 2 with intensity density 3. The emission probability is parameterized on the logit scale as
4
with global slope 5 and spectrum-specific intercept 6 (Li et al., 2013).
Because exact evaluation is combinatorially large, scoring uses the most probable emission configuration:
7
For a candidate set 8 and a uniform prior over candidates, posterior probabilities are approximated by normalized scores,
9
Parameter estimation proceeds by supervised maximum likelihood on known spectrum-peptide pairs, alternating between configuration updates and parameter updates for 0 and the spectrum-specific intercepts 1 (Li et al., 2013).
The framework emphasizes calibration as much as ranking. Calibration curves compare nominal posterior bins with empirical correctness, while target-decoy analysis estimates
2
On the ISB standard mixture, Pep2Prob achieved 96.6% correct top-rank identifications on the full test set, versus 87.6% for the similarity index; among the “confident” subset with 3, 98.7% were correct and 93.1% of spectra were assigned. On spectra whose top-10 SEQUEST candidates included at least one true peptide, the Pep2Prob correct rate was 87.6% overall, versus 80.2% for SEQUEST XCorr and 74.4% for the similarity index, and 98.1% on the 4 subset (Li et al., 2013).
3. Proteotypic peptide detection and absolute protein quantitation
A