Papers
Topics
Authors
Recent
Search
2000 character limit reached

MMPart: Multi-Parton Interactions in Proton Collisions

Updated 12 July 2026
  • MMPart is a framework for multi-parton interactions in proton collisions, linking QCD formulations, phenomenological models, and simulation approaches.
  • It covers both soft/semi-hard underlying-event activities and hard double parton scattering, highlighting observable signatures and effective overlap measures.
  • Recent advances include event-by-event inference using machine-learning regression to better classify MPI activity and probe proton structure.

Searching arXiv for the cited MPI/MMPart papers to ground the article. "MMPart" (Editor's term) denotes the multi-parton-interaction framework in hadron collisions, with particular emphasis on proton–proton events in which more than one parton–parton scattering occurs within a single collision. In the literature represented here, the subject spans soft and semi-hard underlying-event activity, hard double parton scattering (DPS), QCD formulations in terms of double parton distributions and generalized two-parton distributions, Monte Carlo implementations in generators such as PYTHIA and HERWIG, and recent event-by-event inference strategies based on machine-learning regression (Bartalini et al., 2011). In this usage, MMPart is not a separate formalism from multi-parton interactions; rather, it is a compact label for the linked theoretical, phenomenological, and analysis program through which MPI are identified, modeled, and exploited as probes of proton structure and of collective-like phenomena in small systems (Ortiz et al., 2020).

1. Physical content and phenomenological scope

Multi-Parton Interactions (MPI) arise because hadrons are composite objects, so a single pppp collision can contain two or more parton–parton scatterings. The literature distinguishes Single Parton Scattering (SPS) from MPI, and singles out Double Parton Scattering (DPS) as the especially important hard-MPI limit in which two separate hard subprocesses occur simultaneously (Bartalini et al., 2011). At LHC energies, MPI become more important because, at fixed final-state mass, higher collision energy probes smaller parton momentum fractions xx, where parton densities rise rapidly; consequently, the probability for several independent hard scatterings grows (Bartalini et al., 2011).

The subject naturally divides into two regimes. In the soft / semi-hard regime, MPI dominate much of the underlying event (UE) and are central to minimum-bias phenomenology, where perturbative QCD alone is insufficient and phenomenological modeling is required. In the hard regime, perturbative methods can be applied more directly, especially for DPS-sensitive final states such as multijets, W+W+jets, and electroweak-plus-jet channels (Bartalini et al., 2011). This separation is practical rather than absolute: the same event can contain a hard primary interaction, additional semi-hard scatterings, beam-remnant activity, and hadronization effects.

In contemporary small-system studies, MPI are also used as a proxy for the “underlying activity” of a pppp collision. That role became especially prominent because high-multiplicity pppp and pp–Pb collisions at the LHC exhibit long-range angular correlations, strange-hadron enhancement, and mass ordering in spectra, features that resemble phenomena seen in larger systems. Several microscopic mechanisms have been discussed in that context, including hydrodynamics, initial-state correlations, transport, color reconnection, and MPI-driven pictures (Ortiz et al., 2020). This suggests that MPI classification is not only a background-modelling problem but also a method for isolating event classes sensitive to collective-like dynamics.

2. QCD formulation and correlation structure

A central theoretical point is that MPI should be treated as a QCD correlation problem rather than only as a geometric ansatz or generator tune. In the standard DPS formulation, the cross section is written as (Bartalini et al., 2011)

σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),

where Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2) is the double parton distribution, x1,x2x_1,x_2 are momentum fractions, bb is the transverse separation between the two partons, xx0, xx1 are the subprocess scales, and xx2 if xx3, xx4 otherwise (Bartalini et al., 2011). Under the factorized transverse ansatz

xx5

this reduces to a form involving the effective overlap area

xx6

and, with the additional approximation xx7, to the well-known pocket formula

xx8

The literature treats this form as useful but approximate, because it may obscure nontrivial correlations in flavor, spin, color, longitudinal momentum, and transverse geometry (Bartalini et al., 2011).

A more systematic QCD treatment uses the generalized double parton distribution xx9GPD, written as (Dokshitzer, 2012)

W+W+0

where W+W+1 is a transverse momentum mismatch conjugate to the relative transverse separation of the two active partons. In this framework, the double hard-scattering cross section is expressed as

W+W+2

with

W+W+3

Here W+W+4 is not a dynamical cross section in the ordinary sense but a measure of the transverse correlation area, together with longitudinal correlations, of two-parton distributions in the proton (Dokshitzer, 2012).

This QCD formulation also distinguishes mechanisms that are often conflated in phenomenology. The literature separates nonperturbative W+W+5 scattering, mixed W+W+6 scattering in which one proton supplies two partons through perturbative splitting, and the double-perturbative-splitting term W+W+7, which is argued not to be counted as MPI in the same sense because it behaves like a loop correction to the ordinary W+W+8 process and must be subtracted when isolating genuine MPI (Dokshitzer, 2012). A plausible implication is that any operational definition of MMPart that relies only on event topology, without this subtraction logic, risks mixing genuine double scattering with perturbative SPS backgrounds.

3. Experimental observables, UE topology, and DPS-sensitive channels

Because the underlying event cannot be uniquely separated from the hard scatter on an event-by-event basis, experiments analyze activity in topological regions defined relative to a hard reference object. CMS uses the azimuthal regions W+W+9 (“toward”), pppp0 (“transverse”), and pppp1 (“away”), with the transverse region being especially sensitive to UE/MPI because it is least contaminated by the primary hard scattering (Gunnellini, 2013). In both QCD and Drell–Yan events, CMS measured the average number of primary charged particles per unit pppp2 and pppp3, the average scalar sum of transverse momentum per unit pppp4 and pppp5, and the ratio of these two quantities (Gunnellini, 2013). In QCD events, the transverse activity rises rapidly with leading-track pppp6 up to about pppp7 GeV and then saturates; in Drell–Yan events, it rises more slowly and steadily across the dimuon pppp8 range (Gunnellini, 2013).

Hard DPS studies use correlation observables designed to distinguish two independent hard scatterings from a single hard interaction with additional radiation. In the pppp9-jet channel, CMS selected events with exactly one muon with pppp0 and pppp1, missing transverse energy pppp2, pppp3 transverse mass pppp4, and two anti-pppp5 jets with pppp6, each with pppp7 and pppp8 (Gunnellini, 2013). The DPS-sensitive variables are the jet azimuthal separation pppp9, the normalized dijet imbalance

pp0

and

pp1

which measures the azimuthal decorrelation between the pp2 system and the dijet system (Gunnellini, 2013). CMS reported that MadGraph predictions with MPI describe the data well, predictions without MPI do not describe the measurement in normalization or shape, pp3 is much more sensitive than pp4, with deviations of about pp5–pp6 in the DPS-sensitive region when MPI are absent, and PYTHIA 8 underestimates the measurements by a factor of about pp7–pp8 (Gunnellini, 2013).

A complementary DPS-sensitive benchmark is the exclusive four-light-jet final state. CMS used 2010 data with exactly four anti-pp9 jets (σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),0), two “hard” jets with σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),1, two “soft” jets with σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),2, and all jets central, σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),3 (Gunnellini, 2013). The observables include jet-σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),4 spectra, σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),5 for the soft pair, and a normalized transverse-momentum imbalance for the soft pair. Powheg+Pythia Z2 tends to overshoot the data by around σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),6–σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),7, especially in the low-σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),8 region of the hard jets up to about σ(A,B)DPS=m2i,j,k,ldx1dx2dx1dx2d2b  Γij(x1,x2,b;t1,t2)Γkl(x1,x2,b;t1,t2)σ^ikA(x1,x1)σ^jlB(x2,x2),\sigma^{DPS}_{(A,B)} = \frac{m}{2}\sum_{i,j,k,l}\int dx_1dx_2dx_1'dx_2'd^2b\; \Gamma_{ij}(x_1,x_2,b;t_1,t_2)\Gamma_{kl}(x_1',x_2',b;t_1,t_2) \hat{\sigma}^A_{ik}(x_1,x_1')\hat{\sigma}^B_{jl}(x_2,x_2'),9 and across the full Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)0 range of the soft jets; agreement for the leading-jet Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)1 improves significantly when MPI are switched off in Powheg+Pythia6 Z2, indicating that MPI modeling tuned for LO generators does not work equally well with an NLO generator (Gunnellini, 2013). The measurements therefore function less as a clean DPS isolation than as a benchmark exposing the interplay of SPS higher-order radiation, showering, and MPI.

4. Event-by-event MPI inference and machine-learning regression

A recent methodological development within MMPart is the attempt to classify Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)2 events by their actual MPI activity rather than by a single multiplicity estimator. In PYTHIA 8.244, charged-particle production in events with large Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)3, normalized to minimum-bias Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)4 production and then to the corresponding Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)5, exhibits a characteristic structure summarized by

Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)6

For Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)7 GeV/Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)8, this ratio is approximately flat and close to unity, consistent with “binary scaling”; in contrast, for intermediate Γij(x1,x2,b;t1,t2)\Gamma_{ij}(x_1,x_2,b;t_1,t_2)9–8 GeV/x1,x2x_1,x_20, a clear bump appears, its size grows with increasing x1,x2x_1,x_21, and the bump is about a x1,x2x_1,x_22 effect reminiscent of the Cronin-like enhancement seen in x1,x2x_1,x_23–Pb nuclear modification factors (Ortiz et al., 2020). Because PYTHIA does not include parton energy loss, the high-x1,x2x_1,x_24 behavior is interpreted as the expected consequence of independent hard scatterings, whereas the intermediate-x1,x2x_1,x_25 structure is attributed mainly to color reconnection; when color reconnection is turned off, the ratio becomes nearly flat and loses much of this structure (Ortiz et al., 2020).

To access such MPI-sensitive event classes in data, the analysis uses TMVA with boosted decision trees trained in regression mode, specifically gradient-boosted trees with a Huber loss, on PYTHIA 8.244 tune 4C events at x1,x2x_1,x_26 TeV (Ortiz et al., 2020). The target is the actual MPI activity. The three main event-level inputs are forward charged-particle multiplicity in ALICE VZERO-like regions x1,x2x_1,x_27 and x1,x2x_1,x_28, transverse spherocity, and the event-by-event average transverse momentum x1,x2x_1,x_29 (Ortiz et al., 2020). The ALICE-like midrapidity selection is charged particles with bb0 GeV/bb1 and bb2, and spherocity is required only for events with more than two such primary charged particles. Spherocity is defined as

bb3

and is used to separate pencil-like dijet events from isotropic ones (Ortiz et al., 2020).

The regression is tested for robustness by retraining with alternative observables, different bb4 distributions, and different generator tunes, including PYTHIA tune 4C, Monash 2013, 2C, and HERWIG 7.1 (Ortiz et al., 2020). The output reproduces the expected bb5 versus bb6, is largely insensitive to the presence or absence of color reconnection, and, when MPI are turned off in simulation, remains consistent with unity within uncertainties (Ortiz et al., 2020). For the event-classification problem, training with a flat bb7 distribution improves the description of the high-bb8 tail (Ortiz et al., 2020).

Applied to ALICE INELbb9 xx00 data, where INELxx01 means events with at least one primary charged particle within xx02, the regression yields

xx03

and

xx04

The quoted uncertainty includes the systematic spread from event-selection and model-dependence studies, and the extracted values are broadly consistent with the PYTHIA energy dependence (Ortiz et al., 2020). The methodological significance is that the classifier combines forward multiplicity, spherocity, and xx05 to build a more inclusive proxy for the underlying partonic structure than a midrapidity multiplicity cut alone.

5. Event generators, modeling assumptions, and color reconnection

Monte Carlo generators are indispensable to MMPart studies, but the literature treats their MPI sectors as partly perturbative and partly phenomenological. PYTHIA and HERWIG combine perturbative multi-scattering models, parton showers, hadronization, and beam-remnant or soft-interaction models, and both require tuning to UE and minimum-bias data (Bartalini et al., 2011). PYTHIA includes multiple scatterings and, in more advanced implementations, rescattering, joint interactions, and more detailed impact-parameter dependence, whereas HERWIG++ is described as taking a more minimalistic approach, focusing on xx06 MPI with a specific framework for soft activity (Bartalini et al., 2011).

The limitations are explicit. Generator implementations often approximate multi-parton distribution functions by products of single PDFs, assume longitudinal–transverse factorization, and simplify color, spin, and flavor correlations (Bartalini et al., 2011). The review literature emphasizes that these assumptions are not well established and are likely too simple. In particular, color reconnection and color correlations may matter substantially and can resemble or interfere with hadronization modeling (Bartalini et al., 2011).

The regression study in PYTHIA 8.244 offers a concrete example of this dependence. The intermediate-xx07 enhancement in xx08 is mainly attributed to color reconnection, and turning color reconnection off renders the ratio nearly flat (Ortiz et al., 2020). This does not imply that all collective-like features in small systems reduce to color reconnection; rather, it shows that generator-level MPI observables can be highly sensitive to hadronization topology. A plausible implication is that any MMPart-based interpretation of data must distinguish carefully between sensitivity to the number of partonic scatterings and sensitivity to the hadronization model used to map those scatterings into final-state observables.

Generator comparisons in DPS-sensitive measurements support the same caution. In xx09dijet, MadGraph+Pythia6 Z2* with MPI describes the data well, while PYTHIA 8 underestimates the measurement by about xx10–xx11, especially in the DPS-sensitive region, a difference explicitly linked not simply to “missing DPS” but to the limited hard-process structure in PYTHIA 8, which generates only xx12 and xx13 processes (Gunnellini, 2013). In four-jet production, Powheg+Pythia6 Z2 overshoots the data even though several correlation shapes are reproduced reasonably well, showing that normalization and topology are not controlled by a single MPI dial (Gunnellini, 2013).

6. Signatures, misconceptions, and open problems

One recurrent misconception is to treat xx14 as an ordinary interaction cross section. The QCD analysis argues instead that xx15, or equivalently the effective cross-section parameter, measures the transverse correlation area together with longitudinal correlations of two-parton distributions in the proton (Dokshitzer, 2012). Likewise, the generator-oriented review stresses that xx16 is a probe of proton geometry and parton correlations rather than merely a fit constant (Bartalini et al., 2011).

A second misconception is that DPS can be established by total rate alone. Both theory and experiment emphasize geometry and correlation structure. The QCD treatment identifies back-to-back four-jet configurations as the clean MPI signature because ordinary xx17 production tends to give a broader “hedgehog” topology, whereas double hard scattering tends to produce two nearly back-to-back jet pairs with double collinear enhancement (Dokshitzer, 2012). CMS likewise relies on decorrelation observables such as xx18, normalized imbalances, and pairwise xx19 variables rather than on inclusive counting alone (Gunnellini, 2013).

A third misconception is that multiplicity selection is equivalent to MPI selection. The machine-learning study was motivated precisely by the concern that multiplicity estimators can bias a sample toward jet-like or otherwise hard events, obscuring the connection to the number of partonic scatterings (Ortiz et al., 2020). The use of forward multiplicity to reduce autocorrelations with midrapidity spectra, together with spherocity and xx20, reflects an effort to classify events more directly by their underlying partonic structure (Ortiz et al., 2020).

The main open problems follow directly from the cited literature. The review of MPI at the LHC lists the structure and evolution of multi-parton distributions, the validity of longitudinal/transverse factorization, spin/color/flavor correlations, consistent perturbative treatment of multiple emissions and rescattering, and the connection between soft and hard MPI in a unified framework as unresolved issues (Bartalini et al., 2011). The QCD analysis adds the need to separate genuine MPI from perturbative backgrounds such as double splitting (Dokshitzer, 2012). The machine-learning program suggests a further experimental direction: organizing xx21 data by an MPI-sensitive event classifier may help determine whether the collective-like trends seen at the LHC originate from multipartonic initial conditions, hadronization effects, or both (Ortiz et al., 2020).

In that sense, MMPart is best understood not as a single model but as a layered research domain. At its most inclusive, it covers the use of MPI to describe minimum-bias and UE activity; in its perturbative core, it studies DPS through multi-parton distributions and effective overlap measures; in current data analysis, it includes event-shape and machine-learning strategies that attempt to reconstruct the hidden partonic activity of individual xx22 events (Bartalini et al., 2011).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MMPart.