---
title: Skorohod's Functional M1 Topology
url: https://www.emergentmind.com/topics/skorohod-s-functional-m1-topology
type: topic
---

# Skorohod's Functional M1 Topology

toology. Use same spelling as topic perhaps "Skorohod's functional M1 topology" while acknowledging Skorokhod space? Fine.

Need integrate fresh arXiv search. Let's search for relevant papers maybe provided ones.
Skorohod’s functional \(M_1\) topology is a topology on spaces of càdlàg paths in which convergence is organized through the completed graph of a path rather than through exact one-to-one matching of jumps. It is designed for stochastic-process limits with discontinuities, especially when a prelimit path contains rapid successions of jumps or steep local excursions that collapse to a single jump in the limit. In this sense \(M_1\) is weaker than \(J_1\), but it is not merely a permissive substitute: it has its own graph geometry, compactness criteria, continuity theory, and multivariate variants, and it is often the correct mode of convergence for heavy-tailed time series, continuous-time random walks, queueing limits, and stochastic integrals with jump misalignment [2210.16026] [2405.01318].

## 1. Geometric construction on Skorohod space

The ambient state space is the Skorohod space \(D([0,1],\mathbb{R})\), or more generally \(D([0,1],\mathbb{R}^d)\), consisting of right-continuous functions with left limits. For \(x\in D([0,1],\mathbb{R})\), the completed graph is
\[
\Gamma_x=\{(t,z)\in[0,1]\times\mathbb{R}: z=\alpha x(t-)+(1-\alpha)x(t)\text{ for some }\alpha\in[0,1]\}.
\]
Equivalently, \(\Gamma_x\) contains the usual graph together with the vertical segment joining \(x(t-)\) and \(x(t)\) at each jump time. The graph is ordered by
\[
(t_1,z_1)\le (t_2,z_2)
\quad\text{iff}\quad
t_1<t_2,\ \text{or } t_1=t_2 \text{ and } |x(t_1-)-z_1|\le |x(t_2-)-z_2|.
\]

A parametric representation of \(\Gamma_x\) is a continuous, nondecreasing surjection \((u,r)\) from \([0,1]\) onto \(\Gamma_x\), where \(r\) is the time component and \(u\) is the spatial component. Writing \(\Pi(x)\) for the set of such parametrizations, the standard \(M_1\) metric is
\[
d_{M_1}(x,y)
=
\inf_{(u_x,r_x)\in\Pi(x),\,(u_y,r_y)\in\Pi(y)}
\left\{
\|u_x-u_y\|_\infty\vee \|r_x-r_y\|_\infty
\right\}.
\]
This metric induces the \(M_1\) topology. The topology is Polish; some explicit metrics used in expositions are not complete, but equivalent complete metrics exist [2210.16026].

The basic geometric feature is that the parametric curve is allowed to move along the inserted jump segment. Consequently, convergence can occur even when the approximating paths do not reproduce the limit jump at exactly the same time or in exactly the same way; what matters is the ordered alignment of completed graphs, not literal coincidence of discontinuity patterns.

## 2. Functional variants: strong and weak multivariate \(M_1\)

In dimension \(d\ge 2\), two distinct constructions appear. For \(x\in D([0,1],\mathbb{R}^d)\), the weak formulation uses the completed “thick” graph
\[
G_x=\{(t,z)\in[0,1]\times\mathbb{R}^d: z\in [[x(t-),x(t)]]\},
\]
where
\[
[[a,b]]=[a_1,b_1]\times\cdots\times[a_d,b_d].
\]
A weak parametric representation is a continuous, nondecreasing map \((r,u)\) from \([0,1]\) into \(G_x\), with \(r(0)=0\), \(r(1)=1\), and \(u(1)=x(1)\). If \(\Pi_w(x)\) denotes the set of weak parametric representations, the induced weak-\(M_1\) metric is
\[
d_w(x_1,x_2)
=
\inf\left\{
\|r_1-r_2\|_{[0,1]}\vee \|u_1-u_2\|_{[0,1]}:
(r_i,u_i)\in\Pi_w(x_i),\ i=1,2
\right\}.
\]

The strong formulation uses the “thin” graph
\[
\Gamma_x=\{(t,z)\in[0,1]\times\mathbb{R}^d: z\in [x(t-),x(t)]\},
\]
where
\[
[a,b]=\{\lambda a+(1-\lambda)b:0\le \lambda\le 1\}.
\]
Its metric is
\[
d_{M1}(x_1,x_2)
=
\inf\left\{
\|r_1-r_2\|_{[0,1]}\vee \|u_1-u_2\|_{[0,1]}:
(r_i,u_i)\in\Pi(x_i),\ i=1,2
\right\}.
\]

The distinction is structural. Weak \(M_1\) is the product topology of coordinatewise one-dimensional \(M_1\), and in the papers cited it coincides with
\[
d_p(x_1,x_2)=\max_{1\le j\le d} d_{M_1}(x_{1j},x_{2j}).
\]
Strong \(M_1\) requires a common traversal of the vector-valued completed graph and therefore preserves joint jump geometry across coordinates. Weak \(M_1\) permits coordinatewise alignment and is therefore better adapted to asynchronous jumps; strong \(M_1\) is more restrictive and may fail precisely when different coordinates jump on different microscopic time scales [1308.3624] [2405.01318].

A common misconception is that “functional \(M_1\)” is automatically unique in the multivariate case. It is not. In \(\mathbb{R}^d\), the weak and strong topologies differ, and the choice between them is substantive rather than cosmetic.

## 3. Why \(M_1\) differs from \(J_1\)

The \(J_1\) topology is based on small homeomorphic time changes. It is strongest when jumps can be matched one-by-one in both location and magnitude. By contrast, \(M_1\) compares completed graphs and can absorb several same-direction jumps into a single limiting jump. This is the fundamental reason it appears in dependent heavy-tail theory.

The basic prototype is a staircase path. In regularly varying time series with dependence, extremes often occur in clusters. Partial sums then exhibit rapid successions of same-sign jumps over short intervals, and in the scaling limit those clusters collapse to single jumps. Under \(J_1\), such a cluster cannot generally be aligned with a single discontinuity without violating the topology’s matching constraints. Under \(M_1\), the cluster can be traversed through the completed graph and identified with one monotone jump segment in the limit [1001.1345] [2405.01318].

The classical illustrative example is
\[
x_n(t)=\tfrac12\,1_{[\frac12-\frac1n,\frac12)}(t)+1_{[\frac12,1]}(t),
\qquad
x(t)=1_{[\frac12,1]}(t).
\]
Then \(x_n\to x\) in \(M_1\), but not in \(J_1\) and not uniformly. The two nearby jumps in \(x_n\) approximate the single jump of \(x\) through the completed-graph representation rather than through exact jump matching [1001.1345].

This flexibility is not unlimited. \(M_1\) does not ignore arbitrary oscillation. It tolerates steep ramps and clustered monotone jump behavior, but it still controls local deviations from line segments in the completed graph. A plausible implication is that \(M_1\) should be regarded as a geometry of admissible jump aggregation rather than as a generic weak topology for all discontinuous phenomena.

## 4. Compactness, tightness, and convergence criteria

A central quantitative device is the \(M_1\) oscillation modulus. In one formulation,
\[
\omega_\delta''(f)=\sup_{t\in[0,1]} w(f,t,\delta),
\]
with
\[
w(f,t,\delta)
=
\sup_{0\vee t-\delta \le t_1<t_2<t_3<t+\delta\wedge 1}
\inf_{z\in [f(t_1),f(t_3)]}
|f(t_2)-z|.
\]
Equivalently, it measures how far an intermediate value deviates from the line segment between nearby endpoints. For monotone functions this modulus vanishes, which explains why monotone jump processes fit naturally into \(M_1\) [2210.16026].

A subset \(K\subseteq D([0,1],\mathbb{R})\) is compact in \(M_1\) if and only if it is uniformly bounded and satisfies the oscillation and boundary controls
\[
\lim_{\delta\downarrow 0}\sup_{f\in K}\omega_\delta''(f)=0,
\]
\[
\lim_{\delta\downarrow 0}\sup_{f\in K}|f(\delta)-f(0)|=0,
\qquad
\lim_{\delta\downarrow 0}\sup_{f\in K}|f(1-)-f(1-\delta)|=0.
\]
The corresponding tightness criterion for stochastic processes replaces these uniform bounds by convergence in probability. Weak convergence in \(M_1\) is then characterized by \(M_1\)-tightness together with convergence of finite-dimensional distributions at continuity times of the limit [2210.16026].

For monotone functions, \(M_1\) convergence reduces to pointwise convergence on a dense set together with endpoint convergence. This extends, by a cut-and-paste argument, to piecewise monotone functions provided the limit is continuous at the cutting points. That criterion is particularly useful in proofs based on point-process limits and piecewise monotone summation maps [1001.1345].

These criteria clarify another common misconception: \(M_1\) is weaker than \(J_1\), but it is not topologically indiscriminate. Its compactness theory is highly structured and is tailored to graphs that are locally close to ordered line segments.

## 5. Continuity of functional transformations

The functional usefulness of \(M_1\) depends on continuity results for nonlinear maps on path space. One important example is the integral representation
\[
y(t)=x(t)+\int_0^t h(y(s))\,ds,
\]
with \(h\) Lipschitz. This map is continuous on \(D\) endowed with the \(M_1\) topology. The proof uses a refined characterization of \(M_1\) convergence in which the time components of parametric representations are absolutely continuous, have uniformly bounded derivatives, and converge in \(L_1\) [1001.2381].

Heavy-tail applications require additional mapping results. In the self-normalized partial-sum setting, multiplication is continuous in \(M_1\) when jumps do not change sign jointly, and the division map \(h(x,y)=x/y\) is continuous when the denominator lies in the measurable subset \(C_0^{\uparrow}\) of continuous, nondecreasing functions with \(y(0)>0\). These facts permit the continuous mapping theorem for ratio processes such as
\[
\frac{S_{\lfloor n\cdot\rfloor}}{V_n}
\Rightarrow
\frac{L_1(\cdot)}{\sqrt{L_2(1)}}
\]
in \(M_1\) [2405.01318].

For stochastic integrals, the continuity theory is subtler. Under good decompositions for the semimartingale integrators and the asymptotically vanishing consecutive increments condition, one obtains weak convergence of
\[
\left(X^n,\int_0^\cdot H^n_{s-}\,dX^n_s\right)
\]
in \(J_1\) or \(M_1\), depending on the topology used for the input convergence [2309.12197]. More recent work shows that, in the purely \(M_1\) setting, relative compactness of Itô integrals can still be established without the classical AVCI condition, but limit points may contain a jump-product correction term of the form
\[
\int_0^\bullet \tilde H^0_{s-}\,d\tilde X^0_s
+
\sum_{\sigma\in\mathcal{T}(\tilde H^0)}
\tilde\xi^0_\sigma\,\Delta\tilde H^0_\sigma\,\Delta\tilde X^0_\sigma\,1_{[\sigma,\infty)},
\]
where the weights \(\tilde\xi^0_\sigma\) take values in \([0,1]\) and encode the completed-graph alignment of prelimit jumps [2508.21624].

This body of results shows that \(M_1\) is not only a convergence topology for raw paths. It also supports a nontrivial calculus of path transformations, provided the jump geometry of the map is compatible with completed-graph alignment.

## 6. Applications and generalizations

The most developed probabilistic applications concern heavy-tailed limits. For stationary regularly varying sequences with clustered extremes, properly centered partial-sum processes and partial sums of squares converge jointly in weak \(M_1\) to stable Lévy components, and self-normalization by \(V_n=(\sum X_i^2)^{1/2}\) yields an \(M_1\) limit for \(S_{\lfloor n\cdot\rfloor}/V_n\). Earlier univariate work of Basrak–Krizmanić–Segers had already shown that \(M_1\) is the correct topology for stable limits of dependent sequences with infinite variance when clustering causes \(J_1\) to fail [2405.01318] [1001.1345].

In queueing theory, continuity of the \(M_1\) integral representation yields heavy-traffic limits for many-server models whose limit processes have jumps unmatched in the converging sequence, as can occur with bursty arrival processes or service interruptions [1001.2381]. In continuous-time random walks, the linear interpolation map that replaces stairs by segments is continuous in strong \(M_1\) under the matching conditions stated in the paper, and functional limit theorems for continuous-path CTRWs therefore require strong \(M_1\) rather than \(J_1\) in the generic case [1305.4058]. For embedded Markov chains, linear interpolation is the embedding naturally associated with \(M_1\), and \(J_1\) convergence of the step embedding implies convergence of the \(M_1\) embedding in the corresponding topology [1409.4656].

The topology also admits substantial extensions beyond finite-dimensional Euclidean path space. In the strong dual \(E'\) of a countably Hilbertian nuclear space, a functional \(M_1\) topology can be defined on \(D([0,T],E')\) through bounded-set pseudometrics \(d_{B,M1}\). Compactness and tightness then admit Mitoma-type projection criteria: a set or sequence is compact or tight in \(D_{E'}\) under \(M_1\) if and only if all scalar projections onto test functions are compact or tight in scalar \(M_1\) [1509.02855]. In a different direction, an ordered-Hausdorff construction based on filled-in graphs and compatible betweenness extends classical \(M_1\) from fixed-domain real-valued càdlàg functions to paths over general metrisable spaces and varying closed time domains, while recovering the classical topology on \(D([0,1],\mathbb{R})\) with linear betweenness [2301.05637].

These developments suggest a coherent picture. Skorohod’s functional \(M_1\) topology is the topology of completed-graph convergence for càdlàg paths whose limit behavior is governed by jump aggregation, monotone graph traversal, or asynchronous multicomponent discontinuities. Its distinctive value lies not simply in being weaker than \(J_1\), but in encoding a different notion of path geometry—one that is now central in heavy-tailed asymptotics, stochastic integration with misaligned jumps, time-changed processes, and infinite-dimensional weak convergence [2401.13543].

Source: https://www.emergentmind.com/topics/skorohod-s-functional-m1-topology