Papers
Topics
Authors
Recent
Search
2000 character limit reached

Matrix Multiplicative Weights Algorithm

Updated 25 September 2026
  • Matrix multiplicative weights (MMW) are a generalization of multiplicative-weights and exponentiated-gradient methods, replacing probability distributions with positive semidefinite trace-one matrices (density matrices) and scalar normalization with trace normalization.
  • MMW is used in a variety of applications including semidefinite programming, zero-sum and quantum games, online eigenvector problems, and quadratic optimization, offering a convergence rate to Nash equilibria in zero-sum games.
  • Standard MMW has an average regret of order O(√(T logn)), which can be achieved with choice of learning rate η ∝ √log(n)/T; efficient computational approximations are possible.

Matrix multiplicative weights (MMW) is a matrix-valued generalization of multiplicative-weights and exponentiated-gradient methods in which probability distributions are replaced by positive semidefinite trace-one matrices, payoff vectors by Hermitian matrices, and scalar normalization by trace normalization. The canonical decision domain is the spectrahedron

An={X∈Sn:X⪰0, tr⁡(X)=1},\mathcal A_n=\{X\in\mathbb S^n:X\succeq 0,\ \operatorname{tr}(X)=1\},

whose elements are density matrices. MMW is used for online optimization over semidefinite domains, zero-sum and quantum games, semidefinite programming, online eigenvector problems, and quadratic optimization. Its central update maps an accumulated symmetric gain matrix YY to

Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.

1. Mathematical formulation and canonical update

At round tt, an MMW learner selects Xt∈AnX_t\in\mathcal A_n, receives a symmetric gain matrix Gt∈SnG_t\in\mathbb S^n, and obtains payoff

⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).

The regret against the best fixed matrix in An\mathcal A_n is

max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.

Because

max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),

the comparator term equals

YY0

The standard matrix exponentiated-gradient update selects

YY1

where YY2 is the learning rate. If

YY3

then

YY4

The matrix exponential ensures positive semidefiniteness, while trace normalization ensures unit trace.

The associated potential is

YY5

with gradient

YY6

Its convex conjugate is the negative von Neumann entropy, represented on positive definite YY7 by

YY8

The induced Bregman divergence is quantum relative entropy,

YY9

When Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.0, standard MMW gives average regret of order

Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.1

and cumulative regret of order

Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.2

A choice such as

Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.3

produces this scale. The Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.4 term arises from the entropy diameter of the spectrahedron.

MMW is distinct from a self-coupled Hedge map used in some analyses of symmetric games. In that setting,

Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.5

where the same strategy Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.6 supplies both the population distribution and the opponent against which payoffs are evaluated. This is a nonlinear single-population dynamical system, not the usual two-process row-player/column-player matrix MMW algorithm. Its fixed points and KL-divergence identities do not by themselves establish convergence of empirical play to Nash equilibrium (Avramopoulos, 2016).

2. Mirror descent, entropy, and regret analysis

The geometry of MMW is generated by the von Neumann entropy. For density matrices Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.7 and Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.8, the quantum relative entropy is

Pmw(Y)=eYtr⁡(eY).P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.9

In payoff-based analyses, the entropy satisfies the semidefinite Pinsker-type inequality

tt0

This strong convexity supplies the stability term in the mirror-descent analysis.

For an MMW update driven by tt1, the basic energy inequality is

tt2

Summing over rounds yields a matrix regret estimate of the form

tt3

For one player, the entropy diameter can be bounded by tt4; in a two-player setting the corresponding quantity is

tt5

In two-player zero-sum semidefinite games, the averaged iterates

tt6

satisfy a duality-gap estimate

tt7

where tt8 is a Lipschitz or gradient bound. Thus tt9 iterations suffice for an Xt∈AnX_t\in\mathcal A_n0-Nash equilibrium under full matrix-gradient feedback.

Relative entropy also appears in continuous-time and quantum formulations. In a two-player zero-sum quantum game with a fully mixed Nash equilibrium Xt∈AnX_t\in\mathcal A_n1, quantum replicator dynamics preserve

Xt∈AnX_t\in\mathcal A_n2

where

Xt∈AnX_t\in\mathcal A_n3

The conservation law is a continuous-time analogue of the conserved divergence in classical zero-sum replicator dynamics. It constrains the joint evolution of the two density matrices but does not imply that either state has a fixed spectrum or that the trajectory converges to equilibrium (Jain et al., 2022).

3. Computational approximations and low-rank sketches

The dominant computational obstacle in standard MMW is the matrix exponential. General-purpose eigendecomposition costs approximately Xt∈AnX_t\in\mathcal A_n4 in practice, while full matrix exponentiation requires maintaining an Xt∈AnX_t\in\mathcal A_n5 matrix.

A rank-one randomized sketch replaces the full matrix softmax by

Xt∈AnX_t\in\mathcal A_n6

where Xt∈AnX_t\in\mathcal A_n7 is sampled uniformly from the unit sphere. The resulting matrix is positive semidefinite, rank one, and has unit trace. The sketched algorithm plays

Xt∈AnX_t\in\mathcal A_n8

This sketch is not unbiased for the standard MMW projection: Xt∈AnX_t\in\mathcal A_n9 in general. Instead, define

Gt∈SnG_t\in\mathbb S^n0

The averaged projection is the gradient of the convex potential

Gt∈SnG_t\in\mathbb S^n1

Its convex conjugate supplies the regularizer for the averaged algorithm.

The relevant Bregman divergence satisfies the smoothness bound

Gt∈SnG_t\in\mathbb S^n2

and the diameter bound

Gt∈SnG_t\in\mathbb S^n3

These properties allow mirror-descent analysis with only constant-factor degradation relative to standard MMW. The resulting cumulative regret is bounded by

Gt∈SnG_t\in\mathbb S^n4

For Gt∈SnG_t\in\mathbb S^n5, this gives average regret

Gt∈SnG_t\in\mathbb S^n6

The actual random sketch requires an adversary condition: conditional on the past, Gt∈SnG_t\in\mathbb S^n7 must be independent of the fresh random vector Gt∈SnG_t\in\mathbb S^n8. Under this condition,

Gt∈SnG_t\in\mathbb S^n9

so expected regret is controlled by the deterministic averaged-projection analysis. High-probability bounds follow by treating

⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).0

as a martingale difference.

The matrix-exponential–vector product ⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).1 can be approximated using ⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).2 Lanczos iterations. If

⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).3

is the Lanczos tridiagonalization, then

⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).4

approximates ⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).5. The resulting computational cost is

⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).6

where ⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).7 is the matrix-vector multiplication cost and ⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).8 is the floating-point word size. The rank-one method is advantageous when matrix-vector multiplication is substantially cheaper than dense eigendecomposition, particularly for sparse or structured matrices (Carmon et al., 2019).

4. Applications to eigenvectors, semidefinite programming, and quadratic optimization

The spectrahedral domain contains rank-one matrices ⟨Gt,Xt⟩=tr⁡(Gt⊤Xt).\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).9, so MMW directly induces algorithms for online eigenvector problems. The benchmark is

An\mathcal A_n0

The rank-one method selects

An\mathcal A_n1

With exact exponential-vector products, the expected regret is

An\mathcal A_n2

With Lanczos approximation, the total matrix-vector-product complexity for high-probability An\mathcal A_n3-average regret is

An\mathcal A_n4

and the method improves the previous best online-eigenvector complexity by a factor of at least An\mathcal A_n5.

MMW and Oja’s algorithm are related but not equivalent in general. If the input matrices An\mathcal A_n6 share a common orthonormal eigenbasis An\mathcal A_n7, with

An\mathcal A_n8

then the squared eigen-coordinates of Oja’s normalized vector,

An\mathcal A_n9

satisfy

max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.0

After rewriting this expression, max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.1 follows an ordinary multiplicative-weights recursion. Thus Oja’s vector dynamics reduce exactly to scalar MW over eigen-directions in the common-eigenbasis regime. For arbitrary noncommuting matrices, the squared-coordinate dynamics do not close, and this reduction fails (Garber, 2023).

MMW also applies to semidefinite programs in saddle-point form. Let

max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.2

and

max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.3

For symmetric matrices max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.4, define

max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.5

The matrix player uses the rank-one sketch, while the simplex player uses ordinary multiplicative weights. The duality gap is

max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.6

If

max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.7

then the averaged strategies satisfy

max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.8

To achieve expected gap at most max⁡X∈An∑t=1T⟨Gt,X⟩−∑t=1T⟨Gt,Xt⟩.\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle -\sum_{t=1}^T\langle G_t,X_t\rangle.9, it suffices to take

max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),0

5. Quantum games and payoff-based MMW

In quantum games, a mixed strategy is a density matrix

max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),1

For a two-player game, Alice and Bob use density matrices max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),2 and max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),3. A payoff observable max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),4 induces the bilinear payoff

max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),5

The matrix multiplicative-weights updates are

max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),6

and

max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),7

The negative sign in Bob’s update reflects minimization of Alice’s payoff. When all matrices are diagonal in a common basis, MMWU reduces exactly to classical exponential MWU. Noncommutativity is the principal quantum distinction: accumulated payoff operators need not commute with their instantaneous derivatives, and the derivative of max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),8 is generally

max⁡X∈An⟨Y,X⟩=λmax⁡(Y),\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),9

The continuous limit defines quantum replicator dynamics through

YY00

and

YY01

For a fully mixed Nash equilibrium, the total quantum relative entropy is conserved: YY02 The same setting yields Poincaré recurrence for almost every interior initial condition. The proof uses canonical coordinates, volume preservation, boundedness from the relative-entropy invariant, and the Poincaré recurrence theorem. The result establishes recurrence rather than convergence: generic trajectories return arbitrarily close to their initial conditions infinitely often.

Full-information MMW assumes access to the complete payoff-gradient matrix. Payoff-based learning replaces it with an estimator. In minimal-information matrix multiplicative weights, or 3MW, the update is

YY03

The estimator decomposes as

YY04

where YY05 is smoothing bias and YY06 is a martingale-noise term.

For deterministic scalar payoff feedback, a two-point estimator achieves

YY07

with a factor linear in the effective matrix dimension YY08. For a single random payoff-observable realization, a one-point estimator has variance of order YY09, leading to

YY10

The rate difference arises from the bias–variance tradeoff: the smoothing bias is YY11, whereas the one-point estimator variance is YY12. In general non-zero-sum games, a regularized 3MW method has local, high-probability, last-iterate convergence to equilibria satisfying variational stability (Lotidis et al., 2023).

6. Evolutionary, game-theoretic, and dynamical interpretations

MMW also arises exactly in the marginal allele dynamics of sexual evolutionary models. For selection before recombination, let YY13 be the joint population distribution and let

YY14

be the allele marginal. If

YY15

is the conditional fitness of allele YY16, then the evolutionary update satisfies

YY17

which is exactly the parameter-free, correlation-sensitive polynomial-weights update. This identity holds for arbitrary nonnegative fitness matrices, recombination rates, initial distributions, and numbers of loci.

For recombination before selection, the marginal update uses a mixture of correlation-sensitive and independent payoffs: YY18 where

YY19

The resulting update is

YY20

Thus recombination before selection corresponds to an interpolated parameter-free PW process.

The correspondence is between allele marginals and MW strategies, not between the complete joint population distribution and independent mixed strategies. A product distribution satisfies

YY21

whereas linkage disequilibrium is

YY22

Product distributions are generally not preserved, even under weak selection. Consequently, the uncorrelated MWUA based only on marginal distributions need not describe actual evolutionary dynamics, and it can converge to a different equilibrium from the sexual population process. The precise correspondence therefore retains the evolving correlations through conditional payoffs (Meir et al., 2015).

In symmetric bimatrix games, the Hedge map

YY23

has fixed points characterized by equal payoffs on the support of YY24. Interior fixed points are interior symmetric equilibria, while boundary fixed points need not be equilibria of the full game. The map admits a relative-entropy identity

YY25

whose second derivative is an exponentially tilted payoff variance and is therefore nonnegative. These are valid local and variational properties of the Hedge map, but they do not imply convergence of the dynamics to equilibrium or establish a polynomial-time equilibrium algorithm. In particular, the claimed implication that every symmetric game without an equalizer has a weakly dominated pure strategy is not established, and the resulting claim YY26 is not a valid consequence of the argument (Avramopoulos, 2016).

Across these settings, the principal distinction is between exact matrix-valued multiplicative weights, low-rank or payoff-based approximations, and scalar multiplicative-weights reductions available under additional structure. Standard MMW accommodates arbitrary symmetric, including noncommuting, matrix sequences. Rank-one sketches preserve regret through a different averaged mirror map. Oja’s algorithm reduces to ordinary MW only under a common eigenbasis. Quantum MMW extends the geometry to density matrices and yields conservation and recurrence phenomena in zero-sum games. Evolutionary dynamics reproduce MMW exactly at the level of allele marginals while retaining correlations in the underlying joint population.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Matrix Multiplicative Weights.