Matrix Multiplicative Weights Algorithm
- Matrix multiplicative weights (MMW) are a generalization of multiplicative-weights and exponentiated-gradient methods, replacing probability distributions with positive semidefinite trace-one matrices (density matrices) and scalar normalization with trace normalization.
- MMW is used in a variety of applications including semidefinite programming, zero-sum and quantum games, online eigenvector problems, and quadratic optimization, offering a convergence rate to Nash equilibria in zero-sum games.
- Standard MMW has an average regret of order O(√(T logn)), which can be achieved with choice of learning rate η ∝ √log(n)/T; efficient computational approximations are possible.
Matrix multiplicative weights (MMW) is a matrix-valued generalization of multiplicative-weights and exponentiated-gradient methods in which probability distributions are replaced by positive semidefinite trace-one matrices, payoff vectors by Hermitian matrices, and scalar normalization by trace normalization. The canonical decision domain is the spectrahedron
whose elements are density matrices. MMW is used for online optimization over semidefinite domains, zero-sum and quantum games, semidefinite programming, online eigenvector problems, and quadratic optimization. Its central update maps an accumulated symmetric gain matrix to
1. Mathematical formulation and canonical update
At round , an MMW learner selects , receives a symmetric gain matrix , and obtains payoff
The regret against the best fixed matrix in is
Because
the comparator term equals
0
The standard matrix exponentiated-gradient update selects
1
where 2 is the learning rate. If
3
then
4
The matrix exponential ensures positive semidefiniteness, while trace normalization ensures unit trace.
The associated potential is
5
with gradient
6
Its convex conjugate is the negative von Neumann entropy, represented on positive definite 7 by
8
The induced Bregman divergence is quantum relative entropy,
9
When 0, standard MMW gives average regret of order
1
and cumulative regret of order
2
A choice such as
3
produces this scale. The 4 term arises from the entropy diameter of the spectrahedron.
MMW is distinct from a self-coupled Hedge map used in some analyses of symmetric games. In that setting,
5
where the same strategy 6 supplies both the population distribution and the opponent against which payoffs are evaluated. This is a nonlinear single-population dynamical system, not the usual two-process row-player/column-player matrix MMW algorithm. Its fixed points and KL-divergence identities do not by themselves establish convergence of empirical play to Nash equilibrium (Avramopoulos, 2016).
2. Mirror descent, entropy, and regret analysis
The geometry of MMW is generated by the von Neumann entropy. For density matrices 7 and 8, the quantum relative entropy is
9
In payoff-based analyses, the entropy satisfies the semidefinite Pinsker-type inequality
0
This strong convexity supplies the stability term in the mirror-descent analysis.
For an MMW update driven by 1, the basic energy inequality is
2
Summing over rounds yields a matrix regret estimate of the form
3
For one player, the entropy diameter can be bounded by 4; in a two-player setting the corresponding quantity is
5
In two-player zero-sum semidefinite games, the averaged iterates
6
satisfy a duality-gap estimate
7
where 8 is a Lipschitz or gradient bound. Thus 9 iterations suffice for an 0-Nash equilibrium under full matrix-gradient feedback.
Relative entropy also appears in continuous-time and quantum formulations. In a two-player zero-sum quantum game with a fully mixed Nash equilibrium 1, quantum replicator dynamics preserve
2
where
3
The conservation law is a continuous-time analogue of the conserved divergence in classical zero-sum replicator dynamics. It constrains the joint evolution of the two density matrices but does not imply that either state has a fixed spectrum or that the trajectory converges to equilibrium (Jain et al., 2022).
3. Computational approximations and low-rank sketches
The dominant computational obstacle in standard MMW is the matrix exponential. General-purpose eigendecomposition costs approximately 4 in practice, while full matrix exponentiation requires maintaining an 5 matrix.
A rank-one randomized sketch replaces the full matrix softmax by
6
where 7 is sampled uniformly from the unit sphere. The resulting matrix is positive semidefinite, rank one, and has unit trace. The sketched algorithm plays
8
This sketch is not unbiased for the standard MMW projection: 9 in general. Instead, define
0
The averaged projection is the gradient of the convex potential
1
Its convex conjugate supplies the regularizer for the averaged algorithm.
The relevant Bregman divergence satisfies the smoothness bound
2
and the diameter bound
3
These properties allow mirror-descent analysis with only constant-factor degradation relative to standard MMW. The resulting cumulative regret is bounded by
4
For 5, this gives average regret
6
The actual random sketch requires an adversary condition: conditional on the past, 7 must be independent of the fresh random vector 8. Under this condition,
9
so expected regret is controlled by the deterministic averaged-projection analysis. High-probability bounds follow by treating
0
as a martingale difference.
The matrix-exponential–vector product 1 can be approximated using 2 Lanczos iterations. If
3
is the Lanczos tridiagonalization, then
4
approximates 5. The resulting computational cost is
6
where 7 is the matrix-vector multiplication cost and 8 is the floating-point word size. The rank-one method is advantageous when matrix-vector multiplication is substantially cheaper than dense eigendecomposition, particularly for sparse or structured matrices (Carmon et al., 2019).
4. Applications to eigenvectors, semidefinite programming, and quadratic optimization
The spectrahedral domain contains rank-one matrices 9, so MMW directly induces algorithms for online eigenvector problems. The benchmark is
0
The rank-one method selects
1
With exact exponential-vector products, the expected regret is
2
With Lanczos approximation, the total matrix-vector-product complexity for high-probability 3-average regret is
4
and the method improves the previous best online-eigenvector complexity by a factor of at least 5.
MMW and Oja’s algorithm are related but not equivalent in general. If the input matrices 6 share a common orthonormal eigenbasis 7, with
8
then the squared eigen-coordinates of Oja’s normalized vector,
9
satisfy
0
After rewriting this expression, 1 follows an ordinary multiplicative-weights recursion. Thus Oja’s vector dynamics reduce exactly to scalar MW over eigen-directions in the common-eigenbasis regime. For arbitrary noncommuting matrices, the squared-coordinate dynamics do not close, and this reduction fails (Garber, 2023).
MMW also applies to semidefinite programs in saddle-point form. Let
2
and
3
For symmetric matrices 4, define
5
The matrix player uses the rank-one sketch, while the simplex player uses ordinary multiplicative weights. The duality gap is
6
If
7
then the averaged strategies satisfy
8
To achieve expected gap at most 9, it suffices to take
0
5. Quantum games and payoff-based MMW
In quantum games, a mixed strategy is a density matrix
1
For a two-player game, Alice and Bob use density matrices 2 and 3. A payoff observable 4 induces the bilinear payoff
5
The matrix multiplicative-weights updates are
6
and
7
The negative sign in Bob’s update reflects minimization of Alice’s payoff. When all matrices are diagonal in a common basis, MMWU reduces exactly to classical exponential MWU. Noncommutativity is the principal quantum distinction: accumulated payoff operators need not commute with their instantaneous derivatives, and the derivative of 8 is generally
9
The continuous limit defines quantum replicator dynamics through
00
and
01
For a fully mixed Nash equilibrium, the total quantum relative entropy is conserved: 02 The same setting yields Poincaré recurrence for almost every interior initial condition. The proof uses canonical coordinates, volume preservation, boundedness from the relative-entropy invariant, and the Poincaré recurrence theorem. The result establishes recurrence rather than convergence: generic trajectories return arbitrarily close to their initial conditions infinitely often.
Full-information MMW assumes access to the complete payoff-gradient matrix. Payoff-based learning replaces it with an estimator. In minimal-information matrix multiplicative weights, or 3MW, the update is
03
The estimator decomposes as
04
where 05 is smoothing bias and 06 is a martingale-noise term.
For deterministic scalar payoff feedback, a two-point estimator achieves
07
with a factor linear in the effective matrix dimension 08. For a single random payoff-observable realization, a one-point estimator has variance of order 09, leading to
10
The rate difference arises from the bias–variance tradeoff: the smoothing bias is 11, whereas the one-point estimator variance is 12. In general non-zero-sum games, a regularized 3MW method has local, high-probability, last-iterate convergence to equilibria satisfying variational stability (Lotidis et al., 2023).
6. Evolutionary, game-theoretic, and dynamical interpretations
MMW also arises exactly in the marginal allele dynamics of sexual evolutionary models. For selection before recombination, let 13 be the joint population distribution and let
14
be the allele marginal. If
15
is the conditional fitness of allele 16, then the evolutionary update satisfies
17
which is exactly the parameter-free, correlation-sensitive polynomial-weights update. This identity holds for arbitrary nonnegative fitness matrices, recombination rates, initial distributions, and numbers of loci.
For recombination before selection, the marginal update uses a mixture of correlation-sensitive and independent payoffs: 18 where
19
The resulting update is
20
Thus recombination before selection corresponds to an interpolated parameter-free PW process.
The correspondence is between allele marginals and MW strategies, not between the complete joint population distribution and independent mixed strategies. A product distribution satisfies
21
whereas linkage disequilibrium is
22
Product distributions are generally not preserved, even under weak selection. Consequently, the uncorrelated MWUA based only on marginal distributions need not describe actual evolutionary dynamics, and it can converge to a different equilibrium from the sexual population process. The precise correspondence therefore retains the evolving correlations through conditional payoffs (Meir et al., 2015).
In symmetric bimatrix games, the Hedge map
23
has fixed points characterized by equal payoffs on the support of 24. Interior fixed points are interior symmetric equilibria, while boundary fixed points need not be equilibria of the full game. The map admits a relative-entropy identity
25
whose second derivative is an exponentially tilted payoff variance and is therefore nonnegative. These are valid local and variational properties of the Hedge map, but they do not imply convergence of the dynamics to equilibrium or establish a polynomial-time equilibrium algorithm. In particular, the claimed implication that every symmetric game without an equalizer has a weakly dominated pure strategy is not established, and the resulting claim 26 is not a valid consequence of the argument (Avramopoulos, 2016).
Across these settings, the principal distinction is between exact matrix-valued multiplicative weights, low-rank or payoff-based approximations, and scalar multiplicative-weights reductions available under additional structure. Standard MMW accommodates arbitrary symmetric, including noncommuting, matrix sequences. Rank-one sketches preserve regret through a different averaged mirror map. Oja’s algorithm reduces to ordinary MW only under a common eigenbasis. Quantum MMW extends the geometry to density matrices and yields conservation and recurrence phenomena in zero-sum games. Evolutionary dynamics reproduce MMW exactly at the level of allele marginals while retaining correlations in the underlying joint population.