---
title: Matrix Multiplicative Weights Algorithm
url: https://www.emergentmind.com/topics/matrix-multiplicative-weights
type: topic
---

# Matrix Multiplicative Weights Algorithm

Matrix multiplicative weights (MMW) is a matrix-valued generalization of multiplicative-weights and exponentiated-gradient methods in which probability distributions are replaced by positive semidefinite trace-one matrices, payoff vectors by Hermitian matrices, and scalar normalization by trace normalization. The canonical decision domain is the spectrahedron
\[
\mathcal A_n=\{X\in\mathbb S^n:X\succeq 0,\ \operatorname{tr}(X)=1\},
\]
whose elements are density matrices. MMW is used for online optimization over semidefinite domains, zero-sum and quantum games, semidefinite programming, online eigenvector problems, and quadratic optimization. Its central update maps an accumulated symmetric gain matrix \(Y\) to
\[
P_{\mathrm{mw}}(Y)=\frac{e^Y}{\operatorname{tr}(e^Y)}.
\]

## 1. Mathematical formulation and canonical update

At round \(t\), an MMW learner selects \(X_t\in\mathcal A_n\), receives a symmetric gain matrix \(G_t\in\mathbb S^n\), and obtains payoff
\[
\langle G_t,X_t\rangle=\operatorname{tr}(G_t^\top X_t).
\]
The regret against the best fixed matrix in \(\mathcal A_n\) is
\[
\max_{X\in\mathcal A_n}\sum_{t=1}^T\langle G_t,X\rangle
-\sum_{t=1}^T\langle G_t,X_t\rangle.
\]
Because
\[
\max_{X\in\mathcal A_n}\langle Y,X\rangle=\lambda_{\max}(Y),
\]
the comparator term equals
\[
\lambda_{\max}\left(\sum_{t=1}^T G_t\right).
\]

The standard matrix exponentiated-gradient update selects
\[
X_t=P_{\mathrm{mw}}\left(\eta\sum_{i=1}^{t-1}G_i\right),
\]
where \(\eta>0\) is the learning rate. If
\[
Y=Q\operatorname{diag}(\lambda_1,\ldots,\lambda_n)Q^\top,
\]
then
\[
P_{\mathrm{mw}}(Y)
=
Q\operatorname{diag}\left(
\frac{e^{\lambda_1}}{\sum_j e^{\lambda_j}},
\ldots,
\frac{e^{\lambda_n}}{\sum_j e^{\lambda_j}}
\right)Q^\top.
\]
The matrix exponential ensures positive semidefiniteness, while trace normalization ensures unit trace.

The associated potential is
\[
p_{\mathrm{mw}}(Y)=\log\operatorname{tr}(e^Y),
\]
with gradient
\[
\nabla p_{\mathrm{mw}}(Y)=P_{\mathrm{mw}}(Y).
\]
Its convex conjugate is the negative von Neumann entropy, represented on positive definite \(X\in\mathcal A_n\) by
\[
F(X)=\operatorname{tr}(X\log X).
\]
The induced Bregman divergence is quantum relative entropy,
\[
D_F(X\|Z)
=
\operatorname{tr}(X\log X)-\operatorname{tr}(X\log Z).
\]

When \(\|G_t\|_{\mathrm{op}}\leq 1\), standard MMW gives average regret of order
\[
O\left(\sqrt{\frac{\log n}{T}}\right),
\]
and cumulative regret of order
\[
O(\sqrt{T\log n}).
\]
A choice such as
\[
\eta\asymp\sqrt{\frac{\log n}{T}}
\]
produces this scale. The \(\log n\) term arises from the entropy diameter of the spectrahedron.

MMW is distinct from a self-coupled Hedge map used in some analyses of symmetric games. In that setting,
\[
T_i(X)=
\frac{X(i)e^{\alpha(CX)_i}}
{\sum_jX(j)e^{\alpha(CX)_j}},
\]
where the same strategy \(X\) supplies both the population distribution and the opponent against which payoffs are evaluated. This is a nonlinear single-population dynamical system, not the usual two-process row-player/column-player matrix MMW algorithm. Its fixed points and KL-divergence identities do not by themselves establish convergence of empirical play to Nash equilibrium [1609.08934].

## 2. Mirror descent, entropy, and regret analysis

The geometry of MMW is generated by the von Neumann entropy. For density matrices \(X\) and \(Y\), the quantum relative entropy is
\[
D(X\|Y)=\operatorname{tr}\left(X(\log X-\log Y)\right).
\]
In payoff-based analyses, the entropy satisfies the semidefinite Pinsker-type inequality
\[
D(X,Y)\geq \frac12\|X-Y\|_F^2.
\]
This strong convexity supplies the stability term in the mirror-descent analysis.

For an MMW update driven by \(G_t\), the basic energy inequality is
\[
D(Z,X_{t+1})
\leq
D(Z,X_t)
+\eta_t\operatorname{tr}\left(G_t(X_t-Z)\right)
+\frac{\eta_t^2}{2}\|G_t\|_F^2.
\]
Summing over rounds yields a matrix regret estimate of the form
\[
\sum_{t=1}^T\operatorname{tr}\left(G_t(X_t-Z)\right)
\lesssim
\frac{D(Z,X_1)}{\eta}
+\frac{\eta}{2}\sum_{t=1}^T\|G_t\|_F^2.
\]
For one player, the entropy diameter can be bounded by \(\log d\); in a two-player setting the corresponding quantity is
\[
H_{\max}=\log(d_1d_2).
\]

In two-player zero-sum semidefinite games, the averaged iterates
\[
\bar X_T=\frac1T\sum_{t=1}^T X_t
\]
satisfy a duality-gap estimate
\[
\operatorname{Gap}_L(\bar X_T)
\leq
L\sqrt{\frac{2H_{\max}}{T}},
\]
where \(L\) is a Lipschitz or gradient bound. Thus \(T=O(\varepsilon^{-2})\) iterations suffice for an \(\varepsilon\)-Nash equilibrium under full matrix-gradient feedback.

Relative entropy also appears in continuous-time and quantum formulations. In a two-player zero-sum quantum game with a fully mixed Nash equilibrium \((\rho^*,\sigma^*)\), quantum replicator dynamics preserve
\[
S(\rho^*\|\rho(t))+S(\sigma^*\|\sigma(t)),
\]
where
\[
S(\rho\|\sigma)=\operatorname{Tr}\bigl(\rho(\log\rho-\log\sigma)\bigr).
\]
The conservation law is a continuous-time analogue of the conserved divergence in classical zero-sum replicator dynamics. It constrains the joint evolution of the two density matrices but does not imply that either state has a fixed spectrum or that the trajectory converges to equilibrium [2211.01681].

## 3. Computational approximations and low-rank sketches

The dominant computational obstacle in standard MMW is the matrix exponential. General-purpose eigendecomposition costs approximately \(O(n^3)\) in practice, while full matrix exponentiation requires maintaining an \(n\times n\) matrix.

A rank-one randomized sketch replaces the full matrix softmax by
\[
P_u(Y)
=
\frac{e^{Y/2}uu^\top e^{Y/2}}
{u^\top e^Yu}
=
\frac{vv^\top}{v^\top v},
\qquad
v=e^{Y/2}u,
\]
where \(u\) is sampled uniformly from the unit sphere. The resulting matrix is positive semidefinite, rank one, and has unit trace. The sketched algorithm plays
\[
X_t=P_{u_t}\left(\eta\sum_{i=1}^{t-1}G_i\right).
\]

This sketch is not unbiased for the standard MMW projection:
\[
\mathbb E_u[P_u(Y)]\neq P_{\mathrm{mw}}(Y)
\]
in general. Instead, define
\[
P(Y)=\mathbb E_u[P_u(Y)].
\]
The averaged projection is the gradient of the convex potential
\[
p(Y)=\mathbb E_u\log(u^\top e^Yu),
\qquad
\nabla p(Y)=P(Y).
\]
Its convex conjugate supplies the regularizer for the averaged algorithm.

The relevant Bregman divergence satisfies the smoothness bound
\[
V_Y(Y+D)\leq \frac32\|D\|_{\mathrm{op}}^2
\]
and the diameter bound
\[
V_Y(0)-V_Y(Y')\leq \log(4n).
\]
These properties allow mirror-descent analysis with only constant-factor degradation relative to standard MMW. The resulting cumulative regret is bounded by
\[
\lambda_{\max}\left(\sum_{t=1}^TG_t\right)
-\sum_{t=1}^T\langle G_t,X_t\rangle
\leq
\frac{\log(4n)}{\eta}
+\frac{3\eta}{2}\sum_{t=1}^T\|G_t\|_{\mathrm{op}}^2.
\]
For \(\|G_t\|_{\mathrm{op}}\leq 1\), this gives average regret
\[
O\left(\sqrt{\frac{\log n}{T}}\right).
\]

The actual random sketch requires an adversary condition: conditional on the past, \(G_t\) must be independent of the fresh random vector \(u_t\). Under this condition,
\[
\mathbb E[P_{u_t}(Y_t)\mid\mathcal F_{t-1}]=P(Y_t),
\]
so expected regret is controlled by the deterministic averaged-projection analysis. High-probability bounds follow by treating
\[
\langle G_t,P_{u_t}(Y_t)-P(Y_t)\rangle
\]
as a martingale difference.

The matrix-exponential–vector product \(e^{Y/2}u\) can be approximated using \(k\) Lanczos iterations. If
\[
T_k=V\Lambda V^\top
\]
is the Lanczos tridiagonalization, then
\[
\operatorname{exp}_k(A,b)
=
\|b\|_2Q_kV e^\Lambda V^\top e_1
\]
approximates \(e^Ab\). The resulting computational cost is
\[
O\bigl(\operatorname{mv}(A)\,k+k^2B\bigr),
\]
where \(\operatorname{mv}(A)\) is the matrix-vector multiplication cost and \(B\) is the floating-point word size. The rank-one method is advantageous when matrix-vector multiplication is substantially cheaper than dense eigendecomposition, particularly for sparse or structured matrices [1903.02675].

## 4. Applications to eigenvectors, semidefinite programming, and quadratic optimization

The spectrahedral domain contains rank-one matrices \(xx^\top\), so MMW directly induces algorithms for online eigenvector problems. The benchmark is
\[
\max_{\|x\|_2=1}\sum_{t=1}^T x^\top G_tx
=
\lambda_{\max}\left(\sum_{t=1}^TG_t\right).
\]
The rank-one method selects
\[
x_t=
\frac{
e^{\frac{\eta}{2}\sum_{i<t}G_i}u_t
}{
\left\|e^{\frac{\eta}{2}\sum_{i<t}G_i}u_t\right\|_2
},
\qquad
X_t=x_tx_t^\top.
\]
With exact exponential-vector products, the expected regret is
\[
O(\sqrt{T\log n}).
\]
With Lanczos approximation, the total matrix-vector-product complexity for high-probability \(\varepsilon\)-average regret is
\[
O\left(\varepsilon^{-2.5}\log^{2.5}(n/\delta)\right),
\]
and the method improves the previous best online-eigenvector complexity by a factor of at least \(\Omega(\log^5 n)\).

MMW and Oja’s algorithm are related but not equivalent in general. If the input matrices \(A_t\) share a common orthonormal eigenbasis \(v_1,\ldots,v_n\), with
\[
A_tv_i=\lambda_t(i)v_i,
\]
then the squared eigen-coordinates of Oja’s normalized vector,
\[
\psi_t(i)=(v_i^\top\widehat w_t)^2,
\]
satisfy
\[
\psi_{t+1}(i)=\psi_t(i)(1+\mu\lambda_t(i))^2.
\]
After rewriting this expression, \(\psi_t\) follows an ordinary multiplicative-weights recursion. Thus Oja’s vector dynamics reduce exactly to scalar MW over eigen-directions in the common-eigenbasis regime. For arbitrary noncommuting matrices, the squared-coordinate dynamics do not close, and this reduction fails [2310.15559].

MMW also applies to semidefinite programs in saddle-point form. Let
\[
\mathcal A_n=\{X\succeq 0:\operatorname{Tr}(X)=1\}
\]
and
\[
\Delta_m=\{y\in\mathbb R^m:y\geq 0,\ \mathbf 1^\top y=1\}.
\]
For symmetric matrices \(A_1,\ldots,A_m\), define
\[
A^*y=\sum_{i=1}^m y_iA_i.
\]
The matrix player uses the rank-one sketch, while the simplex player uses ordinary multiplicative weights. The duality gap is
\[
\operatorname{Gap}(X,y)
=
\lambda_{\max}(A^*y)-\min_{i\in[m]}\langle A_i,X\rangle.
\]
If
\[
w=\max_i\|A_i\|_{\mathrm{op}},
\]
then the averaged strategies satisfy
\[
\mathbb E[\operatorname{Gap}(X_{\mathrm{avg}},y_{\mathrm{avg}})]
=
O\left(w\sqrt{\frac{\log(mn)}{T}}\right).
\]
To achieve expected gap at most \(\varepsilon\), it suffices to take
\[
T=O\left(\frac{w^2\log(4mn)}{\varepsilon^2}\right).
\]

## 5. Quantum games and payoff-based MMW

In quantum games, a mixed strategy is a density matrix
\[
D(\mathcal H)=\{\rho:\rho\succeq 0,\ \operatorname{Tr}\rho=1\}.
\]
For a two-player game, Alice and Bob use density matrices \(\rho\) and \(\sigma\). A payoff observable \(R\) induces the bilinear payoff
\[
u(\rho,\sigma)
=
\operatorname{Tr}(R(\rho\otimes\sigma))
=
\operatorname{Tr}(\rho\,\Phi(\sigma))
=
\operatorname{Tr}(\sigma\,\Phi^\dagger(\rho)).
\]

The matrix multiplicative-weights updates are
\[
A_j=\exp\left(\mu\sum_{i=0}^{j-1}\Phi(\sigma_i)\right),
\qquad
\rho_j=\frac{A_j}{\operatorname{Tr}(A_j)},
\]
and
\[
B_j=\exp\left(-\mu\sum_{i=0}^{j-1}\Phi^\dagger(\rho_i)\right),
\qquad
\sigma_j=\frac{B_j}{\operatorname{Tr}(B_j)}.
\]
The negative sign in Bob’s update reflects minimization of Alice’s payoff. When all matrices are diagonal in a common basis, MMWU reduces exactly to classical exponential MWU. Noncommutativity is the principal quantum distinction: accumulated payoff operators need not commute with their instantaneous derivatives, and the derivative of \(e^{P(t)}\) is generally
\[
\frac{d}{dt}e^{P(t)}
=
\int_0^1 e^{(1-s)P(t)}\dot P(t)e^{sP(t)}\,ds.
\]

The continuous limit defines quantum replicator dynamics through
\[
\rho(t)=\frac{e^{A(t)}}{\operatorname{Tr}(e^{A(t)})},
\qquad
A(t)=\int_0^t\Phi(\sigma(\tau))\,d\tau,
\]
and
\[
\sigma(t)=\frac{e^{B(t)}}{\operatorname{Tr}(e^{B(t)})},
\qquad
B(t)=-\int_0^t\Phi^\dagger(\rho(\tau))\,d\tau.
\]
For a fully mixed Nash equilibrium, the total quantum relative entropy is conserved:
\[
\frac{d}{dt}
\left[
S(\rho^*\|\rho(t))+S(\sigma^*\|\sigma(t))
\right]
=0.
\]
The same setting yields Poincaré recurrence for almost every interior initial condition. The proof uses canonical coordinates, volume preservation, boundedness from the relative-entropy invariant, and the Poincaré recurrence theorem. The result establishes recurrence rather than convergence: generic trajectories return arbitrarily close to their initial conditions infinitely often.

Full-information MMW assumes access to the complete payoff-gradient matrix. Payoff-based learning replaces it with an estimator. In minimal-information matrix multiplicative weights, or 3MW, the update is
\[
Y_{i,t+1}=Y_{i,t}+\eta_t\widehat V_{i,t},
\qquad
X_{i,t+1}
=
\frac{\exp(Y_{i,t+1})}
{\operatorname{tr}(\exp(Y_{i,t+1}))}.
\]
The estimator decomposes as
\[
\widehat V_t=V(X_t)+b_t+\xi_t,
\]
where \(b_t\) is smoothing bias and \(\xi_t\) is a martingale-noise term.

For deterministic scalar payoff feedback, a two-point estimator achieves
\[
\mathbb E[\operatorname{Gap}_L(\bar X_T)]
=
O(T^{-1/2}),
\]
with a factor linear in the effective matrix dimension \(D=\max_i(d_i^2-1)\). For a single random payoff-observable realization, a one-point estimator has variance of order \(1/\delta_t^2\), leading to
\[
\mathbb E[\operatorname{Gap}_L(\bar X_T)]
=
O(T^{-1/4}).
\]
The rate difference arises from the bias–variance tradeoff: the smoothing bias is \(O(\delta)\), whereas the one-point estimator variance is \(O(\delta^{-2})\). In general non-zero-sum games, a regularized 3MW method has local, high-probability, last-iterate convergence to equilibria satisfying variational stability [2311.02423].

## 6. Evolutionary, game-theoretic, and dynamical interpretations

MMW also arises exactly in the marginal allele dynamics of sexual evolutionary models. For selection before recombination, let \(P^t=(p^t_{ij})\) be the joint population distribution and let
\[
x_i^t=\sum_jp^t_{ij}
\]
be the allele marginal. If
\[
g_i^t=\sum_jP^t(b_j\mid a_i)w_{ij}
\]
is the conditional fitness of allele \(a_i\), then the evolutionary update satisfies
\[
x_i^{t+1}
=
\frac{x_i^tg_i^t}{w^t},
\]
which is exactly the parameter-free, correlation-sensitive polynomial-weights update. This identity holds for arbitrary nonnegative fitness matrices, recombination rates, initial distributions, and numbers of loci.

For recombination before selection, the marginal update uses a mixture of correlation-sensitive and independent payoffs:
\[
g_i^{t,(r)}
=
r\bar g_i^t+(1-r)g_i^t,
\]
where
\[
\bar g_i^t=\sum_jy_j^tw_{ij}.
\]
The resulting update is
\[
x_i^{t+1}
=
\frac{x_i^tg_i^{t,(r)}}{w^R}.
\]
Thus recombination before selection corresponds to an interpolated parameter-free PW process.

The correspondence is between allele marginals and MW strategies, not between the complete joint population distribution and independent mixed strategies. A product distribution satisfies
\[
p_{ij}^t=x_i^ty_j^t,
\]
whereas linkage disequilibrium is
\[
D_{ij}^t=p_{ij}^t-x_i^ty_j^t.
\]
Product distributions are generally not preserved, even under weak selection. Consequently, the uncorrelated MWUA based only on marginal distributions need not describe actual evolutionary dynamics, and it can converge to a different equilibrium from the sexual population process. The precise correspondence therefore retains the evolving correlations through conditional payoffs [1502.05056].

In symmetric bimatrix games, the Hedge map
\[
T_i(X)=
\frac{X(i)e^{\alpha(CX)_i}}
{\sum_jX(j)e^{\alpha(CX)_j}}
\]
has fixed points characterized by equal payoffs on the support of \(X\). Interior fixed points are interior symmetric equilibria, while boundary fixed points need not be equilibria of the full game. The map admits a relative-entropy identity
\[
RE(Y,T_\alpha(X))-RE(Y,X)
=
\log\left(
\sum_jX(j)e^{\alpha((CX)_j-Y\cdot CX)}
\right),
\]
whose second derivative is an exponentially tilted payoff variance and is therefore nonnegative. These are valid local and variational properties of the Hedge map, but they do not imply convergence of the dynamics to equilibrium or establish a polynomial-time equilibrium algorithm. In particular, the claimed implication that every symmetric game without an equalizer has a weakly dominated pure strategy is not established, and the resulting claim \(\mathbf P=\mathbf{PPAD}\) is not a valid consequence of the argument [1609.08934].

Across these settings, the principal distinction is between exact matrix-valued multiplicative weights, low-rank or payoff-based approximations, and scalar multiplicative-weights reductions available under additional structure. Standard MMW accommodates arbitrary symmetric, including noncommuting, matrix sequences. Rank-one sketches preserve regret through a different averaged mirror map. Oja’s algorithm reduces to ordinary MW only under a common eigenbasis. Quantum MMW extends the geometry to density matrices and yields conservation and recurrence phenomena in zero-sum games. Evolutionary dynamics reproduce MMW exactly at the level of allele marginals while retaining correlations in the underlying joint population.

Source: https://www.emergentmind.com/topics/matrix-multiplicative-weights