---
title: Full-History Bradley-Terry (FH-BT)
url: https://www.emergentmind.com/topics/full-history-bradley-terry-fh-bt
type: topic
---

# Full-History Bradley-Terry (FH-BT)

The Full-History Bradley-Terry (FH-BT) framework is a generalization of the classical Bradley-Terry model for paired comparison, designed to unify the strengths of explanatory statistical modeling and predictive, supervised learning for outcomes such as competitive sports results. FH-BT introduces maximum-likelihood learning over the entire history of pairwise outcomes, supports batch and online (Élő-style) updates, integrates covariates and features in a generalized linear model (GLM) framework, and extends naturally to incorporate draws, low-rank structural constraints, and matrix completion interpretations. Its validity and efficacy have been experimentally confirmed on synthetic data and English Premier League outcomes, demonstrating state-of-the-art predictive performance [1701.08055].

## 1. Relationship to Classical Bradley-Terry and Élő Systems

The FH-BT model extends the paradigms of both the classical Bradley-Terry (BT) model and the Élő rating system. In the BT model, each contestant or team $i$ is described by a latent "strength" parameter $\theta_i$. The probability that team $i$ defeats team $j$ is modeled as 
$$
p_{ij} = \sigma(\theta_i - \theta_j + h)
$$
where $\sigma(x) = (1 + e^{-x})^{-1}$ is the logistic sigmoid and $h$ is an intercept accounting for home advantage. However, while the BT model is typically used for explanatory analysis, the Élő rating system is a widely used heuristic baseline for online predictive updating.

FH-BT exploits the mathematical closeness between these models, as well as similarities to logistic regression, to create a unified statistical approach that supports supervised probabilistic prediction and efficient online updates. Unlike the original Élő algorithm—which is not a likelihood-based statistical model—FH-BT provides a principled probabilistic interpretation and a direct connection to GLM and modern matrix completion schemes [1701.08055].

## 2. Maximum Likelihood Fitting Over Full History

The core of FH-BT is the batch ("full-history") maximum-likelihood estimation of the latent strengths $\theta$. For a historical dataset $D = \{(i_k, j_k, Y^{(k)})\}_{k=1}^N$, the model optimizes
$$
\ell(\theta|D) = \sum_{k=1}^N \left[ Y_{i_k j_k}^{(k)} \log \sigma(\theta_{i_k} - \theta_{j_k} + h) + (1 - Y_{i_k j_k}^{(k)}) \log (1 - \sigma(\theta_{i_k} - \theta_{j_k} + h)) \right]
$$
This formulation allows one to compute gradients analytically and apply gradient ascent (or minorization–maximization) to obtain a global optimum $\hat\theta^*$, constituting the "full-history" fit. The batch fitting procedure ensures that latent strengths are informed by the entire sequence of observed pairwise outcomes and not just the most recent matches or updates [1701.08055].

## 3. Online Updating and the K-Factor

FH-BT supports online learning via stochastic gradient ascent. Upon observing a new outcome $(i, j, Y_{ij})$, a single-step update with step-size $K$ ("K-factor") is given by
$$
\theta_i \gets \theta_i + K (Y_{ij} - p_{ij}),\quad \theta_j \gets \theta_j - K (Y_{ij} - p_{ij})
$$
where $p_{ij}$ is the current predicted probability of $i$ defeating $j$. This update mechanism is exactly stochastic gradient ascent on the FH-BT likelihood; the one-sample gradient is $Y_{ij} - p_{ij}$. FH-BT thus generalizes the Élő update, marrying it with the statistical grounding of maximum-likelihood estimation, and enables flexible adaptation as new data arrives [1701.08055].

## 4. Model Extensions: Draws, Features, and Low-Rank Structure

### Draws: Proportional Odds Extension

For outcomes comprising win, draw, and loss (i.e., $Y_{ij} \in \{\text{win},\text{draw},\text{lose}\}$), FH-BT introduces a proportional-odds structure with a shared draw parameter $\phi$:
\[
\begin{align*}
\logit\, P(\text{win}) &= L_{ij}, \\
\logit\, P(\text{win or draw}) &= L_{ij} + \phi \\
P(\text{win}) &= \sigma(L_{ij}), \\
P(\text{draw}) &= \sigma(-L_{ij}) - \sigma(-L_{ij} - \phi), \\
P(\text{lose}) &= \sigma(-L_{ij} - \phi)
\end{align*}
\]
where $L_{ij} = \theta_i - \theta_j + h$.

### Incorporation of Features

FH-BT naturally accommodates additional information by generalizing $L_{ij}$:
\[
L_{ij} = \theta_i - \theta_j + h + \langle\lambda, X_{ij}\rangle
\]
or, more generally,
\[
L_{ij} = \langle\beta, X_{ij}\rangle + \langle\gamma, X_i\rangle + \langle\delta, X_j\rangle + \alpha_{ij}
\]
Structural constraints such as low-rankness, anti-symmetry of $\alpha$, and shared intercepts are used for parsimony and tractability [1701.08055].

### Low-Rank Matrix Structure

The matrix of pairwise log-odds $L \in \mathbb{R}^{Q \times Q}$ is anti-symmetric ($L = -L^\top$) and, in the classical model, of rank two. Extensions permit factorizations such as
\[
L = u v^\top - v u^\top \quad (\text{rank-2 and full two-factor BT}), \\
L = u v^\top - v u^\top + \theta 1^\top - 1 \theta^\top \quad (\text{rank-4 extension})
\]
and nuclear norm regularization for matrix completion perspectives [1701.08055].

## 5. Empirical Validation and Predictive Performance

Empirical evaluation of FH-BT has been conducted on both synthetic and real-world datasets. On synthetic data generated from rank-2 or rank-4 log-odds, FH-BT extensions with appropriately matched rank systematically outperform the rank-2 (classical) BT model in accuracy and log-likelihood (Wilcoxon $p \ll 10^{-4}$). In English Premier League data, a two-stage strategy comprising initial full-history batch fitting followed by K-factor updates outperforms pure online or naive batch re-fit approaches (paired t-test $p \ll 10^{-4}$). Incorporating a single covariate for team promotion significantly improves log-loss over the vanilla FH-BT ($p \approx 0.002$).

The best FH-BT variant achieves approximately 52.8% win/draw/lose accuracy (95% CI [50.6%, 55.0%]) and mean log-loss of approximately $-0.980$, comparing closely to performance based on Bet365 betting odds (54.1%/$-0.967$) [1701.08055].

## 6. Significance and Interpretive Summary

The Full-History Bradley-Terry model provides a statistically principled, flexible, and computationally efficient GLM-style probabilistic predictor for pairwise outcomes. Its latent strengths are learned from the entire match history, supports efficient online adaptations, generalizes naturally to multi-outcome and covariate settings, and offers more expressive structural modelling through low-rank embeddings. This synthesis of explanatory and predictive modeling yields highly competitive results and validates the importance of full-history batch fitting augmented by lightweight online adaptation for practical prediction tasks such as sports outcome modeling [1701.08055].

Source: https://www.emergentmind.com/topics/full-history-bradley-terry-fh-bt