---
title: Derivation of Shannon Entropy
url: https://www.emergentmind.com/topics/shannon-entropy-derivation
type: topic
---

# Derivation of Shannon Entropy

Shannon entropy quantitatively characterizes the expected uncertainty or average informational content associated with a probability distribution over discrete outcomes. Its derivation—foundational to information theory and statistical mechanics—can be constructed rigorously from combinatorics, variational principles, axiomatic arguments, and algebraic frameworks. The uniqueness and universality of the Shannon entropy formula are anchored in symmetry, recursivity, and the asymptotic behavior of large statistical ensembles.

## 1. Combinatorial Derivation via Multiplicities

Consider a system of $N$ independent trials, each producing one of $W$ distinct outcomes labeled $i=1,\ldots,W$. Let $n_i$ denote the count of outcome $i$ in the sequence, constrained by $\sum_{i=1}^W n_i = N$. The central combinatorial object is the number of distinct micro-configurations (sequences) compatible with the macro-state $\{n_i\}$, given by the multinomial coefficient:
$$
M(\{n_i\}) = \frac{N!}{n_1! \, n_2! \cdots n_W!}
$$
This multiplicity measures the volume in configuration space associated with the specified occupation numbers.

The empirical probability of state $i$ is then $p_i = n_i / N$. In the limit $N \to \infty$, typical fluctuations vanish, and the probability of observing $\{n_i\}$ concentrates near values maximizing $P(\{n_i\} | \boldsymbol{\theta})$ with the corresponding multiplicity.

Applying Stirling’s approximation to the factorials:
$$
\ln N! \approx N \ln N - N; \quad \ln n_i! \approx n_i \ln n_i - n_i
$$
the logarithm of the multiplicity reduces to:
$$
\ln M(\{n_i\}) \approx -N \sum_{i=1}^W p_i \ln p_i
$$
Boltzmann postulated that macroscopic entropy should be proportional to the log-multiplicity:
$$
S = k \ln M(\{n_i\}) \approx -k N \sum_{i=1}^W p_i \ln p_i
$$
with $k$ a positive constant (in physics, $k_B$). This leads to the entropy per trial:
$$
s = \frac{S}{N} = -k \sum_{i=1}^W p_i \ln p_i
$$
the Boltzmann–Gibbs–Shannon form [1404.5650, 1504.01407].

## 2. Axiomatic Derivation: Shannon–Khinchin Framework

The uniqueness of the Shannon entropy formula is established by the Shannon–Khinchin (SK) or similar sets of axioms:
- (SK1) **Continuity**: $H$ depends continuously on $\{p_i\}$.
- (SK2) **Maximality**: $H$ is maximal for the uniform distribution, i.e., $p_i=1/W$.
- (SK3) **Expansibility (Null State)**: Adding a state of zero probability leaves $H$ unchanged.
- (SK4) **Recursivity (Additivity/Chain Rule)**: For composite systems, entropy is additive:
  $$
  H(A \cup B) = H(A) + \sum_i p_i(A) H(B|A=i)
  $$
Under these axioms, the only solution (up to a positive multiplicative constant) is:
$$
H(\{p_i\}) = -k \sum_{i=1}^W p_i \ln p_i
$$
This result is achieved independently in classical information theory, combinatorial models, and statistical physics [1404.5650, 1504.01407, 1209.5500].

## 3. Variational and Maximum Entropy Principle Approach

A variational derivation leverages constrained ignorance. Given a probability distribution $p=(p_1,\ldots,p_n)$ constrained by normalization and possibly other functional constraints (such as moments), define a Lagrangian:
$$
\mathcal{L}[p, \lambda] = H[p] - \lambda \left(\sum_{i} p_i - 1\right)
$$
Stationarity under arbitrary $\delta p_i$ subject to normalization yields a functional equation for $H$ that, together with the chain rule property for conditional probabilities, restricts $H$ to the Shannon form. Additional constraints (e.g., fixed mean energy) produce the Gibbs–Boltzmann distribution as the entropy-maximizing solution [2107.05008].

Two derivation strategies arise:
- **Biased Ansatz**: Assume $H[p] = \text{const} + \sum_i p_i h(p_i)$; the chain rule compels $h(p) = -k \ln p$.
- **General Axiomatic Route**: Functional equations from additivity and symmetry directly yield $H[p] = -k \sum_i p_i \ln p_i$.

Both methods converge to the same entropy form and underpin the Maximum Entropy Principle (MaxEnt) for statistical inference [2107.05008].

## 4. Algebraic and Operadic Characterizations

Shannon entropy can be formulated in the context of algebraic structures such as operads. The standard simplex $\Delta_n$ of probability distributions is equipped with partial compositions modeling sequential random processes. A derivation $d_n$ on this operad defined by:
$$
d_n(p)(x) = H(p) = -\sum_{i=1}^n p_i \log p_i
$$
satisfies the Leibniz rule:
$$
d(p \circ_i q) = d(p) \circ_i^R q + p \circ_i^L d(q)
$$
It can be shown that any continuous derivation in this setting is a constant multiple of $H$, as characterized by the Faddeev–Leinster theorem. This algebraic formalism encapsulates the chain rule or grouping property in a categorical framework, further establishing the uniqueness of the Shannon entropy in probabilistic compositional systems [2107.09581].

## 5. Relative Divergence, Grading Functions, and Generalized Contexts

Mathematical entropy arises in the structure of comparing grading functions on linearly ordered sets. Consider grading functions $F,G: W \to \mathbb{R}$ on a totally ordered set $W$; the local divergence between $F$ and $G$ is induced by a logarithmic rate:
$$
h(s,t) = \ln (s/t)
$$
The global divergence over the chain $W$ is:
$$
D(F \| G) = \sum_k \Delta F_k \ln \left(\frac{\Delta G_k}{\Delta F_k}\right)
$$
Specializing $F$ to the cumulative distribution of a probability mass function and $G$ as the position grading, yields the standard Shannon entropy:
$$
-\sum_{k=1}^n p_k \ln p_k
$$
This demonstrates that Shannon entropy is a particular instance of a general divergence measure constrained by smoothness, invariance, and additivity [1903.05240].

## 6. Special Considerations and Extensions

### Internal Entropy and Statistical Mechanics Corrections

In statistical mechanics, when microstates $i$ possess further degeneracy (internal entropy $S_i = k_B \ln w_i$ for multiplicity $w_i$), the total entropy functional becomes:
$$
S_{\text{total}}[p] = -k_B \sum_i p_i \ln p_i + \sum_i p_i S_i
$$
Only when all $S_i$ are equal (or zero by convention) does the Shannon expression suffice. Otherwise, additional terms are required to account for the physical entropy content, especially in non-identically weighted microstates [1209.5500].

### Finite Sample Correction

For finite sample size $N$, a corrected entropy accounts for finite combinatorial freedom:
$$
H_0(\{p_i\}; N) = \frac{1}{N}\left[ \ln \Gamma(N+1) - \sum_i \ln \Gamma(N p_i + 1)\right]
$$
In the $N \to \infty$ limit, this expression reduces to the Shannon entropy. For small $N$, the correction quantifies reduced information per event due to the limited sample size and sets bounds on maximal achievable channel utilization [1504.01407].

### Black Hole Entropy and Information-theoretic Analysis

Shannon entropy applied to the tunneling probability of quantum fields escaping black hole event horizons yields the Bekenstein-Hawking entropy law. The cumulative information loss—expressed as the sum of Shannon information over radiated modes—reproduces the gravitational entropy-area relation, affirming the informational basis of black hole thermodynamics [1008.0946].

## 7. Synthesis and Universality

The derivation of Shannon entropy is robust under distinct mathematical disciplines: combinatorial enumeration, axiomatic characterizations, variational calculus, algebraic operads, and divergence measures. Its formula,
$$
H(\{p_i\}) = -k \sum_{i} p_i \ln p_i
$$
is enforced by core properties—continuity, maximality under equiprobability, and compositional additivity—which are essential for any legitimate quantifier of information or uncertainty. Its appearance across statistical mechanics, information theory, quantum field theory, and categorical algebra underscores its universality and rigidity as a fundamental tool in the quantification of probabilistic ignorance and disorder [1404.5650, 1504.01407, 1209.5500, 1903.05240, 2107.09581, 2107.05008, 1008.0946].

Source: https://www.emergentmind.com/topics/shannon-entropy-derivation