Result 119, Theoretical computer science

The Courtade–Kumar and Hellinger conjectures

Proves the Courtade–Kumar conjecture: among Boolean functions of independent uniform bits, a single coordinate retains the most mutual information after independent bit-flip noise. A stronger theorem treats randomized binary summaries at fixed initial information. The Hellinger conjecture is also proved for every Boolean output bias and noise correlation.

Lean formalization Proof

The bigger picture

Why it matters

Can a clever one-bit summary of many random bits survive noise better than simply keeping one bit? These manuscripts claim the answer is no, establishing sharp limits on information retained by binary summaries.

What changes?

For any number of independent uniform bits under independent symmetric bit-flip noise, the manuscript reports that keeping one coordinate maximizes mutual information between a Boolean summary and the noisy input. A Boolean summary outputs one bit; mutual information measures how much one variable reveals about another. At fixed initial information, the stronger claim makes a randomized coordinate optimal among randomized binary summaries. The Hellinger claim covers every output bias, or imbalance, and noise correlation from minus one to one.

What does that help mathematicians do?

The claimed Hellinger loss is bounded by one minus the square root of one minus the squared noise correlation. Equality holds for coordinates and their sign reversals. This sharp inequality also yields the Courtade-Kumar mutual-information bound, connecting two measures of information. Researchers could therefore rule out any Boolean construction combining many bits that beats a single coordinate in mutual information under these assumptions, regardless of dimension.

Are there practical applications?

The immediate value is foundational: an exact benchmark for studying one-bit compression under symmetric noise, including randomized summaries with a specified initial information budget. This connects directly to theoretical information processing. The supplied abstracts do not establish comparable optimality for correlated inputs, nonuniform bits, other noise models, or summaries longer than one bit.

This section was generated by GPT-6 Astra Medium. This explanation is based on the result summary and manuscript abstracts below. This context is separate from OpenAI's source text.

2 manuscripts

Sharp binary-information contraction on the discrete cube

September 24, 2026 29 pages Main result formalized in Lean

We prove sharp contraction of the information carried by a binary channel under independent symmetric noise on a uniform discrete cube. At fixed initial information, a noisy coordinate retains the most information. The Boolean specialization resolves the Courtade–Kumar conjecture and gives an output-entropy refinement. We also establish a stronger mean-dependent entropy-production bound. The proof combines an explicit three-point optimizer for the local joining problem, two entropy capacities, and a common-output thinning inequality, followed by dimension induction and integration along the noise semigroup.

Cite (BibTeX)
@misc{OAI:Sharp-binary-information-contraction-on-the-discrete-cube-September-24-2026,
  author = {{OpenAI}},
  title = {{Sharp binary-information contraction on the discrete cube}},
  howpublished = {OpenAI Math Release preprint
                  \href{https://github.com/openai/math/blob/main/preprints/Sharp-binary-information-contraction-on-the-discrete-cube-September-24-2026/main.pdf}{OAI:Sharp-binary-information-contraction-on-the-discrete-cube-September-24-2026}},
  year = {2026}
}

Hellinger contraction with arbitrary Boolean output bias

September 24, 2026 50 pages

We prove the Hellinger conjecture for Boolean functions on the uniform discrete cube, with arbitrary output bias. For a Boolean function of mean m and every ρ∈[−1,1]\rho\in[-1,1], the loss 1−m2−E1−(Tρf)2\sqrt{1-m^2}-\mathbb E\sqrt{1-(T_\rho f)^2} is at most 1−1−ρ21-\sqrt{1-\rho^2}, with equality for signed coordinates. The proof combines asymmetric dimension induction, a calibrated noise-semigroup energy estimate, and finite exact arithmetic certificates. The Hellinger inequality also yields the Courtade–Kumar information bound.

Cite (BibTeX)
@misc{OAI:Hellinger-contraction-with-arbitrary-Boolean-output-bias-September-24-2026,
  author = {{OpenAI}},
  title = {{Hellinger contraction with arbitrary Boolean output bias}},
  howpublished = {OpenAI Math Release preprint
                  \href{https://github.com/openai/math/blob/main/preprints/Hellinger-contraction-with-arbitrary-Boolean-output-bias-September-24-2026/main.pdf}{OAI:Hellinger-contraction-with-arbitrary-Boolean-output-bias-September-24-2026}},
  year = {2026}
}

Lean formalization

OpenAI's note on what the formalization covers, from lean/docs/119.md.

The Courtade–Kumar and Hellinger conjectures

The following describes the scope of the Lean formalization related to the following accompanying paper(s):

Scope

The formalization proves sharp contraction of the information carried by a binary channel under independent symmetric noise on a uniform discrete cube. At each fixed initial information level, a noisy coordinate channel attains the maximum retained information. It also proves the refined Boolean bound that accounts for output bias, with strict improvement for nonconstant biased outputs at nonzero noise correlation, and the selected mean-dependent entropy-production inequality.

For Boolean functions this includes the Courtade–Kumar inequality I(f(X);Y)≤1−h2(ε)I(f(X);Y)\le1-h_2(\varepsilon) in bits at crossover probability 0≤ε≤1/20\le\varepsilon\le1/2, where h2h_2 is binary entropy. A coordinate and its complement attain equality.

Comparator links

Result Comparator statement
Courtade–Kumar inequality and attainment CourtadeKumar.lean
Sharp binary-channel contraction and entropy production SoftChannel204.lean

Data from github.com/openai/math at commit adc7f12, committed October 6, 2026 at 21:58 UTC, last checked for changes about 9 hours ago. Titles, subjects, summaries, abstracts and Lean notes are OpenAI's; page counts are read from the PDFs. The map, related results, search, kinds of results and the named-problem index are Emergent Mind's, built with text embeddings and an LLM, and may contain errors.

An Emergent Mind Labs project. Emergent Mind is not affiliated with OpenAI. None of these results has been peer reviewed. Cite the manuscripts themselves, using the BibTeX on each result's page.