- The paper introduces thermodynamic epiplexity per joule for measuring learned environmental structure and empowerment per joule for measuring sensorimotor control capacity, with both requiring explicit energy and information conventions.
- The paper derives a closed-cycle learning-efficiency bound of approximately 3.5 × 10²⁰ bits per joule at 300 K, while showing that open accounting boundaries can falsely decouple information gain from dissipation by hiding pre-initialized low-entropy resources.
- The paper recommends reporting cost-constrained empowerment curves, throughput, coarse-graining, reset protocols, null-policy baselines, and complete energy balances so researchers can compare embodied AI systems without treating Landauer-scale limits as practical performance targets.
Overview and motivation
This paper, by Takahashi and Hayashi (AI Alignment Network / RIKEN AGIS), addresses a definitional gap in the evaluation of embodied AI: how to quantify the energy efficiency of intelligence in physically meaningful units. The authors propose two complementary bits-per-joule metrics—thermodynamic epiplexity per joule (ηE), measuring structural information acquired per unit energy (a "recognition" axis), and empowerment per joule (ηC), measuring sensorimotor channel capacity per unit energetic cost (a "control" axis). The paper's central methodological stance is that bits-per-joule comparisons are only meaningful under explicit accounting conventions: boundary definition, coarse-graining/noise model, horizon/reset protocol, and cost baseline. The stated goal is a reproducible efficiency report rather than a universal intelligence score.
Thermodynamic epiplexity: learning bits per joule
The normative quantity is the mutual information between an agent's internal state W and a benchmark-specified latent environment-instance variable Z: I=I(W;Z). Episode-level acquired epiplexity is defined as the conditional mutual information ΔI=I(Wpost;Z∣Wpre), which is nonnegative and captures newly encoded structure. For continuous variables, an ε-coarse-grained variant via quantization ensures finiteness. The authors are explicit that the choice of Z (and any quotienting of redundant parameterizations) is part of the benchmark specification, not something inferred from data—a candid concession that there may be no unique objective choice of "environment instance" in real-world domains.
When Z is unavailable empirically, the paper adopts compute-bounded MDL epiplexity from Finzi et al. (Finzi et al., 6 Jan 2026) as the operational companion: the model-description length L(MB⋆) of a two-part resource-bounded MDL code, with compression gain ηC0 as a practical surrogate. Importantly, the authors assume no equality between MDL epiplexity and the normative mutual information without additional modeling assumptions.
The thermodynamic core rests on the Goldt–Seifert stochastic-learning inequality [goldt2017thermodynamic]: under bipartite Markov dynamics with local detailed balance, information flow into the learner satisfies ηC1. Combining this with the data processing inequality under the Markov relation ηC2, the closed-cycle corollary states:
ηC3
The authors emphasize that this Landauer-scale figure is a far-above-hardware benchmark yardstick, not an attainable target; contemporary training stacks sit many orders of magnitude below it due to algorithmic convergence and engineering losses rather than fundamental physics.
The open-boundary decoupling result
A key conceptual contribution is Proposition 1, a simple construction showing that without boundary closure, information gain and in-boundary dissipation need not be linked. If an ηC4-bit register enters the accounting boundary pre-initialized to zero (with its preparation cost unmetered) and reversible CNOT gates copy ηC5 into it quasistatically, then ηC6 while ηC7 for arbitrary ηC8. The "missing cost" resides entirely in the unmetered negentropy resource crossing the boundary. This result motivates the paper's insistence on charging for externally prepared low-entropy resources (fresh memory provisioning and maintenance) when comparing against Landauer benchmarks, and on reporting bits/J alongside throughput (bits/s), since the quasistatic limit requires diverging wall-clock time. This is an accounting caution rather than a claim that practical systems can evade the second law; boundary closure restores the Landauer-scaled benchmark for repeated operation.
Empowerment per joule: control efficiency
On the control axis, the paper defines empowerment as the channel capacity ηC9 over horizon W0, following Klyubin et al., interpreted deliberately as a system-level capability of the embodiment rather than a score of a particular controller. The efficiency metric is cast as a capacity-per-unit-cost objective in the sense of Verdú:
W1
Two reporting conventions are distinguished—total on-boundary energy versus incremental energy above a declared null policy—with the recommendation to report both plus the baseline decomposition, since interpreting W2 without it is "generally ambiguous." The primary recommended object is the cost-constrained empowerment curve W3 and its marginal slope, which resolves the ill-posedness of asking "the energy needed to achieve empowerment." Finite-resolution conventions (W4-empowerment) are required because idealized noiseless mutual informations can diverge; these must be held fixed across compared systems, mirroring the epiplexity coarse-graining requirement.
Trade-offs and the unified framework
In closed-cycle operation, learning and control draw on shared dissipation budgets. The paper offers a coarse budget statement—W5—explicitly flagged as a conceptual yardstick, not a universal identity, valid only when marginal state distributions return across episodes and no unaccounted free-energy resources are injected. Systems may therefore exhibit trade-offs between high W6 and high W7, which joint reporting makes visible. The unified framework comes with a minimum reporting checklist covering accounting boundary, energy-balance terms (W8, W9, Z0), null-policy baselines, coarse-graining/noise, horizon/reset, time/throughput, and estimator details.
Scaling laws and quantum considerations
Connecting to empirical practice, under power-law scaling Z1 with training energy proportional to compute, the marginal compression gain per joule decays as Z2. The authors state plainly that observed diminishing returns in large-scale training are dominated by algorithmic and engineering factors, not proximity to thermodynamic limits—the energy-adjusted analysis makes those dominant factors explicit but does not imply physical boundedness. On quantum computing, the paper argues quantum hardware does not circumvent the limits: attainable structure information remains constrained by information available in agent-accessible experience, and Landauer-scale benchmarks reapply once qubit initialization, measurement, and fault-tolerance resets are inside the boundary. Quantum advantages are confined to near-reversible execution trajectories and reduced logical step counts for certain problem classes, with measured efficiency still potentially dominated by cryogenic and control-electronics overheads.
Limitations and open questions
Several limitations are conceded explicitly. The normative metric depends on a chosen latent variable Z3 whose specification is conventional; the closed-cycle bound requires assumptions (bipartite Markov dynamics, local detailed balance, Z4 steady state) that may not hold for real learners; translation from Z5 bounds to measured-Z6 claims is a reporting convention unless the energy balance terms are negligible or measured; and the free-energy-principle connection acknowledged in the discussion is noted as controversial. The coarse information-budget statement combining epiplexity and empowerment is heuristic. The paper leaves open the concrete instantiation of its checklist in fully specified benchmarks with validated estimators—an immediate next step the authors identify themselves.
Conclusion
The paper contributes a disciplined two-axis vocabulary—recognition efficiency (Z7) and control efficiency (Z8)—for physical intelligence, derives a Landauer-scale closed-cycle benchmark for learning efficiency as a corollary of standard thermodynamic-learning inequalities, and demonstrates via an explicit decoupling construction why boundary closure and cost conventions are indispensable for meaningful bits-per-joule comparison. Its principal message is methodological: without explicit accounting of boundaries, coarse-graining, horizons, and low-entropy resources, information-gain and dissipation can be arbitrarily decoupled, rendering efficiency claims vacuous.