Mechanisms underlying the competitiveness of One-Hot protein representations

Clarify the mechanisms by which One-Hot protein sequence encoding achieves predictive performance comparable to or better than protein language model embeddings in protein fitness prediction, including the roles of representation sparsity, inductive bias, latent geometry, and interactions with Bayesian learning.

Background

The benchmark found that simple One-Hot encoding repeatedly performed comparably to large pretrained protein LLM embeddings and sometimes outperformed them, particularly in high-data regimes. This result challenges the prevailing assumption that protein LLMs generally provide superior representations for protein fitness prediction.

The discussion identifies several possible explanations—representation sparsity, inductive bias, latent geometry, and interactions with Bayesian learning—but does not determine which mechanisms account for the observed performance. The authors explicitly state that these mechanisms require further investigation.

References

The mechanisms underlying this observation remain unclear and may involve differences in representation sparsity, inductive bias, latent geometry, or interactions with Bayesian learning, all of which require further investigation.

Multitask Bayesian Neural Networks for Multiparameter Protein Engineering  (2608.18604 - Herrera-Rocha et al., 19 Aug 2026) in Section 3, Discussion