Mechanisms underlying the competitiveness of One-Hot protein representations
Clarify the mechanisms by which One-Hot protein sequence encoding achieves predictive performance comparable to or better than protein language model embeddings in protein fitness prediction, including the roles of representation sparsity, inductive bias, latent geometry, and interactions with Bayesian learning.
References
The mechanisms underlying this observation remain unclear and may involve differences in representation sparsity, inductive bias, latent geometry, or interactions with Bayesian learning, all of which require further investigation.
— Multitask Bayesian Neural Networks for Multiparameter Protein Engineering
(2608.18604 - Herrera-Rocha et al., 19 Aug 2026) in Section 3, Discussion