Papers
Topics
Authors
Recent
Search
2000 character limit reached

On learning functions over biological sequence space: relating Gaussian process priors, regularization, and gauge fixing

Published 26 Apr 2025 in cs.LG, q-bio.GN, and stat.ML | (2504.19034v1)

Abstract: Mappings from biological sequences (DNA, RNA, protein) to quantitative measures of sequence functionality play an important role in contemporary biology. We are interested in the related tasks of (i) inferring predictive sequence-to-function maps and (ii) decomposing sequence-function maps to elucidate the contributions of individual subsequences. Because each sequence-function map can be written as a weighted sum over subsequences in multiple ways, meaningfully interpreting these weights requires "gauge-fixing," i.e., defining a unique representation for each map. Recent work has established that most existing gauge-fixed representations arise as the unique solutions to L2L_2-regularized regression in an overparameterized "weight space" where the choice of regularizer defines the gauge. Here, we establish the relationship between regularized regression in overparameterized weight space and Gaussian process approaches that operate in "function space," i.e. the space of all real-valued functions on a finite set of sequences. We disentangle how weight space regularizers both impose an implicit prior on the learned function and restrict the optimal weights to a particular gauge. We also show how to construct regularizers that correspond to arbitrary explicit Gaussian process priors combined with a wide variety of gauges. Next, we derive the distribution of gauge-fixed weights implied by the Gaussian process posterior and demonstrate that even for long sequences this distribution can be efficiently computed for product-kernel priors using a kernel trick. Finally, we characterize the implicit function space priors associated with the most common weight space regularizers. Overall, our framework unifies and extends our ability to infer and interpret sequence-function relationships.

Summary

Exploring Sequence-to-Function Mappings via Regularization and Gaussian Processes

The paper "On learning functions over biological sequence space: relating Gaussian process priors, regularization, and gauge fixing" addresses the intricacies of modeling and interpreting sequence-to-function relationships in biological sequences, such as DNA, RNA, and proteins. A fundamental goal in this area is to understand how variations in sequences influence observable biological functions, a challenge complicated by the high dimensionality and non-identifiability of these mappings.

Summary of Contributions

The authors focus on two primary tasks: inferring predictive sequence-to-function maps and decomposing these maps to understand the contributions of individual subsequences. To achieve unique representations, they employ "gauge fixing," a method that involves choosing a specific representation for each sequence-function map. This paper makes several key contributions:

  1. Integrated Framework: The authors link regularized regression in weight space with Gaussian process regression in function space, establishing that each regularizer in weight space induces a prior in function space. They prove that for any given Gaussian process prior over function space and any gauge, there exists a regularizer that corresponds to the specified gauge.
  2. Efficient Computation: They introduce a kernel trick that allows for the efficient computation of gauge-fixed weights from Gaussian process distributions, even for long sequences. This trick leverages product-kernel priors, enabling the computation of posterior distributions without explicitly calculating the entire sequence-function map.
  3. Exploration of Regularizers: The paper investigates common regularizers and their corresponding function space priors, identifying how regularizers affect correlation decay among sequences, thus shaping the inferred sequence-function relationship.

Theoretical Implications

The theoretical results offer significant advancements in unifying methodologies for inferring sequence-to-function relationships. They enhance understanding of how different regularization schemas interplay with Bayesian approaches in both weight and function space. This unified framework potentially simplifies model selection and interpretation in biological modeling, offering a robust toolkit for assessing high-dimensional biological data.

Practical Implications and Future Directions

Practically, these methodologies enable more precise control over the representation of sequence-function maps, offering improved capabilities in areas such as protein engineering and genetic mapping. By allowing researchers to select a gauge that best suits their interpretative needs, the insights derived from data are enriched significantly.

Future work may explore extending these methods to non-linear models or exploring other forms of regularization beyond L2L_2. Further investigations could consider adapting these techniques to broader applications beyond biological contexts, such as materials science or computational genomics, where similar high-dimensional mappings are prevalent.

In summary, this paper contributes a unified framework that enhances the ability to interpret complex biological interactions through a rigorous mathematical approach, potentially transforming methodologies for predictive modeling in sequence analysis.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 4 tweets with 0 likes about this paper.