Papers
Topics
Authors
Recent
Search
2000 character limit reached

Machine-Learning-Assisted Directed Evolution (MLDE)

Updated 7 September 2026
  • Machine-Learning-Assisted Directed Evolution (MLDE) is a framework that uses machine learning models to enhance the efficiency of directed evolution, guiding the creation and testing of protein variants based on sequence-functio.
  • By utilizing both functional and nonfunctional variants, MLDE can intelligently explore large sequence spaces, reducing experimental costs and improving outcomes over conventional methods in applications such as enzyme engineering, stability improvements, and drug design.
  • Recursive model updating and robust acquisition strategies enable MLDE to optimize biological sequences and phenotypes within the constraints of experimental resources, significantly enhancing the practical aspects of protein engineering.

Machine-learning-assisted directed evolution (MLDE) is an experimental–computational framework that augments conventional directed evolution with models learned from protein sequence–function data. A surrogate model approximates the relationship between sequence or sequence-derived representation and an experimentally measured property, then guides the construction and testing of subsequent variants. The target may be activity, stability, specificity, binding, fluorescence, expression, localization, productivity, enantioselectivity, or another phenotype. Unlike conventional directed evolution, which generally retains selected variants while discarding much of the information from unsuccessful ones, MLDE can use measurements from both functional and nonfunctional variants to guide search across a combinatorial fitness landscape (Yang et al., 2018).

1. Conceptual foundations and objectives

The central abstraction of MLDE is a noisy sequence–function mapping,

y=f(x)+ε,y=f(x)+\varepsilon,

where xx is a protein sequence or representation, ff is the unknown biological mapping, yy is the measured phenotype, and ε\varepsilon represents measurement noise, biological variability, batch effects, and other unmodeled factors. Given a dataset

D={(xi,yi)}i=1N,\mathcal{D}=\{(x_i,y_i)\}_{i=1}^{N},

a supervised model f^θ\hat f_\theta is trained and used to predict untested sequences,

y^=f^θ(x).\hat y=\hat f_\theta(x).

The purpose of the model is not necessarily to reconstruct the physical mechanism of the protein. It must instead provide predictions sufficiently useful for experimental decision-making. The practical objective is therefore sequential optimization under limited synthesis and assay budgets rather than globally accurate reconstruction of the fitness landscape.

Protein sequence space is combinatorial. A protein of length LL has approximately 20L20^L possible amino-acid sequences. For a 300-residue protein, the number of single amino-acid substitutions is

xx0

whereas the number of two-substitution variants is

xx1

For xx2 simultaneously randomized positions, the theoretical amino-acid library contains xx3 sequences; even four positions yield xx4 variants. Exhaustive screening is therefore generally impossible, particularly when assays are slow, expensive, or unavailable at high throughput.

Conventional directed evolution typically cycles through diversification, screening or selection, and propagation of selected variants. MLDE inserts model fitting and computational candidate selection between experimental rounds:

  1. define the engineering objective;
  2. design and measure an initial library;
  3. encode sequences numerically;
  4. fit and validate a sequence–function model;
  5. select candidates using predicted fitness, uncertainty, diversity, or information gain;
  6. experimentally test the selected variants;
  7. add the new measurements to the training set;
  8. retrain and iterate.

The distinction between finding an optimal protein and finding a sufficient protein is consequential. A recent retrospective argues that practical directed evolution generally seeks a protein meeting a project-specific threshold within constraints on time, DNA synthesis, screening, labor, and risk, whereas much MLDE work optimizes the identity or predicted fitness of individual variants (Wittmann, 2 Sep 2026). This motivates resource-aware and distribution-focused objectives in which the desired output may be an overview-feasible library distribution rather than a list of individually ranked sequences.

2. Experimental–computational workflow

Objective definition and initial library design

The first step is to specify the phenotype to optimize. Examples include enzyme productivity, thermostability, activity, substrate specificity, enantioselectivity, fluorescence, membrane localization, channelrhodopsin conductance, expression, binding affinity, solubility, and viral or cellular function (Yang et al., 2018).

The initial training set defines the region of sequence space from which the model can learn. Three broad sampling strategies are commonly distinguished:

  • Random sampling: variants are sampled from an available library without explicit information optimization.
  • Mutation-focused sampling: variants are selected so that individual mutations appear in multiple sequence contexts.
  • Diversity- or information-focused sampling: variants are selected to cover sequence space, maximize mutation coverage, or reduce predictive uncertainty.

Variant-generation methods include random mutagenesis, site-saturation mutagenesis, structure-guided mutation, sequence-homology-guided mutation, recombination of previously identified mutations, recombination of homologous fragments, and direct synthesis of designed sequences.

Combinatorial MLDE explicitly explores multiple positions simultaneously. Wu and colleagues implemented a sample–model–enumerate–restrict–test workflow in which a broad input library is measured, several regressors are trained, the complete combinatorial library is scored in silico, and model-enriched degenerate-codon libraries are experimentally screened (Wu et al., 2019). For xx5 positions, the theoretical space is xx6. In the enzyme application, NDT codons encoded 12 amino acids,

xx7

and SwiftLib was used to construct restricted libraries enriched for amino acids frequently appearing among high-scoring model predictions.

The translation from model-ranked sequences to a degenerate library is a compromise. Independent recombination of amino-acid frequencies can lose predicted higher-order combinations and introduce nonoptimal variants, but it reduces DNA synthesis and cloning costs. Direct synthesis of individually ranked sequences is preferable when affordable.

Sequence–function data generation

Selected variants are synthesized or constructed, expressed, and assayed. The dataset may contain quantitative phenotypes or binary functional labels. Negative data are important because they constrain deleterious mutations, functional boundaries, epistatic interactions, and regions of sequence space to avoid.

Assay choice determines the phenotype learned by the model. Apparent binding may depend on expression, purification, assay conditions, and cellular context; whole-cell activity may conflate catalytic activity with expression, heme loading, or background product formation. MLDE therefore learns the experimental assay as implemented, not necessarily an intrinsic molecular property.

Throughput varies substantially. Deep mutational scanning and sequencing-linked assays can generate approximately xx8–xx9 measurements, whereas chromatographic activity assays may generate only ff0–ff1 measurements. MLDE is most valuable when screening is the bottleneck, but those same settings often provide the smallest training sets (Johnston et al., 2023).

Iterative model updating

At iteration ff2, a selected batch ff3 is measured and incorporated into the training data:

ff4

The model is then retrained, candidates are rescored, and a new batch is selected. Iteration can end when the engineering target is reached, resources are exhausted, or the best observed function no longer improves.

The framework is an active-learning loop rather than a one-time prediction task. Its effectiveness depends on the relationship between model uncertainty, candidate-generation constraints, assay noise, and the distribution of variants actually synthesized.

3. Representations and predictive models

Sequence representations

One-hot encoding represents each residue using a categorical vector. For a sequence of length ff5 and alphabet size ff6, the representation has approximately ff7 dimensions. It avoids imposing an artificial numerical ordering on amino acids and is often a strong baseline for small, local datasets, but it is sparse and does not intrinsically encode biochemical similarity.

Mutation encoding represents substitutions relative to a parent. In the described 20-dimensional form, the original amino acid is encoded as ff8, the substituted amino acid as ff9, and all others as zero. Physicochemical representations use descriptors such as charge, hydrophobicity, and volume. Evolutionary representations include BLOSUM matrices, multiple-sequence alignments, hidden Markov models, Potts models, and related covariance models.

Protein-LLMs provide learned sequence embeddings from large unlabeled databases. UniRep used global unsupervised pretraining on more than 20 million proteins followed by target-family adaptation through evotuning (Narayanan et al., 2021). ESM-2 was pretrained on more than 60 million protein sequences and can provide contextual residue representations for downstream structure-aware models (Tan et al., 2023). Protein-language-model scores can also be used zero-shot, through pseudo-log-likelihood or related sequence-probability measures, although evolutionary plausibility is not identical to the desired engineered phenotype (Qiu et al., 2023).

Structure-aware representations

Structure-aware encodings incorporate spatial relationships, residue contacts, local chemical environments, predicted secondary structure, structural neighborhoods, or three-dimensional coordinates. In the cytochrome P450 thermostability case study, a structure-aware one-hot representation was more predictive than primary-sequence one-hot encoding (Yang et al., 2018).

Graph neural networks represent residues or atoms as nodes and contacts or geometric relationships as edges. Equivariant graph neural networks preserve appropriate behavior under translations and rotations. The 13 framework cascades a frozen ESM-2 sequence representation with six equivariant graph neural-network layers operating on a residue graph with 20-nearest-neighbor connectivity and geometric edge features (Tan et al., 2023).

Topological data analysis provides multiscale descriptors from structural point clouds and complexes. Persistent homology tracks the birth and death of connected components, loops, and voids over a filtration. Persistent Laplacians and spectral graph methods augment topological invariants with nonzero eigenvalues that can encode shape evolution. Persistent path, sheaf, hypergraph, hyperdigraph, and evolutionary de Rham–Hodge constructions are proposed as structure-aware representations for heterogeneous, higher-order, or directed relationships (Qiu et al., 2023).

Model classes

Model choice depends on dataset size, epistasis, representation, and uncertainty requirements.

  • Linear and ridge models: computationally efficient and interpretable; appropriate when effects are approximately additive or data are sparse.
  • Partial least squares regression: projects high-dimensional inputs and outputs into latent spaces; useful when the number of variables exceeds the number of observations.
  • Decision trees and ensembles: random forests and boosted trees are strong baselines for small biological datasets, particularly below approximately yy0 examples.
  • Kernel methods: support-vector machines and kernel ridge regression use sequence or structure similarity functions.
  • Gaussian processes: provide predictive means and uncertainty estimates, with exact regression cost scaling approximately as yy1.
  • Neural networks: multilayer perceptrons, CNNs, RNNs, transformers, and graph neural networks become more attractive as labeled datasets increase or as pretrained representations are available.
  • Generative models: VAEs, GANs, autoregressive models, and related systems model sequence distributions and may generate candidate proteins, although their general engineering utility requires experimental validation.
  • Ensembles and committees: model disagreement can approximate epistemic uncertainty, but uncertainty estimates require calibration.

A critical methodological issue is that average held-out prediction error may not reflect MLDE utility. Candidate selection concentrates on the predicted upper tail, often outside the training distribution. A model can have acceptable average correlation but poor calibration where synthesis decisions are made.

4. Acquisition, exploration, and optimization strategies

Mutation-level and greedy selection

Linear or additive models can classify mutations as beneficial, neutral, or deleterious. Beneficial mutations may be retained, deleterious mutations removed, and neutral mutations retested. If the candidate space is small enough, the model can score all feasible sequences and select the highest-scoring variants.

Greedy selection is simple but can repeatedly sample near-duplicates, overexploit model errors, and become trapped in local optima. It is particularly vulnerable when the training data are biased around a few parents or when epistatic combinations are poorly represented.

Bayesian optimization

When the surrogate provides uncertainty, Bayesian optimization balances exploitation and exploration. For a Gaussian process with posterior mean yy2 and standard deviation yy3, an upper-confidence-bound rule is

yy4

where yy5 controls the exploration weight. Expected improvement and probability of improvement are alternative acquisition functions.

The P450 thermostability case used Gaussian processes to predict both continuous yy6 and functionality. Expected mutual information selected variants that were informative about the recombination landscape while retaining a high probability of being functional. Subsequent GP-UCB rounds alternated exploration and exploitation; two of five sequences in a final exploitation round were more thermostable than previously observed (Yang et al., 2018).

ODBO combines a low-dimensional, function-value-based representation with information-covering initialization, XGBOD outlier prescreening, Gaussian-process Bayesian optimization, optional RobustGP regression, and TuRBO trust regions (Cheng et al., 2022). In the GB1(4) benchmark, the best ODBO protocol recovered the global optimum in all ten trials with 40 initial samples and 50 sequential iterations. Its function-value-based representation was four-dimensional, compared with 80 dimensions for one-hot encoding and 76 for physicochemical encoding. The benchmark result was obtained by replaying an existing dataset rather than by a new wet-laboratory campaign.

Thompson sampling and bandit formulations

TS-DE formulates MLDE as Bayesian bandit optimization with evolutionary action constraints. A posterior over a sequence–function model is updated from noisy measurements; a plausible function is sampled from that posterior; mutation, recombination, and selection then generate experimentally testable sequences. In the stated linear model,

yy7

with Gaussian prior and Gaussian measurement noise, the theoretical Bayesian regret scales as

yy8

under the specified assumptions and sufficiently large population size. The result has square-root dependence on the total number of evaluated individuals but does not cover nonlinear protein fitness maps or realistic biological mutation operators (Yuan et al., 2022).

Guided proposal and product-of-experts methods

Plug and Play Directed Evolution combines unsupervised evolutionary models with supervised activity predictors in a product-of-experts distribution,

yy9

Potts models and protein LLMs constrain sequences toward evolutionary plausibility, while supervised predictors favor target activity. Gradient-informed discrete MCMC proposes amino-acid substitutions using gradients through continuous one-hot extensions. The method was evaluated in silico on PABP, UBE4B, and GFP landscapes, including ESM2 models up to 650 million parameters, but the reported experiments did not include wet-laboratory validation (Emami et al., 2022).

Diversity can be treated as an explicit acquisition objective. Candidate batches may combine high predicted fitness, high uncertainty, high expected information, sequence diversity, structural diversity, and functionality probability. This is important because a batch of highly correlated sequences can provide less information than a more diverse batch with slightly lower predicted scores.

Conformational Rank Conditioned Committees extend this principle to antibody design by assigning separate neural-network committees to ranked predicted conformations. Within-rank disagreement estimates epistemic uncertainty, while between-rank disagreement estimates conformational uncertainty. The approach was evaluated for SARS-CoV-2 antibody docking, but the specific incremental benefit of rank conditioning over a pooled committee was not established by a fully controlled ablation (Adler et al., 28 Oct 2025).

5. Empirical evidence and representative applications

PLS-guided enzyme productivity

Fox and colleagues applied partial least squares regression to halohydrin dehalogenase-catalyzed cyanation. The campaign tested 519,045 variants over 18 rounds and reported approximately a 4,000-fold improvement in volumetric productivity. The model estimated mutation contributions and classified them as beneficial, neutral, or deleterious. No individual round produced more than a threefold improvement, illustrating that substantial cumulative gains can coexist with large experimental burdens (Yang et al., 2018).

Combinatorial GB1 benchmark

The GB1 benchmark contained 160,000 four-position sequences, of which 149,361 were experimentally measured. MLDE used 470 randomly selected sequences for training, predicted the complete combinatorial space, and treated the top 100 predictions as the experimental set, giving a total burden of 570 variants. At this equal burden, MLDE achieved expected fitness 6.42 and reached the global maximum in 8.17% of simulations, compared with 5.41 and 4.91% for a single-mutant walk.

The held-out Pearson correlation was only ε\varepsilon0, demonstrating that useful upper-tail ranking does not require globally accurate fitness prediction. The method was probabilistic rather than guaranteed to outperform conventional evolution in every run (Wu et al., 2019).

Stereodivergent enzyme engineering

A putative nitric oxide dioxygenase from Rhodothermus marinus was evolved for stereodivergent carbene Si–H insertion. Two rounds fixed seven mutations across the targeted positions and produced variants with 93% ee for one enantiomer and 79% ee for the other. The campaign generated 445 sequence–function relationships, tested 360 predicted-library variants, and performed 805 total experimental tests.

The result illustrates that MLDE can identify multiple sequence solutions to the same function rather than merely recombining the best amino acid observed independently at each site. It also demonstrates that successful enrichment can occur despite weak validation metrics in one lineage.

P450 thermostability

Romero and colleagues trained Gaussian-process models on 242 chimeric P450s, using ε\varepsilon1 and binary function as targets. Information-guided selection tested sequences averaging 106 mutations from the closest parent; 26 of 30 selected variants were functional. Subsequent exploration and exploitation identified variants more thermostable than any previously observed, although initial GP-UCB rounds improved landscape coverage without increasing the maximum ε\varepsilon2 (Yang et al., 2018).

ODBO benchmark results

ODBO was evaluated on GB1(4), GB1(55), Ube4b/BRCA1 RING, and avGFP benchmark datasets. On GB1(4), ODBO with TuRBO and RobustGP achieved average maximum fitness ε\varepsilon3 after 50 iterations, compared with ε\varepsilon4 for random search and ε\varepsilon5 for TuRBO with a conventional GP. On GB1(55), the best ODBO configuration achieved average maximum enrichment ε\varepsilon6 but did not recover the true maximum. On Ube4b and avGFP, ODBO approached the dataset maxima.

These were benchmark replay experiments using existing measurements. They demonstrate closed-loop selection under pre-existing labels but do not constitute new wet-laboratory evidence (Cheng et al., 2022).

Zero-shot and structure-aware prediction

The 13 framework combines frozen ESM-2 sequence embeddings with an equivariant graph neural network trained through mutation-and-denoising pretraining. It achieved ProteinGym Spearman correlations of 0.424 for single mutations, 0.395 for double mutations, and 0.426 across all variants. For MLDE, its principal demonstrated role is cold-start mutation ranking and library design rather than a complete experimental model-update loop (Tan et al., 2023).

A related review reports that persistent Laplacian representations achieved mean Spearman correlations of 0.280, 0.457, 0.525, and 0.564 with 24, 96, 168, and 240 labels, respectively, in a TopFit comparison. These results suggest that structure-aware spectral representations can be competitive in low-label regimes, although representation performance depends on dataset, structural input, and model configuration (Qiu et al., 2023).

Antibody MLDE and conformational uncertainty

The RCC antibody framework used ImmuneBuilder to generate four conformations per antibody, HADDOCK3 for docking, AbMAP embeddings, and rank-specific DNN committees. In the reported comparison, DNN MLDE increased the mean docking-score value from an initial 69.83 to 93.95 after the second batch, a reported 34.5% increase. XGBoost MLDE reached 90.47 and bioinformatics-only directed evolution reached 75.21.

These results concern computational docking scores, supplemented by limited cell-free AlphaLISA follow-up. They do not provide a controlled quantitative ablation proving that rank conditioning itself outperforms pooled committees.

6. Limitations, controversies, and future directions

Data scarcity and bias

Protein-engineering datasets are small, biased, and concentrated around a few parents or mutation sites. Natural sequence databases are also biased by phylogeny, evolutionary history, sampling, and functional constraints. A model may learn local relationships without acquiring globally valid rules.

Public datasets often combine measurements obtained under different expression systems, assay conditions, organisms, and normalization procedures. Missing metadata can cause models to learn experimental context rather than intrinsic molecular function. Careful curation, deduplication, functional filtering, and preservation of negative examples are therefore essential.

Epistasis and extrapolation

Additive models assume

ε\varepsilon7

Epistasis requires interaction terms,

ε\varepsilon8

The number of possible interactions grows rapidly, while training datasets often contain mostly single and double mutants. A mutation that is beneficial in one background may be neutral or deleterious in another. Models trained near wild type may be unreliable for distant recombinations, high-order mutants, or generated sequences.

Uncertainty and calibration

Predictive variance can reflect model uncertainty, sparse data, or disagreement among models, but it does not necessarily capture irreducible biological or experimental noise. Uncertainty estimates must be calibrated if they are used in acquisition functions. Gaussian processes, ensembles, Bayesian neural networks, conformal prediction, and rank-conditioned committees provide alternative approaches, but no single method is universally reliable under adaptive distribution shift.

Model quality versus experimental utility

High prediction accuracy is not sufficient evidence of practical benefit. MLDE must be evaluated on prospective high-fitness discovery, assay cost, synthesis cost, number of experimental cycles, time to a sufficient protein, and probability of success under a fixed budget. The retrospective critique of MLDE emphasizes that DNA synthesis and library-construction costs are frequently omitted, even though they can dominate the economics of targeted variant design (Wittmann, 2 Sep 2026).

Distribution-focused MLDE addresses this problem by optimizing synthesis-feasible library parameters: degenerate-codon composition, mutation spectra, recombination boundaries, mutable positions, error rates, and other factors controlling the distribution of generated variants. DeCOIL, variational synthesis, and policy-gradient library design are examples of approaches that optimize experimentally realizable distributions rather than individually ordering every model-generated sequence.

Structural and generative limitations

Structure-based methods depend on experimental or predicted structures. Static structures may not represent conformational ensembles, ligand-induced rearrangements, allostery, solvent effects, or dynamics. AlphaFold2 and related predictors can be unreliable for proteins with few homologs, flexible regions, or substantial sequence divergence.

Generative models can produce sequence-divergent variants, but generated sequences may misfold, fail to express, be inactive, toxic, or exploit predictor errors. Protein-LLMs provide evolutionary priors rather than direct mechanistic guarantees. Their natural-sequence bias can be problematic for non-natural functions or objectives not optimized by natural selection.

Practical and methodological standards

Robust MLDE campaigns should:

  • compare against random mutagenesis, site-saturation mutagenesis, recombination, and expert-designed libraries;
  • preserve and sequence unsuccessful variants when possible;
  • separate training, validation, and prospective test data;
  • report high-tail enrichment rather than only average prediction error;
  • quantify uncertainty calibration and batch diversity;
  • account for synthesis, cloning, screening, sequencing, and turnaround costs;
  • use replicate assays and explicit noise models;
  • test generalization across sequence distance and protein families;
  • report complete protocols, budgets, hardware, and stopping rules;
  • validate computationally selected variants experimentally.

The broader research direction is a closed-loop, multimodal, resource-aware system combining protein-LLMs, structural and topological representations, assay-specific supervised models, uncertainty-aware acquisition, generative or combinatorial library design, and experimentally validated iteration. The intended endpoint is not a universally accurate predictor or guaranteed global optimum, but a system that allocates finite experimental resources toward a sufficiently good protein with an acceptable probability, cost, and time to success.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Machine-Learning-Assisted Directed Evolution (MLDE).