Papers
Topics
Authors
Recent
Search
2000 character limit reached

SimpleDesign: Multimodal Protein Design Model

Updated 5 September 2026
  • SimpleDesign operates directly in data space, modeling protein sequences with categorical discrete denoising and structures with continuous velocity regression.
  • The model utilizes a shared Transformer with modality-specific processing, generating over 1.8 million sequence-structure pairs and achieving competitive performance in sequence generation, structural generation, and joint codesign.
  • Despite not overtaking specialized geometric models in structural consistency, SimpleDesign offers a simplified, efficient approach that minimizes complexity and error propagation.

SimpleDesign is a multimodal generative model for joint protein sequence–structure codesign. It generates amino-acid sequences and three-dimensional protein backbones within a single-stage, end-to-end framework operating directly in data space rather than through discrete sequence or structure tokenizers. The model combines masked discrete denoising for sequences with continuous velocity regression for structures and uses a shared Transformer with modality-specific processing. Trained on more than 2 million sequence–structure pairs, SimpleDesign achieves competitive codesignability and sequence-generation performance relative to tokenized multimodal protein LLMs, while remaining weaker than specialized geometric flow models on some structural-consistency benchmarks (Lu et al., 3 Sep 2026).

1. Motivation and design objective

Protein function depends on the interaction between amino-acid sequence and three-dimensional structure. A generative model for protein design must therefore represent both categorical sequence information and continuous geometric information, while maintaining consistency between the two modalities.

Many multimodal protein-generation systems use a multi-stage pipeline. Sequence and structure autoencoders or tokenizers are first trained to produce latent representations; a generative model is then trained over those representations; finally, generated latent codes are decoded back into sequences or structures. This design introduces several sources of complexity:

  • Error propagation: the generative model learns the distribution of imperfect latent reconstructions rather than the original data.
  • Information loss: continuous structural coordinates must be compressed into a finite vocabulary.
  • Objective misalignment: tokenizer reconstruction objectives may not coincide with sequence–structure generation objectives.
  • Additional engineering: modality-specific tokenizers, decoders, vocabularies, and pretraining stages are required.
  • Weak end-to-end coupling: sequence and structure consistency is mediated through separately learned representations.

SimpleDesign addresses these issues by modeling amino-acid sequences and backbone coordinates directly. Its main hypothesis is that structure tokenization is not necessary for competitive sequence–structure codesign. The model uses one joint training objective, one multimodal Transformer trunk, and two modality-specific prediction objectives:

  1. masked discrete prediction for amino-acid sequences;
  2. continuous velocity regression for structural coordinates.

The model represents a protein sequence as

a=(a(1),…,a(L)),a(i)∈V,a=(a^{(1)},\ldots,a^{(L)}),\qquad a^{(i)}\in\mathcal V,

where LL is the residue count and V\mathcal V is the amino-acid vocabulary. The structure is represented by CαC_\alpha coordinates

x=(x(1),…,x(L))∈RL×3,x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},

where x(i)x^{(i)} is the three-dimensional coordinate of residue ii. The target is the joint distribution

pθ(a,x)≈qdata(a,x),p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),

where qdataq_{\mathrm{data}} denotes the empirical distribution of paired sequences and structures.

2. Joint corruption and training objectives

SimpleDesign uses independent corruption processes for sequence and structure. The sequence corruption time is t∈[0,1]t\in[0,1], while the structural interpolation time is LL0. These variables are sampled independently, allowing the model to represent folding-like, inverse-folding-like, and joint codesign regimes.

When LL1 and LL2, the sequence is mostly visible while the structure is highly noisy, approximating structure prediction from sequence. When LL3 and LL4, the structure is visible while the sequence is highly masked, approximating inverse folding. Intermediate values represent joint sequence–structure denoising.

Sequence denoising

For each residue, a Bernoulli variable determines whether the amino acid is masked:

LL5

The corrupted sequence is

LL6

The sequence head predicts a categorical distribution over the amino-acid vocabulary at masked positions. The masked cross-entropy objective is

LL7

with LL8. The mask indicator restricts the loss to corrupted positions, while the denominator prevents division by zero when no position is masked. The factor LL9 downweights examples with high masking rates in the reported formulation.

Structural velocity regression

Structures are perturbed using a linear interpolation between Gaussian noise and the clean coordinates:

V\mathcal V0

V\mathcal V1

The ideal velocity along this path is

V\mathcal V2

The structural objective is

V\mathcal V3

where V\mathcal V4 is the predicted structural velocity. During generation, numerical integration of this vector field transports Gaussian noise toward the protein-structure distribution.

Because protein structures are invariant under global translation and rotation, SimpleDesign applies Kabsch alignment before constructing structural targets. The procedure centers predicted and reference coordinates, computes a covariance matrix, obtains a singular-value decomposition, constructs the optimal rotation, corrects reflections when necessary, and aligns the reference structure. This prevents coordinate-level error from penalizing equivalent structures that differ only by rigid transformation.

Single-stage multimodal objective

For paired data V\mathcal V5, the total loss combines the sequence and structure objectives:

V\mathcal V6

The reported implementation uses V\mathcal V7. The sequence time is sampled uniformly:

V\mathcal V8

whereas the structural time uses

V\mathcal V9

This distribution emphasizes later, less-noisy structural states while retaining some coverage of highly corrupted structures.

3. Architecture

Direct continuous structure representation

SimpleDesign does not convert structures into discrete tokens. Each residue coordinate CαC_\alpha0 is passed through Fourier feature encoding, followed by a linear projection and layer normalization:

CαC_\alpha1

The sequence stream uses a learnable amino-acid embedding:

CαC_\alpha2

For a protein of length CαC_\alpha3, the sequence and structure representations are concatenated into a joint stream:

CαC_\alpha4

The Transformer therefore receives CαC_\alpha5 residue-level representations. Sequence and structure tokens share residue indices, providing an alignment signal for

CαC_\alpha6

The model uses additive sinusoidal positional encodings and rotary positional embeddings inside attention layers. Global self-attention allows sequence representations to attend to structural representations and vice versa, without requiring a separate cross-attention module.

Mixture-of-Transformer

The default architecture is a Mixture-of-Transformer (MoT). MoT retains global joint self-attention while assigning modality-specific parameters to the two streams:

  • sequence-specific query, key, and value projections;
  • structure-specific query, key, and value projections;
  • modality-specific layer normalization;
  • modality-specific feed-forward networks.

For modality CαC_\alpha7, the projections can be written as

CαC_\alpha8

The resulting representations participate in shared attention:

CαC_\alpha9

MoT is intended to balance two extremes. A completely shared Transformer may underrepresent the distinction between categorical and continuous modalities, while two independent networks would lose direct cross-modal reasoning. MoT provides modality-specific processing while preserving joint attention.

The ablation results indicate that MoT is not essential to the central contribution. A vanilla shared Transformer is highly competitive and sometimes performs better. Accordingly, the principal methodological contribution is the tokenizer-free joint objective rather than the MoT parameterization itself.

Output heads

The sequence head is an MLP followed by layer normalization and a projection to the amino-acid vocabulary:

x=(x(1),…,x(L))∈RL×3,x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},0

x=(x(1),…,x(L))∈RL×3,x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},1

The final sequence projection is tied to the input amino-acid embedding matrix.

The structure head uses an MLP with adaptive layer normalization conditioned on the structural time x=(x(1),…,x(L))∈RL×3,x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},2:

x=(x(1),…,x(L))∈RL×3,x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},3

The resulting representation is projected into three-dimensional velocity vectors:

x=(x(1),…,x(L))∈RL×3,x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},4

The architecture is not explicitly SE(3)-equivariant. Instead, the implementation uses random rigid rotations and translations during training and Kabsch alignment for structural supervision.

4. Data, initialization, and optimization

SimpleDesign is trained primarily on AFESM, a resource derived from the AlphaFold Database and the ESM Metagenomic Atlas. The source resource contains more than 800 million predicted structures. The authors cluster sequences and structures into approximately 5 million non-singleton structural clusters and retain one representative from each cluster.

The training set is restricted to proteins with lengths satisfying

x=(x(1),…,x(L))∈RL×3,x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},5

and pLDDT greater than $85. No filtering is performed based on secondary-structure or coil content. The resulting dataset contains 1,807,333 training proteins, with 1,000 proteins held out for validation.

The model is subsequently fine-tuned on a filtered AFDB-derived SwissProt subset containing 442,511 high-quality sequence–structure pairs.

The Transformer and sequence embeddings are initialized from ESM2-650M weights. For MoT, sequence-side attention, normalization, and feed-forward components can be initialized from ESM2, while structure-specific components are randomly initialized.

Reported optimization settings include:

  • Optimizer: AdamW.
  • Learning rate: $x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},$6 for the standard Transformer and $x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},$7 for MoT.
  • Weight decay: none.
  • Warm-up: 5,000 steps from $x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},$8 to the target learning rate.
  • Gradient clipping: norm $x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},$9.
  • Training duration: 300,000 AFESM steps.
  • Fine-tuning: 50,000 SwissProt steps.
  • Hardware: 64 NVIDIA H100 80GB GPUs.
  • Effective outer batch size: 128.
  • Gradient accumulation: 2.
  • Inner replica batch size: 16.

Coordinates are rescaled using

$x^{(i)}$0

Samples are repeatedly presented under different masking, noising, and rigid transformations.

5. Inference and generation

Sequence generation

Sequence generation begins with a fully masked sequence. At each denoising step, the model produces logits $x^{(i)}$1 for every residue. Special-token logits are suppressed. Gumbel noise is added:

$x^{(i)}$2

where

$x^{(i)}$3

and $x^{(i)}$4.

Temperature scaling is then applied, with the temperature annealed approximately from $x^{(i)}$5 to $x^{(i)}$6. Tokens are sampled categorically, and the highest-confidence positions are progressively unmasked. Residues whose frequency exceeds $x^{(i)}$7 are remasked and resampled to prevent domination by a single amino acid.

Structure generation

Structure generation initializes coordinates with Gaussian noise:

$x^{(i)}$8

The model’s velocity field is converted into a score-like quantity:

$x^{(i)}$9

A stochastic flow is simulated using an Euler–Maruyama discretization:

$i$0

where

$i$1

and $i$2 controls stochasticity. Coordinates are centered and randomly rotated during sampling, then rescaled to their original units. The generated structure is a $i$3-only backbone rather than an all-atom structure.

Joint codesign

Joint generation begins with a masked sequence and Gaussian structural coordinates. The sequence and structure streams are repeatedly passed through the joint Transformer. At each iteration, the model:

  1. predicts sequence logits;
  2. samples and selectively unmasks sequence residues;
  3. predicts structural velocity;
  4. advances the coordinates by one integration step;
  5. repeats until both modalities approach their clean states.

The sequence time follows an approximately linear schedule, while the structural time uses a log-spaced schedule with more steps near the clean-data endpoint. The reported inference parameter $i$4 controls a fidelity–diversity trade-off. Sequence-only results use $i$5, joint codesign results use $i$6 and $i$7, and structure-only results commonly use $i$8.

The mathematically described generation path begins near $i$9, corresponding to a masked sequence and Gaussian structure, and proceeds toward $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$0, the clean data state.

6. Evaluation and limitations

The main benchmark settings generate 100 samples at lengths

$p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$1

Unconditional sequence generation

Sequence generation is evaluated using ProGen2 perplexity, ESMFold pLDDT, MMseqs2 diversity, and sequence novelty relative to SwissProt.

Method PPL $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$2 pLDDT $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$3 MMseqs diversity Novelty
EvoDiff $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$4 $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$5 1.00 0.49
DPLM $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$6 $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$7 0.82 0.49
ESM3, sequence to structure $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$8 $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$9 0.58 0.45
DPLM2 $q_{\mathrm{data}}$0 $q_{\mathrm{data}}$1 0.56 0.90
SimpleDesign $q_{\mathrm{data}}$2 $q_{\mathrm{data}}$3 0.50 0.80

SimpleDesign produces protein-like and foldable sequences, although it is not the best sequence generator on every metric. Progressive structural realization can constrain sequence diversity.

Unconditional structure generation

Structures are evaluated after inverse folding with ProteinMPNN and refolding with ESMFold. The protocols PMPNN1 and PMPNN8 use one and eight inverse-folded sequences, respectively. Metrics include designability under scRMSD and scTM thresholds, pairwise TM-score similarity, FoldSeek cluster diversity, and novelty.

Method PMPNN1 designability PMPNN8 designability PMPNN8 TM similarity PMPNN8 FoldSeek diversity
ESM3 sequence to structure 0.17 / 0.19 0.24 / 0.27 0.39 / 0.34 0.41 / 0.50
DPLM2 0.31 / 0.48 0.52 / 0.66 0.28 / 0.27 0.47 / 0.44
SimpleDesign 0.44 / 0.63 0.60 / 0.78 0.29 / 0.30 0.27 / 0.23
MultiFlow 0.86 / 0.90 0.95 / 0.98 0.33 / 0.33 0.52 / 0.52
La-proteina, no triangle 0.84 / 0.86 0.95 / 0.97 0.33 / 0.32 0.61 / 0.61

SimpleDesign substantially outperforms the tokenized multimodal protein-language-model baselines for structural generation, but specialized geometric models remain stronger in designability.

Sequence–structure codesign

Codesignability measures the proportion of generated sequence–structure pairs whose folded sequence agrees with the generated structure.

Method Codesignability TM similarity FoldSeek diversity Novelty
ESM3 sequence to structure 0.09 / 0.11 0.30 / 0.29 0.59 / 0.61 0.91
DPLM2 0.30 / 0.46 0.29 / 0.28 0.51 / 0.39 0.95 / 0.96
SimpleDesign, qdataq_{\mathrm{data}}4 0.53 / 0.74 0.31 / 0.30 0.18 / 0.14 0.97 / 0.97
SimpleDesign, qdataq_{\mathrm{data}}5 0.36 / 0.55 0.29 / 0.30 0.30 / 0.26 0.98 / 0.97
MultiFlow 0.76 / 0.80 0.34 / 0.34 0.54 / 0.52 0.83 / 0.83
La-proteina, triangle 0.77 / 0.79 0.36 / 0.36 0.31 / 0.31 0.85 / 0.85

SimpleDesign is stronger than ESM3 and competitive with or better than DPLM2 on codesignability. MultiFlow and La-proteina achieve higher structural consistency. The results also show a trade-off between consistency and diversity: lower qdataq_{\mathrm{data}}6 can produce higher codesignability while changing structural diversity behavior.

Ablation findings

The vanilla Transformer can outperform MoT in reported SwissProt-fine-tuned codesignability results. At qdataq_{\mathrm{data}}7, MoT obtains codesignability of qdataq_{\mathrm{data}}8, compared with qdataq_{\mathrm{data}}9 for the vanilla Transformer. At t∈[0,1]t\in[0,1]0, the corresponding values are t∈[0,1]t\in[0,1]1 and t∈[0,1]t\in[0,1]2; at t∈[0,1]t\in[0,1]3, they are t∈[0,1]t\in[0,1]4 and t∈[0,1]t\in[0,1]5.

SwissProt fine-tuning improves sequence–structure consistency but generally reduces FoldSeek diversity, indicating a quality–diversity trade-off.

Scope and limitations

SimpleDesign occupies an intermediate position between tokenized multimodal protein LLMs and specialized geometric flow systems. Its principal limitations are:

  • Geometric inductive bias: the architecture is not explicitly SE(3)-equivariant.
  • Structural representation: outputs contain only t∈[0,1]t\in[0,1]6 coordinates, not all-atom structures.
  • Structural fidelity: specialized geometric models remain stronger on some designability and consistency metrics.
  • Evaluation scope: the reported evaluation is computational and does not establish experimental folding, function, binding, stability, or safety.
  • Length and representation limits: the principal generation range is 100–500 residues, and the model generates backbone structures rather than complete molecular structures.
  • Diversity sensitivity: diversity varies with t∈[0,1]t\in[0,1]7, data source, and metric.
  • Biological validation: computational plausibility does not demonstrate biological activity.

The central result is therefore methodological rather than a claim of universal superiority. SimpleDesign shows that a multimodal Transformer can jointly model categorical sequence denoising and continuous structural flow without a separately trained structure vocabulary. This reduces pipeline complexity and avoids latent reconstruction error, while retaining competitive sequence–structure codesign performance. Specialized geometric models remain preferable when structural fidelity and designability are the dominant objectives.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SimpleDesign.