SimpleDesign: Multimodal Protein Design Model
- SimpleDesign operates directly in data space, modeling protein sequences with categorical discrete denoising and structures with continuous velocity regression.
- The model utilizes a shared Transformer with modality-specific processing, generating over 1.8 million sequence-structure pairs and achieving competitive performance in sequence generation, structural generation, and joint codesign.
- Despite not overtaking specialized geometric models in structural consistency, SimpleDesign offers a simplified, efficient approach that minimizes complexity and error propagation.
SimpleDesign is a multimodal generative model for joint protein sequence–structure codesign. It generates amino-acid sequences and three-dimensional protein backbones within a single-stage, end-to-end framework operating directly in data space rather than through discrete sequence or structure tokenizers. The model combines masked discrete denoising for sequences with continuous velocity regression for structures and uses a shared Transformer with modality-specific processing. Trained on more than 2 million sequence–structure pairs, SimpleDesign achieves competitive codesignability and sequence-generation performance relative to tokenized multimodal protein LLMs, while remaining weaker than specialized geometric flow models on some structural-consistency benchmarks (Lu et al., 3 Sep 2026).
1. Motivation and design objective
Protein function depends on the interaction between amino-acid sequence and three-dimensional structure. A generative model for protein design must therefore represent both categorical sequence information and continuous geometric information, while maintaining consistency between the two modalities.
Many multimodal protein-generation systems use a multi-stage pipeline. Sequence and structure autoencoders or tokenizers are first trained to produce latent representations; a generative model is then trained over those representations; finally, generated latent codes are decoded back into sequences or structures. This design introduces several sources of complexity:
- Error propagation: the generative model learns the distribution of imperfect latent reconstructions rather than the original data.
- Information loss: continuous structural coordinates must be compressed into a finite vocabulary.
- Objective misalignment: tokenizer reconstruction objectives may not coincide with sequence–structure generation objectives.
- Additional engineering: modality-specific tokenizers, decoders, vocabularies, and pretraining stages are required.
- Weak end-to-end coupling: sequence and structure consistency is mediated through separately learned representations.
SimpleDesign addresses these issues by modeling amino-acid sequences and backbone coordinates directly. Its main hypothesis is that structure tokenization is not necessary for competitive sequence–structure codesign. The model uses one joint training objective, one multimodal Transformer trunk, and two modality-specific prediction objectives:
- masked discrete prediction for amino-acid sequences;
- continuous velocity regression for structural coordinates.
The model represents a protein sequence as
where is the residue count and is the amino-acid vocabulary. The structure is represented by coordinates
where is the three-dimensional coordinate of residue . The target is the joint distribution
where denotes the empirical distribution of paired sequences and structures.
2. Joint corruption and training objectives
SimpleDesign uses independent corruption processes for sequence and structure. The sequence corruption time is , while the structural interpolation time is 0. These variables are sampled independently, allowing the model to represent folding-like, inverse-folding-like, and joint codesign regimes.
When 1 and 2, the sequence is mostly visible while the structure is highly noisy, approximating structure prediction from sequence. When 3 and 4, the structure is visible while the sequence is highly masked, approximating inverse folding. Intermediate values represent joint sequence–structure denoising.
Sequence denoising
For each residue, a Bernoulli variable determines whether the amino acid is masked:
5
The corrupted sequence is
6
The sequence head predicts a categorical distribution over the amino-acid vocabulary at masked positions. The masked cross-entropy objective is
7
with 8. The mask indicator restricts the loss to corrupted positions, while the denominator prevents division by zero when no position is masked. The factor 9 downweights examples with high masking rates in the reported formulation.
Structural velocity regression
Structures are perturbed using a linear interpolation between Gaussian noise and the clean coordinates:
0
1
The ideal velocity along this path is
2
The structural objective is
3
where 4 is the predicted structural velocity. During generation, numerical integration of this vector field transports Gaussian noise toward the protein-structure distribution.
Because protein structures are invariant under global translation and rotation, SimpleDesign applies Kabsch alignment before constructing structural targets. The procedure centers predicted and reference coordinates, computes a covariance matrix, obtains a singular-value decomposition, constructs the optimal rotation, corrects reflections when necessary, and aligns the reference structure. This prevents coordinate-level error from penalizing equivalent structures that differ only by rigid transformation.
Single-stage multimodal objective
For paired data 5, the total loss combines the sequence and structure objectives:
6
The reported implementation uses 7. The sequence time is sampled uniformly:
8
whereas the structural time uses
9
This distribution emphasizes later, less-noisy structural states while retaining some coverage of highly corrupted structures.
3. Architecture
Direct continuous structure representation
SimpleDesign does not convert structures into discrete tokens. Each residue coordinate 0 is passed through Fourier feature encoding, followed by a linear projection and layer normalization:
1
The sequence stream uses a learnable amino-acid embedding:
2
For a protein of length 3, the sequence and structure representations are concatenated into a joint stream:
4
The Transformer therefore receives 5 residue-level representations. Sequence and structure tokens share residue indices, providing an alignment signal for
6
The model uses additive sinusoidal positional encodings and rotary positional embeddings inside attention layers. Global self-attention allows sequence representations to attend to structural representations and vice versa, without requiring a separate cross-attention module.
Mixture-of-Transformer
The default architecture is a Mixture-of-Transformer (MoT). MoT retains global joint self-attention while assigning modality-specific parameters to the two streams:
- sequence-specific query, key, and value projections;
- structure-specific query, key, and value projections;
- modality-specific layer normalization;
- modality-specific feed-forward networks.
For modality 7, the projections can be written as
8
The resulting representations participate in shared attention:
9
MoT is intended to balance two extremes. A completely shared Transformer may underrepresent the distinction between categorical and continuous modalities, while two independent networks would lose direct cross-modal reasoning. MoT provides modality-specific processing while preserving joint attention.
The ablation results indicate that MoT is not essential to the central contribution. A vanilla shared Transformer is highly competitive and sometimes performs better. Accordingly, the principal methodological contribution is the tokenizer-free joint objective rather than the MoT parameterization itself.
Output heads
The sequence head is an MLP followed by layer normalization and a projection to the amino-acid vocabulary:
0
1
The final sequence projection is tied to the input amino-acid embedding matrix.
The structure head uses an MLP with adaptive layer normalization conditioned on the structural time 2:
3
The resulting representation is projected into three-dimensional velocity vectors:
4
The architecture is not explicitly SE(3)-equivariant. Instead, the implementation uses random rigid rotations and translations during training and Kabsch alignment for structural supervision.
4. Data, initialization, and optimization
SimpleDesign is trained primarily on AFESM, a resource derived from the AlphaFold Database and the ESM Metagenomic Atlas. The source resource contains more than 800 million predicted structures. The authors cluster sequences and structures into approximately 5 million non-singleton structural clusters and retain one representative from each cluster.
The training set is restricted to proteins with lengths satisfying
5
and pLDDT greater than $85. No filtering is performed based on secondary-structure or coil content. The resulting dataset contains 1,807,333 training proteins, with 1,000 proteins held out for validation.
The model is subsequently fine-tuned on a filtered AFDB-derived SwissProt subset containing 442,511 high-quality sequence–structure pairs.
The Transformer and sequence embeddings are initialized from ESM2-650M weights. For MoT, sequence-side attention, normalization, and feed-forward components can be initialized from ESM2, while structure-specific components are randomly initialized.
Reported optimization settings include:
- Optimizer: AdamW.
- Learning rate: $x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},$6 for the standard Transformer and $x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},$7 for MoT.
- Weight decay: none.
- Warm-up: 5,000 steps from $x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},$8 to the target learning rate.
- Gradient clipping: norm $x=(x^{(1)},\ldots,x^{(L)})\in\mathbb R^{L\times 3},$9.
- Training duration: 300,000 AFESM steps.
- Fine-tuning: 50,000 SwissProt steps.
- Hardware: 64 NVIDIA H100 80GB GPUs.
- Effective outer batch size: 128.
- Gradient accumulation: 2.
- Inner replica batch size: 16.
Coordinates are rescaled using
$x^{(i)}$0
Samples are repeatedly presented under different masking, noising, and rigid transformations.
5. Inference and generation
Sequence generation
Sequence generation begins with a fully masked sequence. At each denoising step, the model produces logits $x^{(i)}$1 for every residue. Special-token logits are suppressed. Gumbel noise is added:
$x^{(i)}$2
where
$x^{(i)}$3
and $x^{(i)}$4.
Temperature scaling is then applied, with the temperature annealed approximately from $x^{(i)}$5 to $x^{(i)}$6. Tokens are sampled categorically, and the highest-confidence positions are progressively unmasked. Residues whose frequency exceeds $x^{(i)}$7 are remasked and resampled to prevent domination by a single amino acid.
Structure generation
Structure generation initializes coordinates with Gaussian noise:
$x^{(i)}$8
The model’s velocity field is converted into a score-like quantity:
$x^{(i)}$9
A stochastic flow is simulated using an Euler–Maruyama discretization:
$i$0
where
$i$1
and $i$2 controls stochasticity. Coordinates are centered and randomly rotated during sampling, then rescaled to their original units. The generated structure is a $i$3-only backbone rather than an all-atom structure.
Joint codesign
Joint generation begins with a masked sequence and Gaussian structural coordinates. The sequence and structure streams are repeatedly passed through the joint Transformer. At each iteration, the model:
- predicts sequence logits;
- samples and selectively unmasks sequence residues;
- predicts structural velocity;
- advances the coordinates by one integration step;
- repeats until both modalities approach their clean states.
The sequence time follows an approximately linear schedule, while the structural time uses a log-spaced schedule with more steps near the clean-data endpoint. The reported inference parameter $i$4 controls a fidelity–diversity trade-off. Sequence-only results use $i$5, joint codesign results use $i$6 and $i$7, and structure-only results commonly use $i$8.
The mathematically described generation path begins near $i$9, corresponding to a masked sequence and Gaussian structure, and proceeds toward $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$0, the clean data state.
6. Evaluation and limitations
The main benchmark settings generate 100 samples at lengths
$p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$1
Unconditional sequence generation
Sequence generation is evaluated using ProGen2 perplexity, ESMFold pLDDT, MMseqs2 diversity, and sequence novelty relative to SwissProt.
| Method | PPL $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$2 | pLDDT $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$3 | MMseqs diversity | Novelty |
|---|---|---|---|---|
| EvoDiff | $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$4 | $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$5 | 1.00 | 0.49 |
| DPLM | $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$6 | $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$7 | 0.82 | 0.49 |
| ESM3, sequence to structure | $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$8 | $p_\theta(a,x)\approx q_{\mathrm{data}}(a,x),$9 | 0.58 | 0.45 |
| DPLM2 | $q_{\mathrm{data}}$0 | $q_{\mathrm{data}}$1 | 0.56 | 0.90 |
| SimpleDesign | $q_{\mathrm{data}}$2 | $q_{\mathrm{data}}$3 | 0.50 | 0.80 |
SimpleDesign produces protein-like and foldable sequences, although it is not the best sequence generator on every metric. Progressive structural realization can constrain sequence diversity.
Unconditional structure generation
Structures are evaluated after inverse folding with ProteinMPNN and refolding with ESMFold. The protocols PMPNN1 and PMPNN8 use one and eight inverse-folded sequences, respectively. Metrics include designability under scRMSD and scTM thresholds, pairwise TM-score similarity, FoldSeek cluster diversity, and novelty.
| Method | PMPNN1 designability | PMPNN8 designability | PMPNN8 TM similarity | PMPNN8 FoldSeek diversity |
|---|---|---|---|---|
| ESM3 sequence to structure | 0.17 / 0.19 | 0.24 / 0.27 | 0.39 / 0.34 | 0.41 / 0.50 |
| DPLM2 | 0.31 / 0.48 | 0.52 / 0.66 | 0.28 / 0.27 | 0.47 / 0.44 |
| SimpleDesign | 0.44 / 0.63 | 0.60 / 0.78 | 0.29 / 0.30 | 0.27 / 0.23 |
| MultiFlow | 0.86 / 0.90 | 0.95 / 0.98 | 0.33 / 0.33 | 0.52 / 0.52 |
| La-proteina, no triangle | 0.84 / 0.86 | 0.95 / 0.97 | 0.33 / 0.32 | 0.61 / 0.61 |
SimpleDesign substantially outperforms the tokenized multimodal protein-language-model baselines for structural generation, but specialized geometric models remain stronger in designability.
Sequence–structure codesign
Codesignability measures the proportion of generated sequence–structure pairs whose folded sequence agrees with the generated structure.
| Method | Codesignability | TM similarity | FoldSeek diversity | Novelty |
|---|---|---|---|---|
| ESM3 sequence to structure | 0.09 / 0.11 | 0.30 / 0.29 | 0.59 / 0.61 | 0.91 |
| DPLM2 | 0.30 / 0.46 | 0.29 / 0.28 | 0.51 / 0.39 | 0.95 / 0.96 |
| SimpleDesign, 4 | 0.53 / 0.74 | 0.31 / 0.30 | 0.18 / 0.14 | 0.97 / 0.97 |
| SimpleDesign, 5 | 0.36 / 0.55 | 0.29 / 0.30 | 0.30 / 0.26 | 0.98 / 0.97 |
| MultiFlow | 0.76 / 0.80 | 0.34 / 0.34 | 0.54 / 0.52 | 0.83 / 0.83 |
| La-proteina, triangle | 0.77 / 0.79 | 0.36 / 0.36 | 0.31 / 0.31 | 0.85 / 0.85 |
SimpleDesign is stronger than ESM3 and competitive with or better than DPLM2 on codesignability. MultiFlow and La-proteina achieve higher structural consistency. The results also show a trade-off between consistency and diversity: lower 6 can produce higher codesignability while changing structural diversity behavior.
Ablation findings
The vanilla Transformer can outperform MoT in reported SwissProt-fine-tuned codesignability results. At 7, MoT obtains codesignability of 8, compared with 9 for the vanilla Transformer. At 0, the corresponding values are 1 and 2; at 3, they are 4 and 5.
SwissProt fine-tuning improves sequence–structure consistency but generally reduces FoldSeek diversity, indicating a quality–diversity trade-off.
Scope and limitations
SimpleDesign occupies an intermediate position between tokenized multimodal protein LLMs and specialized geometric flow systems. Its principal limitations are:
- Geometric inductive bias: the architecture is not explicitly SE(3)-equivariant.
- Structural representation: outputs contain only 6 coordinates, not all-atom structures.
- Structural fidelity: specialized geometric models remain stronger on some designability and consistency metrics.
- Evaluation scope: the reported evaluation is computational and does not establish experimental folding, function, binding, stability, or safety.
- Length and representation limits: the principal generation range is 100–500 residues, and the model generates backbone structures rather than complete molecular structures.
- Diversity sensitivity: diversity varies with 7, data source, and metric.
- Biological validation: computational plausibility does not demonstrate biological activity.
The central result is therefore methodological rather than a claim of universal superiority. SimpleDesign shows that a multimodal Transformer can jointly model categorical sequence denoising and continuous structural flow without a separately trained structure vocabulary. This reduces pipeline complexity and avoids latent reconstruction error, while retaining competitive sequence–structure codesign performance. Specialized geometric models remain preferable when structural fidelity and designability are the dominant objectives.