Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Annotation-guided Protein Design with Multi-Level Domain Alignment (2404.16866v4)

Published 18 Apr 2024 in q-bio.QM, cs.AI, and cs.LG

Abstract: The core challenge of de novo protein design lies in creating proteins with specific functions or properties, guided by certain conditions. Current models explore to generate protein using structural and evolutionary guidance, which only provide indirect conditions concerning functions and properties. However, textual annotations of proteins, especially the annotations for protein domains, which directly describe the protein's high-level functionalities, properties, and their correlation with target amino acid sequences, remain unexplored in the context of protein design tasks. In this paper, we propose Protein-Annotation Alignment Generation, PAAG, a multi-modality protein design framework that integrates the textual annotations extracted from protein database for controllable generation in sequence space. Specifically, within a multi-level alignment module, PAAG can explicitly generate proteins containing specific domains conditioned on the corresponding domain annotations, and can even design novel proteins with flexible combinations of different kinds of annotations. Our experimental results underscore the superiority of the aligned protein representations from PAAG over 7 prediction tasks. Furthermore, PAAG demonstrates a significant increase in generation success rate (24.7% vs 4.7% in zinc finger, and 54.3% vs 22.0% in the immunoglobulin domain) in comparison to the existing model. We anticipate that PAAG will broaden the horizons of protein design by leveraging the knowledge from between textual annotation and proteins.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Chaohao Yuan (5 papers)
  2. Songyou Li (3 papers)
  3. Geyan Ye (4 papers)
  4. Yikun Zhang (18 papers)
  5. Long-Kai Huang (14 papers)
  6. Wenbing Huang (95 papers)
  7. Wei Liu (1135 papers)
  8. Jianhua Yao (50 papers)
  9. Yu Rong (146 papers)
Citations (1)

Summary

We haven't generated a summary for this paper yet.