Language model driven: a PROTAC generation pipeline with dual\n constraints of structure and property
Preprint 2024 en
Authors
JS
Jinsong Shao
QG
Qiang Gong
ZY
Zeyu Yin
Abstract
1 min read
The imperfect modeling of ternary complexes has limited the application of\ncomputer-aided drug discovery tools in PROTAC research and development. In this\nstudy, an AI-assisted approach for PROTAC molecule design pipeline named\nLM-PROTAC was developed, which stands for language model driven Proteolysis\nTargeting Chimera, by embedding a transformer-based generative model with dual\nconstraints on structure and properties, referred to as the DCT. This study\nutilized the fragmentation representation of molecules and developed a language\nmodel driven pipeline. Firstly, a language model driven affinity model for\nprotein compounds to screen molecular fragments with high affinity for the\ntarget protein. Secondly, structural and physicochemical properties of these\nfragments were constrained during the generation process to meet specific\nscenario requirements. Finally, a two-round screening of the preliminary\ngenerated molecules using a multidimensional property prediction model to\ngenerate a batch of PROTAC molecules capable of degrading disease-relevant\ntarget proteins for validation in vitro experiments, thus achieving a complete\nsolution for AI-assisted PROTAC drug generation. Taking the tumor key target\nWnt3a as an example, the LM-PROTAC pipeline successfully generated PROTAC\nmolecules capable of inhibiting Wnt3a. The results show that DCT can\nefficiently generate PROTAC that targets and hydrolyses Wnt3a.\n
Discussion(0)
No comments yet. Be the first to comment.