High-Throughput Prediction of Protein–Protein Interactions Uncovers Hidden Molecular Networks in Biosynthetic Gene Clusters
Source: PubMed Central Open Access, NCBI / U.S. National Library of Medicine
Biosynthetic gene clusters (BGCs) are contiguous genomic regions that encode diverse proteins responsible for natural product biosynthesis. These proteins collectively produce various secondary metabolites with complex chemical structure, including antibiotics and mycotoxins, yet the complete biosynthetic pathways have been experimentally elucidated for only a limited number of compounds. Recently, protein–protein interactions within BGCs have been recognized as key determinants of intermediate transfer, enzymatic regulation, and structural stability. However, many BGCs still contain proteins of unknown function that cannot be predicted by conventional sequence-based bioinformatics tools, hindering a comprehensive understanding of their biosynthetic pathways. To address this challenge, we built a high-throughput complex prediction pipeline by replacing AlphaFold3’s multiple sequence alignment generation with a faster MMSeqs2. We systematically screened 487,828 protein pairs derived from 2,437 BGCs registered in the Minimum Information about a Biosynthetic Gene cluster database and predicted 15,438 heteromeric interactions with an interface predicted TM score ≥ 0.6. Among them, 1,390 protein pairs exhibited structural homology with a root mean square deviation ≤ 2.0 Å. Our analysis demonstrates that AF3-based complex prediction matches experimental results with high confidence for most proteins and reveals many uncharacterized but novel heterocomplexes within BGCs. These findi
Abstract
Biosynthetic gene clusters (BGCs) are contiguous genomic regions that encode diverse proteins responsible for natural product biosynthesis. These proteins collectively produce various secondary metabolites with complex chemical structure, including antibiotics and mycotoxins, yet the complete biosynthetic pathways have been experimentally elucidated for only a limited number of compounds. Recently, protein–protein interactions within BGCs have been recognized as key determinants of intermediate transfer, enzymatic regulation, and structural stability. However, many BGCs still contain proteins of unknown function that cannot be predicted by conventional sequence-based bioinformatics tools, hindering a comprehensive understanding of their biosynthetic pathways. To address this challenge, we built a high-throughput complex prediction pipeline by replacing AlphaFold3’s multiple sequence alignment generation with a faster MMSeqs2. We systematically screened 487,828 protein pairs derived from 2,437 BGCs registered in the Minimum Information about a Biosynthetic Gene cluster database and predicted 15,438 heteromeric interactions with an interface predicted TM score ≥ 0.6. Among them, 1,390 protein pairs exhibited structural homology with a root mean square deviation ≤ 2.0 Å. Our analysis demonstrates that AF3-based complex prediction matches experimental results with high confidence for most proteins and reveals many uncharacterized but novel heterocomplexes within BGCs. These findings will facilitate experimental verification of unidentified enzymatic reactions leading to the final products.
