Leveraging latent space models for enzyme discovery and sampling
2026-09-01
The exponential growth in available protein sequence data has broadened enzyme discovery opportunities but simultaneously highlighted a significant gap between sequence and function information. Traditional tools like phylogenetic trees and sequence similarity networks (SSNs) are widely adopted for sampling enzymes for novel transformations. However, their utility suffers from inherent limitations, which are exacerbated for large enzyme families. Phylogenetic trees, while useful for studying evolutionary relationships, become computationally intensive and difficult to visualize for larger protein datasets. SSNs, on the other hand, are sensitive to user-defined thresholds for sequence clustering and easily fail to capture more distant relationships between clusters. Additionally, both tools are alignment based and cannot capture higher-order interactions between residues. In this study, we address these limitations by optimizing a variational autoencoder (VAE)-based latent space model to visualize and explore enzyme sequence–function landscapes. By training our models on simulated datasets and real enzyme families, such as cyclases and flavin-dependent monooxygenases (FDMOs), we demonstrated that the optimized latent space effectively preserves phylogenetic relationships and enables high-resolution clustering for functionally distinct enzymes. The models further outperform traditional SSNs in capturing local and global relationships in a continuous two-dimensional space, enabling the discovery of multiple uncharacterized FDMOs for oxidative dearomatization and decarboxylative hydroxylation that illustrates their application. Our findings show that low-dimensional latent spaces can serve as valuable tools for enzyme discovery, allowing for interpolation and extrapolation to guide novel enzyme sampling for biocatalytic reactions.