Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023).
Hayes, T. et al. Simulating 500 million years of evolution with a language model. Science 387, 850–858 (2025).
Su, J. et al. Democratizing protein language model training, sharing and collaboration. Nat. Biotechnol. https://doi.org/10.1038/s41587-025-02859-7 (2025).
Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023).
Alamdari, S. et al. Protein generation with evolutionary diffusion: sequence is all you need. Preprint at bioRxiv https://doi.org/10.1101/2023.09.11.556673 (2024).
Lin, Y. & AlQuraishi, M. Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds. In Proc. 40th International Conference on Machine Learning 20978–21002 (JMLR, 2023).
Lin, Y., Lee, M., Zhang, Z. & AlQuraishi, M. Out of many, one: designing and scaffolding proteins at the scale of the structural universe with Genie 2. Preprint at https://doi.org/10.48550/arxiv.2405.15489 (2024).
Lorch, M. in Biochemistry: A Very Short Introduction 34–51 (Oxford Univ. Press, 2021).
Madani, A. et al. Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 41, 1099–1106 (2023).
Pokharel, S., Pratyush, P., Heinzinger, M., Newman, R. H. & Kc, D. B. Improving protein succinylation sites prediction using embeddings from protein language model. Sci. Rep. 12, 16933 (2022).
Brandes, N., Ofer, D., Peleg, Y., Rappoport, N. & Linial, M. ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics 38, 2102–2110 (2022).
Sledzieski, S., Singh, R., Cowen, L. & Berger, B. D-SCRIPT translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein–protein interactions. Cell Syst. 12, 969–982.e6 (2021).
Singh, R., Devkota, K., Sledzieski, S., Berger, B. & Cowen, L. Topsy-Turvy: integrating a global view into sequence-based PPI prediction. Bioinformatics 38, i264–i272 (2022).
Singh, R., Sledzieski, S., Bryson, B., Cowen, L. & Berger, B. Contrastive learning in protein language space predicts interactions between drugs and protein targets. Proc. Natl Acad. Sci. USA 120, e2220778120 (2023).
Bhat, S. et al. De novo design of peptide binders to conformationally diverse targets with contrastive language modeling. Sci. Adv. 11, eadr8638 (2025).
Frey, N. C. et al. Protein discovery with discrete walk-jump sampling. Preprint at https://doi.org/10.48550/arxiv.2306.12360 (2023).
Cohen, T. & Schneidman-Duhovny, D. Epitope-specific antibody design using diffusion models on the latent space of ESM embeddings. In Workshop on Generative and Experimental Perspectives for Biomolecular Design https://openreview.net/pdf?id=r561kIH4lE (ICLR, 2024).
Grimmett, G. R. & Stirzaker, D. R. Probability and Random Processes (Oxford Univ. Press, 2001).
Serfling, R. J. Contributions to central limit theory for dependent variables. Ann. Math. Statist. 39, 1158–1175 (1968).
Song, J., Meng, C. & Ermon, S. Denoising diffusion implicit models. In International Conference on Learning Representations https://openreview.net/pdf?id=St1giarCHLP (ICLR, 2023).
Song, Y., Durkan, C., Murray, I. & Ermon, S. Maximum likelihood training of score-based diffusion models. In Proc. 35th International Conference on Neural Information Processing Systems 1415–1428 (2021).
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M. & Le, M. Flow matching for generative modeling. In International Conference on Learning Representations https://openreview.net/pdf?id=PqvMRDCJT9t (ICLR, 2023).
Shcherbakova, D. M. & Verkhusha, V. V. Chromophore chemistry of fluorescent proteins controlled by light. Curr. Opin. Chem. Biol. 20, 60–68 (2014).
Jani, V., Sonavane, U. & Joshi, R. Insight into structural dynamics involved in activation mechanism of full length KRAS wild type and P-loop mutants. Heliyon 10, e36161 (2024).
Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024).
Wilson, C. J., Choy, W.-Y. & Karttunen, M. AlphaFold2: a role for disordered protein/region prediction? IJMS 23, 4591 (2022).
Mariani, V., Biasini, M., Barbato, A. & Schwede, T. lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests. Bioinformatics 29, 2722–2728 (2013).
Zhang, Y. TM-align: a protein structure alignment algorithm based on the TM-score. Nucleic Acids Res. 33, 2302–2309 (2005).
Wu, R. et al. High-resolution de novo structure prediction from primary sequence. Preprint at bioRxiv https://doi.org/10.1101/2022.07.21.500999 (2022).
Passaro, S. et al. Boltz-2: towards accurate and efficient binding affinity prediction. Preprint at bioRxiv https://doi.org/10.1101/2025.06.14.659707 (2025).
Bateman, A. The Pfam protein families database. Nucleic Acids Res. 32, 138D–141D (2004).
Yu, T. et al. Enzyme function prediction using contrastive learning. Science 379, 1358–1363 (2023).
Shaner, N. C., Patterson, G. H. & Davidson, M. W. Advances in fluorescent protein technology. J. Cell Sci. 120, 4247–4260 (2007).
Rappoport, J. Z. & Simon, S. M. A functional GFP fusion for imaging clathrin-mediated endocytosis. Traffic 9, 1250–1255 (2008).
Skube, S. B., Chaverri, J. M. & Goodson, H. V. Effect of GFP tags on the localization of EB1 and EB1 fragments in vivo. Cytoskeleton 67, 1–12 (2010).
Zhou, Z. K., Hong, K. Huang, B. & Narlikar, G. J. Understanding how genetically encoded tags affect phase separation by heterochromatin protein HP1α. Cell Rep. Methods 5, 101029 (2025).
Cubitt, A. B. et al. Understanding, improving and using green fluorescent proteins. Trends Biochem. Sci. 20, 448–455 (1995).
Rodriguez, E. A. et al. The growing and glowing toolbox of fluorescent and photoactive proteins. Trends Biochem. Sci. 42, 111–129 (2017).
Sarkisyan, K. S. et al. Local fitness landscape of the green fluorescent protein. Nature 533, 397–401 (2016).
Lambert, T. J. FPbase: a community-editable fluorescent protein database. Nat. Methods 16, 277–278 (2019).
Barondeau, D. P., Putnam, C. D., Kassmann, C. J., Tainer, J. A. & Getzoff, E. D. Mechanism and energetics of green fluorescent protein chromophore synthesis revealed by trapped intermediate structures. Proc. Natl Acad. Sci. USA 100, 12111–12116 (2003).
Kim, D. I. et al. An improved smaller biotin ligase for BioID proximity labeling. Mol. Biol. Cell 27, 1188–1196 (2016).
Branon, T. C. et al. Efficient proximity labeling in living cells and organisms with TurboID. Nat. Biotechnol. 36, 880–887 (2018).
Kubitz, L. et al. Engineering of ultraID, a compact and hyperactive enzyme for proximity-dependent biotinylation in living cells. Commun. Biol. 5, 657 (2022).
Pudžiuvelytė, I. et al. TemStaPro: protein thermostability prediction using sequence representations from protein language models. Bioinformatics 40, btae157 (2024).
Cotet, T.-S. et al. Crowdsourced protein design: lessons From the Adaptyv EGFR binder competition. Preprint at bioRxiv https://doi.org/10.1101/2025.04.17.648362 (2025).
Su, J. et al. A trimodal protein language model enables advanced protein searches. Nat. Biotechnol. https://doi.org/10.1038/s41587-025-02836-0 (2025).
Geffner, T. et al. La-Proteina: atomistic protein generation via partially latent flow matching. Preprint at https://doi.org/10.48550/arxiv.2507.09466 (2025).
Wang, X. et al. Diffusion language models are versatile protein learners. In Proc. 41st International Conference on Machine Learning 52309–52333 (JMLR, 2024).
Hie, B. L. et al. Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol. 42, 275–283 (2024).
Devkota, K. Raygun benchmarking + training dataset. Zenodo https://doi.org/10.5281/zenodo.19546626 (2026).

