FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications
Kieran Didi, Sarah Alamdari, Alex Xijie Lu, Bruce James Wittmann, Kadina E Johnston, Ava P Amini, Ali Madani, Maya Czeneszew
ICML 2026 spotlight
Abstract (source: OpenReview · © authors)
Machine learning methods that predict protein fitness from sequence remain sensitive to changes in data distributions, limiting generalization across common conditions encountered in protein engineering. Practically, protein engineers are thus left wondering about the effective utility of ML tools. The FLIP benchmark established protocols for testing generalization under some domain shifts, but it was limited to measurements of stability, binding, and viral capsid viability. We introduce FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns. Evaluating a suite of benchmark models across these datasets and suites reveals that simpler models often matched or outperformed fine-tuned protein language models on FLIP2, challenging the utility of existing transfer learning techniques. Provenance for all datasets has been recorded and we redistribute all data CC-BY 4.0 to facilitate continued progress.
Keywords
Metadata from BioTender-max/icml2026-ai-bio (CC0-1.0). Phở does not host any PDF; links point back to the source.
Related
LipoPU: Pocket-level Prediction of Lipid-Protein Interactions via Positive-Unlabeled Learning
Yuxing Wang, Wenyi Zhang, Yilong Zou, Jing Huang
Computational identification of lipid-binding proteins is critical for both fundamental research and therapeutic development. Existing models are typically trained in a fully…
MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation Effects
Ruihan Guo, Chaoran Cheng, Zhanghan Ni, Neil He +2
Protein mutation effect prediction is fundamental to protein engineering and disease variant interpretation, yet experimentally measured mutation data remain accurate but…
PPI Candidate Ranking: Large-Scale Evaluation of a Domain Knowledge–Guided Pipeline
Maria Emilia Russo, Federico Di Valerio, Alessia Borghini, Alessio Ragno +1
Computational approaches have become central to Protein–Protein Interaction (PPI) research, complementing experimental techniques that remain costly and incomplete. While modern…