Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision
Josh Qixuan Sun, Morteza Babaie, Wenyang hou, Mark Crowley, David Young
ICML 2026 regular
Tóm tắt (nguồn: OpenReview · © tác giả)
Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that integrates scarce quantitative expression data with large-scale weak positive supervision from immunization data. We adapt Direct Preference Optimization (DPO) to protein language models by introducing a union-masked log-likelihood approximation and IMGT-based alignment, enabling efficient training on variable-length sequences. Evaluating on a diverse internal dataset of 1254 labeled sequences and 4 million unlabeled camelid-derived antibodies, we show that our method consistently outperforms baselines on most metrics. Our results demonstrate that preference learning can effectively learn from weak supervision, providing a scalable solution for antibody expressibility optimization in data-constrained settings. Project page: https://kisoji-biotechnology-inc.github.io/Preference-Expression-Ranking/.
Từ khoá
Metadata từ BioTender-max/icml2026-ai-bio (CC0-1.0). Phở không lưu trữ bản PDF; link trỏ về nguồn gốc.
Cùng chủ đề
Conditionally Site-Independent Neural Evolution of Antibody Sequences
Stephen Zhewen Lu, Aakarsh Vermani, Kohei Sanno, Jiarui Lu +3
Common deep learning approaches for antibody engineering focus on modeling the marginal distribution of sequences. By treating sequences as independent samples, however, these…
EpiCoCo: De Novo Epitope Generation via MHC-Context Co-Modeling and Contrastive Affinity Guidance
Haoyang Luan, Gufeng Yu, Letian Chen, Zhenran Xiao +3
The *de novo* generation of high-affinity epitopes tailored to specific major histocompatibility complex (MHC) proteins is a pivotal challenge in computational immunotherapy.…
Explicit representation of germline and non-germline residues improves antibody language modeling
Jeonghyeon Kim, Nathaniel Blalock, Ameya Kulkarni, Kensuke Nakamura +1
Antibodies originate from germline templates and are diversified by somatic hypermutation, producing sequences in which conserved germline residues scaffold structure while rare…