When an antibody changes format: predicting synthesis before the next batch
Topics: Multimodal learning · Biomedical AI · Antibody engineering
This project grew out of my internship at BigHat Biosciences. I’m grateful to my coauthors and the wider BigHat team for their collaboration, including the DS/ML team’s helpful discussions and suggestions.
An antibody is a protein that recognizes and binds a specific molecular target. Engineers reuse its binding domains in different designs, but changing the format can alter folding, stability, or how well the protein can be produced. Before measuring the new design’s binding or other properties, we need enough usable protein to run the assays.
Our paper, Improved Therapeutic Antibody Reformatting through Multimodal Machine Learning, studies this bottleneck when converting full-length immunoglobulin G antibodies (IgGs) into single-chain variable fragments (scFvs). We presented the work at the NeurIPS 2025 AI4Science and FM4LS workshops. We ask whether sequence, predicted structure, and biophysical features can help identify which designs are likely to synthesize, especially for a starting antibody the model has never seen.
The surprising finding was that, for synthesis classification, logistic regression on simple one-hot sequence features outperformed the tested frozen protein-language-model embeddings with neural prediction heads. Adding explicit structural and biophysical information improved transfer to new antibody families further. This reformatting task shows how a small model with carefully chosen features can outperform a more complex predictor using pretrained representations.
Biology detour: what are we reformatting?
Biology vocabulary alert: this is where the acronyms multiply.
A full-length IgG has the familiar Y-shaped antibody architecture, built from heavy and light chains. The variable domains at its tips, VH and VL, work together to form a target-binding site. An scFv keeps this variable-domain pair and joins the two domains with a peptide linker to make a single chain. Here, reformatting means moving the starting antibody’s variable domains from the IgG architecture into this smaller scFv arrangement.
scFvs can support high-throughput antibody screening and become components of more complex therapeutic designs. Their smaller format changes the context in which the domains must fold and function, so successful production and binding have to be checked experimentally.
The full-length IgG is at the upper left, and the scFv is near the upper middle. We focus on converting IgG-derived variable domains into the linked, single-chain scFv format.
We predict two experimental outcomes for each proposed scFv:
- Synthesis pass/fail is a classification task. A logistic classifier predicts whether the construct will synthesize adequately.
- Synthesis yield is a regression task. A separate linear model predicts a numerical yield in ng/µL.
These predictions are intended to help prioritize which designs advance to the next experimental batch. The model scores proposed constructs; experimental assays establish their actual properties.
What must the model generalize to?
Changing the linker, domain order, or amino-acid sequence produces a different scFv construct. We identify each one by
\[\mathrm{signature}=(V_H,V_L,\mathrm{linker},\mathrm{orientation}).\]After aggregating repeated measurements of identical signatures, the dataset contains 1,477 constructs. Those derived from the same starting antibody form a parental family and often differ by only a few mutations. Predicting a new construct within a familiar family is therefore different from predicting outcomes for an entirely new family.
We separate these situations in our evaluation:
- scFv split: all families appear in training, but the test constructs have unseen signatures.
- Target-family split: training includes a small experimental batch from the new family; the model predicts its remaining constructs.
- Parental-family split: entire families are held out, so the model receives no experimental training data from them.
The last split tests the hardest deployment question: can the model help before the first batch from a new antibody has been measured?
The amount of experimental information available for a new antibody determines the appropriate split. Panels (a), (b), and (c) show parental-family, target-family, and scFv splits.
Representing both the sequence and its new context
Sequence features describe the design directly. We align VH and VL using the AHo numbering scheme, one-hot encode their amino acids, and append encodings of the linker and domain order. Alignment lets the model compare corresponding residue positions across antibodies.
Structural features describe how the domains may change after reformatting. We predict both the parental IgG and scFv structures with Boltz-2, align them, and concatenate their residue-level Cα coordinates with gap indicators. This retains local spatial detail. Global root-mean-square deviation (RMSD) provides a coarser summary of the structural difference.
Biophysical features add properties such as surface hydrophobicity and charge patches, computed from the predicted structures using NaturalAntibody. These descriptors give the classifier information about the protein’s physical context before experimental measurements are available.
We concatenate the features into $z=[x_{\mathrm{seq}};x_{\mathrm{struct}};x_{\mathrm{bio}}]$ and fit a regularized logistic classifier:
\[\Pr(y=1\mid z)=\sigma(b+\beta^\top z), \qquad \sigma(a)=\frac{1}{1+e^{-a}}.\]Here $y$ is the binary synthesis-outcome label, and the learned weights $\beta$ combine evidence across the feature groups. Separate linear regressors predict synthesis yield. All representations are precomputed and frozen, allowing us to test their value with a simple downstream model.
The model combines residue-level sequence and geometry with broader physical descriptors.
What screening would change in one measured family
Our Fam1 case study makes the prediction task concrete. Its held-out set contains 55 variants, of which 34 pass synthesis. The classifier retains all 34 successful variants and also accepts five failures.
In this retrospective example, using our model to screen the batch would have avoided 16 of 55 synthesis attempts, about 29%, while retaining every successful variant. The fraction of successful experiments would rise from
\[\frac{34}{55}=61.8\% \quad\text{to}\quad \frac{34}{39}=87.2\%.\]That illustrates the value of screening before synthesis: fewer unsuccessful experiments and more wet-lab capacity for promising designs.
The largest gains appear across families
Across ten random folds of the parental-family split, the multimodal classifier reaches 88.92 ± 14.93 AUROC, compared with 66.35 ± 10.73 for sequence-only logistic regression. AUPRC rises from 59.21 ± 15.65 to 85.68 ± 20.94. These scores use a 0–100 scale, and the reported variation is the standard deviation across folds. The improvement is substantial, although performance varies widely with the held-out families. Table 2 reports the full comparison.
Within known families, the AUROC gain is smaller: 89.46 to 92.93. This difference is important. The added features are most useful when the model must transfer beyond the sequence context already represented in training. With a small pilot batch from Fam2, AUROC also rises from 71.81 to 82.96.
Multimodal features improve the mean classification scores in every reported split, with the largest gains for held-out parental families. Fam1–Fam3 are target-family settings with a small pilot batch available; the ± values show variability across folds.
Frozen protein-language-model embeddings do not outperform the simple one-hot baseline in our experiments. On the parental-family split, AbLang embeddings with an MLP reach 62.58 AUROC, compared with the one-hot baseline’s 66.35. In this comparison, we evaluate fixed representations with trained prediction heads; we do not test end-to-end fine-tuning.
The one-hot sequence baseline leads the tested frozen-embedding models under both splits. Structural coordinates alone perform poorly, motivating a combination of sequence and structure.
In the modality ablation in Appendix E.1, we specifically vary sequence features, residue-level structural coordinates, and global RMSD. Sequence plus coordinates performs best, reaching 89.2 AUROC and 86.3 AUPRC. RMSD alone has an AUROC of about 50.0, and adding it to sequence plus coordinates changes AUROC from 89.2 to 89.0. The useful complementarity is between sequence context and detailed geometry.
Sequence plus structure is the strongest combination across these three metrics. Global RMSD adds little once the residue-level coordinates are present; error bars show standard deviations.
Our family-level analysis in Appendix E.2 helps explain why a single distance measure can be inadequate: within some families, greater structural deviation correlates with higher yield; within others, the relationship reverses. Sequence context and residue-level geometry preserve information that an overall displacement score loses.
What remains to be tested
With 1,477 constructs, our dataset leaves limited scope for training higher-capacity models. Extending it would help us test how the gains change with more experimental data. We primarily study synthesis and yield here, with an additional SEC-purity analysis in the appendix; binding and therapeutic performance remain separate questions.
Our main finding is that explicit sequence and structural context can help a lightweight model transfer to unseen antibody families. Equally important, the deployment-aware splits make that transfer visible. A model intended for a new antibody program needs to be evaluated on that new-family question.
Reference
Jiayi Xin, Aniruddh Raghu, Nick Bhattacharya, Adam Carr, Melanie Montgomery, and Hunter Elliott. 2025. Improved Therapeutic Antibody Reformatting through Multimodal Machine Learning. arXiv:2509.19604. NeurIPS 2025 AI4Science Workshop and Multi-modal Foundation Models and Large Language Models for Life Sciences Workshop.
The format overview, split diagram, model pipeline, and modality ablation are Figures 1, 2, 5, and 10 from our paper. The main classification and frozen-embedding comparisons are Tables 2 and 1.






