Your AlphaFold model is blue from end to end. Median pLDDT is 92, the core looks compact, and the active-site residues sit exactly where the literature says they should. Then the construct expresses into inclusion bodies, the soluble fraction aggregates during concentration, and the purified material has no activity. It feels as if the model and the experiment are describing different proteins.
They are not. They are answering different questions.
pLDDT estimates local coordinate accuracy conditional on the model’s predicted structure. Your bench experiment measures whether a particular sequence, construct, host, buffer, cofactor state, oligomeric state, and concentration produce a usable molecule. High pLDDT can support a structural hypothesis. It cannot certify expression, solubility, activity, monodispersity, or the biological state you captured.
Key Takeaways
- High pLDDT means “locally self-consistent coordinates,” not “bench-ready protein.” It does not score expression yield, folding kinetics, aggregation, activity, or formulation stability.
- Read pLDDT together with PAE, model-to-model agreement, sequence liabilities, and biological context. A confident domain can still be placed incorrectly or represent only one state.
- Most high-confidence/poor-behavior mismatches fall into six buckets: kinetic folding failure, exposed aggregation patches, missing partners or cofactors, wrong oligomer/state, construct-context problems, and assay/formulation mismatch.
- A failed purification does not falsify the fold. It often falsifies the construct or experimental environment.
- Use orthogonal experiments matched to the claim. SEC-MALS tests assembly, DSF or nanoDSF tests thermal transitions, CD tests secondary structure, activity assays test function, and HDX-MS/SAXS/NMR test dynamics or ensembles.
- Do not “fix” every failure with stabilizing mutations. First determine whether you are stabilizing the wrong state or masking a missing biological dependency.
What pLDDT Actually Certifies
AlphaFold’s predicted Local Distance Difference Test score, or pLDDT, estimates how accurately a residue’s local atomic neighborhood would match an experimental structure. In AlphaFold’s original validation, pLDDT correlated with experimental local accuracy across held-out structures; values above 90 were associated with high local backbone and side-chain accuracy, while values above 70 generally indicated a correct backbone (Jumper et al., 2021; Tunyasuvunakool et al., 2021).
That is a powerful result. It is also narrower than the way pLDDT is often used.
pLDDT does not directly estimate:
- whether the protein expresses in your host;
- whether it reaches the predicted fold on a useful timescale;
- whether it remains soluble at 5 or 20 mg/mL;
- whether an exposed hydrophobic patch drives self-association;
- whether a cofactor, membrane, ligand, PTM, or binding partner is required;
- whether the model represents the active, inactive, apo, holo, open, or closed state;
- whether your chosen termini preserve the native domain;
- whether the protein survives the purification buffer.
The central category error is using a structure-accuracy confidence metric as a developability score.
Read Four Layers, Not One Color Scale
Before you compare the model with the bench, separate four layers of evidence.
| Layer | Question | Useful readout | What it cannot establish alone |
|---|---|---|---|
| Local structure | Is this local fold plausible? | pLDDT, side-chain environment | Domain placement, dynamics, activity |
| Global arrangement | Are domains or chains positioned reliably? | PAE, pTM/ipTM, seed agreement | Biological relevance of the state |
| Sequence behavior | Will the construct express and remain soluble? | Disorder, aggregation, hydrophobicity, topology, charge, PTM context | Atomic structure |
| Experimental state | Is this the molecule needed for the assay? | SEC-MALS, activity, CD, DSF, native MS, SAXS, HDX-MS, NMR | Every alternative state |
pLDDT is local
A domain can have pLDDT 90 while its orientation relative to another domain is uncertain. The corresponding PAE block will show low expected error within each domain and high expected error between them. If your assay depends on the inter-domain distance, local confidence is not the relevant metric.
PAE is relational
Predicted Aligned Error estimates how uncertain the position of residue j is when the model is aligned on residue i. Compact low-error blocks often identify domains. High inter-block error warns that the hinge or relative orientation is not fixed by the prediction.
Agreement is a robustness test
Run multiple seeds, recycles, MSA perturbations, or alternative predictors. If all runs produce the same local fold and global arrangement, the hypothesis is stronger. If high-confidence models disagree on a hinge, interface, or topology, the disagreement is information—not noise to hide by selecting rank 1.
Our companion article on the six cases where pLDDT misleads goes deeper on confidence failure modes. Here the question is what to do after the tube tells you the model was not enough.
Mismatch 1: The Fold Is Plausible, but Folding Kinetics Fail
AlphaFold predicts an endpoint. Expression asks whether the nascent chain can reach that endpoint while emerging from a ribosome in a crowded cytosol.
A protein may have a thermodynamically plausible native fold and still:
- expose aggregation-prone segments during translation;
- require slower synthesis or specific chaperones;
- form non-native disulfides;
- misassemble before a partner becomes available;
- become trapped in a kinetic intermediate.
This explains the apparently paradoxical combination of high pLDDT and inclusion bodies. The model may be reasonable for the correctly folded fraction, while the host produces almost none of that fraction.
Bench signature
- strong total-expression band;
- weak soluble fraction;
- sensitivity to induction temperature or expression rate;
- rescue by chaperone coexpression, periplasmic targeting, a fusion partner, or a different host.
Correct response
Change the folding pathway before changing the predicted native structure. Test lower induction temperature, reduced expression rate, alternative compartment/host, a small construct panel, and an N-terminal solubility fusion. If tag cleavage causes immediate precipitation, the tag was a solubility crutch rather than a folding solution; see the solubility-tag paradox.
Mismatch 2: A Correct Fold Can Still Have a Bad Surface
High pLDDT can coexist with an exposed hydrophobic or low-charge surface that drives concentration-dependent self-association. The monomeric fold may be locally accurate. The missing variable is intermolecular behavior.
Look for:
- contiguous exposed hydrophobic patches;
- unpaired β-edge strands;
- long low-complexity or disordered appendages attached to a confident core;
- a pI close to the working pH;
- native oligomerization surfaces left unsatisfied by a monomeric construct;
- exposed transmembrane or amphipathic elements misclassified as soluble.
Bench signature
- good behavior at low concentration and precipitation during concentration;
- reversible self-association by SEC-MALS or AUC;
- strong dependence on salt, pH, ligand, or concentration;
- particles or a rising baseline by DLS before visible precipitation.
Correct response
Map the liability to sequence and structure. Test buffer pH away from the pI, salt and osmolyte panels, ligand/cofactor occupancy, removal of a disordered tail, or restoration of the native partner. Mutation should be the later branch: altering a surface patch can also disrupt binding or allostery.
Mismatch 3: The Model Omitted Something the Protein Needs
Many proteins are not autonomous folding units. A confident AlphaFold monomer can represent the structure a chain adopts when supported by context that the model and the purification omitted.
Common missing elements include:
- metal ions and organic cofactors;
- lipids or sterols;
- obligate subunits;
- ligands that stabilize one state;
- glycosylation or disulfide formation;
- chaperones required for assembly but not present in the final complex.
AlphaFold’s human-proteome analysis found that low-confidence regions often correspond to disorder or to segments that become structured through heterotypic contacts. The reverse warning also matters: a high-confidence core does not tell you whether missing context is required for activity or stability (Tunyasuvunakool et al., 2021).
Bench signature
- folded-looking CD spectrum but absent activity;
- rescue by metal, nucleotide, ligand, lipid, or partner;
- correct mass but heterogeneous oligomeric state;
- expression host has incompatible PTM capability;
- instability appears only after removing a copurifying molecule.
Correct response
Build a dependency checklist from function and homology. Use ICP-MS or color/UV spectra for metals and cofactors, native MS for assemblies, glycan/disulfide analysis where relevant, and activity-guided supplementation. Do not interpret “pure apo protein” as the biologically correct endpoint by default.
Mismatch 4: AlphaFold Chose One State; the Assay Needs Another
Proteins are ensembles. A structure predictor usually returns a small set of discrete models biased toward familiar, well-represented conformations. Enzymes close around substrates. Transporters alternate access. Kinases switch regulatory elements. GPCRs move between inactive and active ensembles.
High confidence can therefore mean “this is a convincing member of the learned structural family,” not “this is the state populated in your buffer.”
This is especially dangerous when you use a high-confidence model to:
- lock a hinge with a disulfide;
- mutate an apparent cavity that closes in the active state;
- delete a regulatory segment because it looks detached;
- design an assay around one inter-domain distance;
- interpret a ligand-insensitive preparation as misfolded.
Fold-switching benchmarks provide a sharp demonstration. For proteins with two experimentally established folds, AlphaFold confidence metrics can favor a predicted conformation not supported by experiment and fail to distinguish accurate from inaccurate alternatives (Chakravarty et al., 2024).
Correct response
Sample deliberately: vary MSA depth, seeds, templates, ligands, oligomeric states, and construct context. Then test the ensemble with HDX-MS, SAXS, smFRET, DEER, NMR, or state-selective ligands. Treat a model as a prior that experiments update.
Mismatch 5: The Construct Is Not the Biological Protein
The model may be excellent for the input sequence and irrelevant to the molecule you made.
Check every construct-level change:
- Were native termini truncated through a helix, β-strand, signal peptide, or degron?
- Did the expression vector add scar residues?
- Is the tag on the wrong side of a membrane?
- Does the fusion occlude a binding site or oligomerization surface?
- Did codon optimization change translation kinetics or create an unintended sequence error?
- Is the purified band a proteolytic fragment rather than the designed construct?
- Does the host install the required PTMs?
AlphaFold often models the retained domain confidently even when the boundary destabilizes its first or last structural element. A two-residue shift can change soluble yield dramatically without changing the predicted core.
Correct response
Sequence-verify the clone, confirm intact mass, map proteolysis, and screen boundaries in parallel. Compare the isolated domain with the full-length or partner-bound construct. If a membrane protein is involved, verify topology experimentally or with an independent topology predictor before interpreting tag accessibility.
Mismatch 6: Your Buffer and Assay Select a Failure State
Purification is a sequence of perturbations. A protein can be correctly folded after expression and damaged by:
- pH near its pI;
- loss of a cofactor during dialysis;
- redox conditions incompatible with disulfides;
- dilution below a required ligand or detergent concentration;
- adsorption to plastic or concentrator membranes;
- repeated freeze–thaw;
- concentration far above physiological levels;
- an assay reporter that aggregates or interferes.
Bench signature
- activity decreases with time after purification;
- fresh material works but frozen material does not;
- SEC peak remains but activity falls;
- different assay formats disagree;
- adding back a ligand, metal, lipid, or reducing agent rescues behavior.
Correct response
Run a time-resolved mass balance. Measure recovery and function after each step rather than only at the end. The step after which the loss appears is usually more informative than another structure prediction.
A Bench-Reconciliation Workflow
Step 1: Define the failed claim
“The protein misbehaves” is not a diagnosis. Choose one:
- not expressed;
- expressed but insoluble;
- soluble but heterogeneous;
- homogeneous but inactive;
- active but unstable;
- stable at low concentration but not at working concentration.
Each failure maps to different evidence.
Step 2: Confirm molecule identity
Use intact mass or peptide mapping, sequence verification, and N-/C-terminal analysis if proteolysis is suspected. Confirm the tag and cleavage state. A beautiful model cannot explain a sample that is not the intended molecule.
Step 3: Re-read the model at the relevant scale
- Local failure: inspect pLDDT and side-chain environment.
- Domain-arrangement failure: inspect PAE and alternative seeds.
- Assembly failure: model the biological oligomer and inspect interface confidence.
- Membrane failure: verify topology and lipid/detergent context.
- Dynamics failure: generate and compare plausible states rather than one rank.
Step 4: Add sequence-behavior tracks
Overlay disorder, aggregation propensity, hydrophobicity, transmembrane topology, PTMs, and binding sites. Ask whether the bench failure sits on a visible sequence liability or an unsatisfied biological interface.
Step 5: Choose the cheapest discriminating experiment
| Hypothesis | Low-cost discriminator | Stronger follow-up |
|---|---|---|
| Mostly inclusion bodies | soluble/total expression matrix | chaperone or host comparison |
| Folded but aggregated | analytical SEC, DLS, concentration series | SEC-MALS or AUC |
| Missing cofactor | activity/add-back panel, UV spectrum | ICP-MS, native MS, co-crystal/cryo-EM |
| Wrong oligomer | calibrated SEC | SEC-MALS, native MS, AUC |
| Multi-state protein | ligand-dependent DSF/activity | HDX-MS, SAXS, NMR, smFRET |
| Disordered appendage | limited proteolysis | NMR/SAXS, construct panel |
Step 6: Change one causal variable
Do not simultaneously change host, boundaries, tag, buffer, and ligand. Use a compact matrix that isolates the dominant variable, then iterate.
Worked Example: High-Confidence Core, Inactive Preparation
Consider a two-domain enzyme with pLDDT above 90 inside both domains, high PAE between them, and a metal-binding motif in the active site. The protein expresses solubly, elutes as one SEC peak, and shows no activity.
A weak workflow mutates active-site residues because the model “must be wrong.” A stronger workflow asks three narrower questions:
- Is the molecule intact? Intact mass confirms the construct.
- Is the assembly correct? SEC-MALS confirms the expected monomer.
- Is the chemical state correct? Metal add-back restores activity, while the SEC profile barely changes.
The model correctly described a local fold. The purification removed the functional state. pLDDT and the bench were never in conflict.
The Economics: Match Evidence to the Decision
| Decision | Insufficient evidence | Minimum useful evidence |
|---|---|---|
| Order a construct | High pLDDT alone | pLDDT + PAE + boundaries + sequence-liability tracks |
| Scale expression | Soluble microculture band | soluble yield + monodispersity + identity |
| Start structural campaign | One sharp SEC peak | identity + assembly + stability + functional state |
| Design mutations | One static model | conserved context + state/PAE analysis + function-preserving constraints |
The costliest error is not distrusting AlphaFold. It is asking AlphaFold to authorize an expensive experiment that depends on properties AlphaFold does not predict.
Bottom Line
High pLDDT tells you that a local structural hypothesis is credible. It does not tell you that your construct will express, fold efficiently, remain soluble, assemble correctly, retain its cofactor, occupy the right state, or survive your buffer. When the model is blue and the tube is cloudy, narrow the failed claim, inspect the right confidence layer, add sequence behavior and biological context, and run the cheapest experiment that can distinguish the competing causes.
Bench reality does not invalidate the prediction. It completes it.
Reconciling Prediction and Experiment With Orbion
Orbion combines AlphaFold2 structure and confidence tracks with AstraUNFOLD topology, disorder, and amyloidogenicity predictions; AstraSUIT experimental-suitability context; AstraBIND pocket predictions; PTM annotations; and the PAE Insight Engine. The Design module compares construct boundaries and fusion strategies against those liabilities, while Bench generates protocols with explicit QC gates and acceptance criteria.
The goal is not a single “confidence” number. It is a traceable chain from sequence to structural hypothesis to construct to experiment.
References
- Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021). doi:10.1038/s41586-021-03819-2
- Tunyasuvunakool, K. et al. Highly accurate protein structure prediction for the human proteome. Nature 596, 590–596 (2021). doi:10.1038/s41586-021-03828-1
- Chakravarty, D. et al. AlphaFold predictions of fold-switched conformations are driven by structure memorization. Nature Communications 15, 7296 (2024). doi:10.1038/s41467-024-51801-z
- Sala, D. et al. AlphaFold prediction of structural ensembles of disordered proteins. Nature Communications 16, 1837 (2025). doi:10.1038/s41467-025-56572-9



