Blog
Structural Biology

AlphaFold Gives You One Structure. Your Protein Has Several.

Aug 12, 2026 · 21 min read

You typed in a sequence, waited a few minutes, and got back a single, beautiful model with a 90-plus pLDDT core. It looks definitive. It looks like the answer. So you open it in your viewer, find the pocket, and start planning a screen. But the protein sitting in your fridge is not a statue. It is a machine that moves: a kinase that swings its activation loop in and out, a transporter that opens to one side of the membrane and then the other, a receptor that flips between an inactive and an active state depending on what is bound. AlphaFold handed you one frame of that movie and labeled it with high confidence. The danger is not that the frame is wrong. The danger is that it is right, but it is the wrong state for the experiment you are about to run.

This is a different failure mode from low confidence or missing PTMs. Your model can be locally accurate everywhere and still be biologically incomplete, because a single set of coordinates cannot represent a protein that populates more than one conformation. This guide is about spotting that specific problem: when your target is multi-state, when AlphaFold collapsed those states into one, and what to do about it before you commit to a construct or an assay.

Key Takeaways

  • AlphaFold2 returns one structure, usually the dominant or ground state. The training objective rewards matching a single deposited structure per sequence, so the model averages away alternative functional conformations rather than enumerating them.
  • Many important targets are inherently multi-state: GPCRs (active/inactive), kinases (DFG-in/DFG-out), transporters (inward-open/outward-open), and any protein with a hinge, a lid, or a cryptic pocket.
  • The confidence metrics leak signal about heterogeneity. Bimodal or "checkerboard" PAE blocks, low inter-domain PAE that implies a hinge, low pLDDT confined to linkers, and disagreement across seeds all hint that more than one state exists.
  • You can coax AlphaFold to sample alternative states by subsampling the MSA (reduced-depth AlphaFold), running many seeds, clustering the alignment (AF-Cluster), or using ensemble-generation methods (AlphaFlow). None of these is a physics simulation, and none is guaranteed to find the biologically relevant state.
  • In drug discovery, state matters more than resolution. Docking against the inactive conformation of a target you want to activate, or against a closed cryptic pocket, wastes the entire campaign no matter how good the structure looks.

Introduction

AlphaFold2 changed what "we have a structure" means. For most single-domain, well-conserved proteins, the prediction is close enough to experiment that you can plan real work from it (Jumper et al., 2021). That reliability is exactly what makes the conformational-state problem so easy to miss. When the model is usually right, you stop asking whether "right" is the same as "complete."

A protein's function is very often carried by its motion, not its resting shape. An enzyme that never closed its lid would never catalyze. A receptor locked in one state would never signal. A transporter frozen open to one side would never move a substrate across the membrane. These proteins exist as ensembles: populations of interconverting conformations, weighted by their free energies. Experimental structural biology has always known this, which is why the Protein Data Bank contains multiple structures of the same receptor in different states, captured by trapping each one with a different ligand, mutation, or fusion partner.

AlphaFold, trained on that same PDB, does not reproduce the ensemble. It reproduces the single most predictable structure consistent with the evolutionary signal in the alignment. Understanding why it does this, and how to read the tells it leaves behind, is the difference between using a model as a starting hypothesis and mistaking it for the ground truth.

Why the Training Objective Yields One State

AlphaFold2 was trained to predict a single set of 3D coordinates from a sequence and its multiple sequence alignment (MSA), with a loss that compares the prediction to one deposited structure (Jumper et al., 2021). Nothing in that objective asks the model to enumerate alternatives. If a sequence maps to several real conformations in the PDB, the optimal single answer under the loss is something like the most evolutionarily supported, most frequently observed structure. The model learns to output that.

Two consequences follow directly:

  1. The MSA drives the answer toward the dominant state. Coevolution signal in the alignment encodes contacts that are conserved across the family. Contacts present in the most common state dominate that signal, so a deep, diverse MSA pushes the prediction toward the ground state and suppresses minority conformations (Del Alamo et al., 2022).
  2. Alternative states get averaged, not represented. When a region genuinely samples multiple conformations, a single-coordinate model cannot show both. It either commits to the dominant one or produces a compromise geometry that matches neither, which is one reason flexible loops score low on pLDDT (Ruff and Pappu, 2021).

This is not a bug that a bigger model fixes for free. It is a property of asking "what is the structure?" of an object that has several. Saldaño et al. (2022) showed systematically that AlphaFold's accuracy drops for proteins with high conformational diversity, and that the single model it returns tends to align with one end of the conformational range rather than the middle.

Multi-State Proteins You Will Actually Meet

Conformational heterogeneity is not exotic. It is the norm for many of the highest-value drug targets.

Target classThe states that matterWhy one model misleads
GPCRsInactive vs. active (agonist + transducer bound)AlphaFold usually returns the inactive-like state; the active state has a large outward shift of TM6
KinasesDFG-in (active) vs. DFG-out (inactive), αC-in/outType I vs. type II inhibitor design depends entirely on which one you model
TransportersInward-open, outward-open, occludedThe alternating-access cycle is the function; a single snapshot hides the mechanism
Ion channelsOpen, closed, inactivatedPore geometry differs by angstroms that decide conductance and drug block
Multi-domain enzymesOpen (apo) vs. closed (substrate-bound)Cryptic and interdomain pockets are shut in the resting state
Fold-switching proteinsTwo distinct folds from one sequenceAlphaFold typically predicts only the dominant fold and misses the switch entirely

The last row deserves emphasis. Some sequences genuinely adopt two different folds. Chakravarty and Porter (2022) evaluated AlphaFold2 on known fold-switching proteins and found it predicted a single dominant fold in most cases, missing the alternative conformation even when both are well characterized experimentally. If your target does something functionally interesting through motion, assume there is more than one state until you have ruled it out.

Signal 1: Bimodal and Blocky PAE

The Predicted Aligned Error (PAE) matrix is the single richest source of heterogeneity cues in a standard AlphaFold run. We covered how to read it for domain boundaries elsewhere; here the focus is on what it says about motion.

A rigid, single-state multi-domain protein produces sharp blue blocks on the diagonal with confident (blue) off-diagonal patches connecting domains that sit in a fixed arrangement. When you instead see:

  • Blue diagonal blocks with red or yellow off-diagonal regions between them, AlphaFold is telling you each domain folds confidently but their relative orientation is uncertain. That uncertainty is frequently not noise. It is the signature of a hinge or an interdomain motion that the model cannot pin to one angle.
  • A "checkerboard" or bimodal look, where inter-domain PAE is neither confidently low nor uniformly high, is consistent with two or more preferred relative orientations that the network is torn between.

The interpretation rule is simple: high inter-domain PAE between two individually confident domains means the arrangement is either flexible, multi-state, or both. It does not mean the domains are wrong. It means you should not trust the single relative pose you were handed, and you should suspect that the real protein visits more than one.

Signal 2: The Low-PAE Hinge

There is a subtler pattern worth naming on its own. When two domains show low inter-domain PAE in one region of the matrix that transitions sharply to high PAE, the low-PAE stripe often marks a hinge axis: a small region the model is confident about because it anchors the domains, flanked by uncertainty about how far the domains rotate around it.

Hinges are where open/closed and apo/holo transitions happen. A protein with a clean hinge signature (confident domains, a narrow confident contact region, uncertain global arrangement) is a protein that almost certainly closes around a substrate. If your assay depends on the closed state (for example, docking into an interdomain cleft), a single open-state model will send you after a pocket that does not exist in the geometry you modeled.

Signal 3: Low pLDDT That Tracks Function, Not Disorder

pLDDT below 50 usually means disorder (Ruff and Pappu, 2021). But not every low-confidence region is disordered. Two functional patterns masquerade as low pLDDT:

  • Flexible loops that sample multiple conformations. An activation loop or a lid that toggles between positions has no single dominant geometry, so AlphaFold hedges and pLDDT drops. The region is not structureless; it is multi-structured.
  • State-dependent segments. A helix that is ordered in one state and unwound in another (for example, parts of TM6 in some receptor transitions) can receive intermediate confidence because the model is blending states.

The tell is location. Disorder tends to sit at termini and in long linkers. When low or intermediate pLDDT falls precisely on a known functional element (an activation loop, a switch region, a gating helix), read it as a flexibility flag, not a disorder flag. The distinction changes what you do next: you do not truncate a functional switch the way you would truncate a disordered tail.

Signal 4: Disagreement Across Seeds and Models

AlphaFold2 produces five models per run and uses random seeds internally. For a single-state protein, the five models agree closely in the confident regions. When a protein is multi-state, the models often disagree in exactly the region that moves: one seed places a loop in, another places it out; one puts a domain closed, another leaves it ajar.

Run the same sequence with several seeds and overlay the results. Convergence in the moving region argues for one dominant state. Divergence argues that AlphaFold itself is uncertain which state to commit to, which is a strong hint that more than one is physically accessible. This is close to free: it needs no new method, only more sampling and a superposition. Del Alamo et al. (2022) built their state-sampling protocol on exactly this principle, combining reduced MSAs with many seeds to widen the range of conformations AlphaFold explores.

How to Sample Alternative States

If the signals above suggest a second state, you have several ways to try to make AlphaFold show it. All of them are approximations. None simulates dynamics or gives you populations or free energies. They perturb the inputs so the network relaxes toward conformations it otherwise suppresses.

Reduced-MSA (MSA subsampling)

The most established approach. Because the depth of the MSA pushes AlphaFold toward the dominant state, deliberately shallow alignments loosen that constraint and let alternative states surface. Del Alamo et al. (2022) showed that subsampling the MSA to a few dozen sequences, combined with many random seeds, let AlphaFold2 predict both the inward- and outward-facing states of transporters and both active-like and inactive-like states of GPCRs. The tradeoff is real: shallower MSAs mean lower average confidence and more junk models, so you must generate many and cluster the output.

Sequence-clustering the MSA (AF-Cluster)

Instead of thinning the alignment randomly, cluster it by sequence similarity and predict from each cluster. Different subfamilies can carry coevolution signals for different states. Wayment-Steele et al. (2024) used this idea to predict multiple conformations of the fold-switching protein KaiB and metamorphic proteins that a full-depth MSA collapses to one fold. This works best when the family is large and diverse enough that subgroups genuinely encode different states.

Many seeds, plain AlphaFold

Even without touching the MSA, running dozens to hundreds of seeds and clustering the results can surface minority states, especially when paired with modest MSA reduction. This is the cheapest first move and pairs naturally with the seed-disagreement diagnostic above.

State-targeted MSA masking (SPEACH_AF)

Rather than reducing depth globally, mask coevolution signal at specific positions to bias the prediction away from the contacts that define the dominant state. Stein and Mchaourab (2022) introduced this to sample conformational heterogeneity in a more targeted way than random subsampling, which helps when you know which region should move.

A newer class of methods retrains or repurposes structure predictors to emit ensembles directly. AlphaFlow (Jing et al., 2024) fine-tunes AlphaFold with a flow-matching objective so it generates a distribution of structures rather than a single point estimate, approximating the kind of diversity you would otherwise chase with molecular dynamics. These methods are promising but young; treat their ensembles as hypotheses to test, not as calibrated populations.

MethodWhat you changeBest whenMain limitation
Reduced MSAMSA depth (shallow)Transporters, GPCRs, apo/holoLower confidence, many junk models
AF-ClusterMSA split by sequence clusterFold switchers, large diverse familiesNeeds deep, diverse alignment
Many seedsRandom seed countCheap first pass, loop/hinge motionMay stay stuck in dominant state
SPEACH_AFMasked coevolution at chosen sitesYou know which region movesRequires prior mechanistic insight
AlphaFlowThe model itself (ensemble output)Want a distribution, not one stateYoung; populations not calibrated

A shared caveat runs through all of these: sampling a second geometry is not the same as sampling the right second geometry, and none of these methods tells you the relative populations. Always validate a sampled state against orthogonal evidence (a known ligand-bound structure, functional mutagenesis, HDX-MS, DEER, or SAXS) before you build a program on it.

Why This Matters for Drug Discovery

The practical cost of the one-state trap is highest in structure-based drug design, because the entire premise of SBDD is that you are targeting a real pocket in a real conformation.

  • You may be dosing against the wrong state. If you want to inhibit a kinase and you model the DFG-in active state, you will miss the type II chemical matter that only binds DFG-out. Model the wrong receptor state and your virtual screen enriches for compounds that stabilize a conformation you did not want.
  • Cryptic pockets are closed in the ground state. Many of the most valuable allosteric sites simply do not exist in the resting structure; they open transiently or upon binding. A single apo model shows a smooth, undruggable surface where a real pocket breathes open. Conformational sampling is often the only computational way to see it before a crystallographer traps it.
  • Selectivity lives in the differences between states. Two related targets can look similar in one state and diverge sharply in another. If you only have one state per target, you cannot reason about the state-specific differences that selectivity usually exploits.

The correct mental model: for a dynamic target, the question is not "how accurate is my structure?" but "which structure, out of several the protein actually adopts, am I looking at, and is it the one my program needs?"

Case Study: A GPCR Modeled in the Wrong State

Problem. A team is pursuing an agonist for a class A GPCR. They run the receptor through AlphaFold2, get a clean 7-transmembrane bundle with a high-pLDDT core, and start docking their agonist series into the orthosteric pocket. Hit rates are poor and structure-activity relationships make no sense: compounds that are functionally active in cells dock no better than inactive ones.

The model looked authoritative, so the state question never came up. But class A GPCRs undergo a large conformational change on activation, most visibly a substantial outward movement of the cytoplasmic end of TM6 that opens the transducer-coupling cavity. AlphaFold, trained on a PDB where inactive-state receptor structures dominate and reinforced by an MSA that encodes the more conserved inactive contacts, had returned an inactive-like state (Heo and Feig, 2022). Docking agonists into an inactive orthosteric pocket cannot reproduce the geometry that agonist binding stabilizes, so the SAR collapsed.

Solution. The team applied a state-sampling protocol: reduced-depth MSAs combined with many seeds, following the reduced-MSA approach validated for receptors and transporters (Del Alamo et al., 2022). Instead of one model they generated an ensemble, then clustered it. Two populations emerged: an inactive-like cluster resembling the original prediction and a second cluster with the characteristic TM6 outward shift, an active-like state. They cross-checked the active-like cluster against a published agonist-bound structure of a close homolog and confirmed the TM6 position and key microswitch residues were consistent.

Outcome. Re-docking the agonist series into the active-like model recovered a sensible SAR: cellularly active compounds now scored and posed in line with their functional potency, and the pocket accommodated substituents that had clashed in the inactive model. The team had not run a single new experiment beyond re-doing the prediction correctly. The lesson generalizes past GPCRs: whenever the biology depends on a state transition, sample the states before you trust the pocket.

Practical Tools: Does My Target Have Multiple States, and Did AlphaFold Collapse Them?

Use this as a decision workflow. Each step is answerable from a standard AlphaFold run plus public knowledge of your target.

STEP 1 - Prior probability: is this target class dynamic?
  □ GPCR, kinase, transporter, channel, or multi-domain enzyme?
  □ Known apo/holo, open/closed, or active/inactive states in the PDB
    for this protein or a close homolog?
  □ Literature mentions a hinge, lid, gating helix, cryptic pocket,
    or fold switch?
      → Any box checked: treat as multi-state until proven otherwise.
        Proceed to Step 2.
      → None checked and small single-domain: single state is likely
        adequate. Stop.

STEP 2 - Read the confidence metrics for heterogeneity tells
  □ PAE: individually confident domains with red/yellow between them
    (uncertain relative arrangement)?
  □ PAE: a narrow low-PAE stripe between domains (hinge signature)?
  □ pLDDT: low/intermediate confidence sitting ON a functional element
    (activation loop, switch, gating helix), not just termini/linkers?
      → Any yes: AlphaFold is signaling flexibility or multiple states.
        Proceed to Step 3.

STEP 3 - Test AlphaFold's own certainty (cheap)
  □ Run 20+ seeds of the same sequence, full MSA.
  □ Superpose the models; inspect the suspect region.
      → Models converge everywhere: one dominant state is likely. Note
        it and proceed with caution.
      → Models diverge in the moving region: AlphaFold is unsure which
        state to commit to. Proceed to Step 4.

STEP 4 - Actively sample alternative states
  □ Reduced-MSA + many seeds (start here for GPCRs/transporters).
  □ AF-Cluster if the family is large and diverse (fold switches).
  □ SPEACH_AF if you know which region should move.
  □ AlphaFlow if you want a distribution rather than discrete states.
  □ Cluster the output; identify distinct conformational populations.

STEP 5 - Validate before you build on it
  □ Compare each sampled state to a known ligand/state-trapped structure
    of the protein or a homolog.
  □ Cross-check with functional data (mutagenesis of switch residues).
  □ Where available, use HDX-MS, DEER, or SAXS as orthogonal evidence.
      → Confirmed second state: use the state your program requires.
      → Cannot validate: label as hypothesis, do not commit a campaign.

Two rules keep you out of trouble. First, the absence of a heterogeneity signal in the confidence metrics is weak evidence, not proof, of a single state: a deep MSA can hide a real second state behind high confidence. Second, match the state to the experiment. Modeling the state you can predict most confidently is worthless if your assay needs the state you did not model.

Economic Analysis

The cost of the one-state trap is almost entirely downstream, which is what makes it dangerous. The prediction is free and fast; the wasted work is neither.

ScenarioWhere the cost landsRough magnitude
Virtual screen against the wrong receptor/kinase stateSynthesis and assay of enriched-but-irrelevant hitsWeeks to months of medchem plus reagent cost
Docking into a closed cryptic pocketAn entire allosteric campaign chasing a site that never opened in the modelMonths, sometimes a program restart
Construct designed around a single-state interdomain arrangementExpression/crystallization attempts on a geometry the protein does not favorMultiple purification and screening rounds
Missed fold switchFundamentally wrong mechanistic hypothesisHard to bound; can misdirect a project for a long time

Set that against the cost of the diagnostic side of the workflow. Checking prior state knowledge, reading the PAE and pLDDT tells, and running a seed-divergence test costs an afternoon and no new infrastructure. Even full state sampling (reduced-MSA ensembles plus clustering) is compute-bounded and cheap next to one round of chemistry. The asymmetry is stark: a few hours of state triage protects months of wet-lab spend. The expensive mistake is not running the extra sampling. It is trusting a confident single model and finding out which state it was after the assays come back flat.

The Bottom Line

QuestionAnswer
Does AlphaFold give me the whole conformational picture?No. It returns one structure, biased toward the dominant/ground state.
Is a confident single model automatically the right state?No. Local accuracy and biological completeness are different things.
How do I know my target is multi-state?Target class, known states in the PDB, and heterogeneity tells in PAE/pLDDT/seed disagreement.
Can I make AlphaFold show other states?Sometimes: reduced MSA, AF-Cluster, SPEACH_AF, seeds, or ensemble methods (AlphaFlow).
Do those methods give populations or dynamics?No. They surface candidate states; you still must validate and cannot read off free energies.
Why care in drug discovery?You can dose against the wrong state or a closed pocket and waste the campaign.

Bottom line: AlphaFold answers "what is the most predictable structure of this sequence?" For a protein that moves, that is not the same question as "what are the states my experiment depends on?" Read the confidence metrics for heterogeneity, test the model's own certainty across seeds, sample alternative states when the signals warrant it, and validate before you commit. A single beautiful model is a hypothesis about one frame of a movie, not the movie.

Spot Hidden States Before You Commit

This is exactly the interpretation layer we built into Orbion's Characterization module. The PAE Insight Engine reads the AlphaFold2 confidence metrics for you and derives domain boundaries, hinge regions, conformational-heterogeneity cues, and dynamics hints directly from the PAE matrix, so you do not have to eyeball raw heatmaps to notice that two confident domains have an uncertain relationship or that a low-PAE stripe marks a hinge. Combined with per-residue disorder from AstraUNFOLD (to separate genuine disorder from functional flexibility) and the interactive 3D viewer, it helps you catch the tell-tale signs that a single model is hiding more than one state, before you lock in a construct boundary or point an assay at a pocket that may be closed.

To be clear about what this is and is not: Orbion surfaces heterogeneity cues from AlphaFold2 confidence metrics. It does not simulate full conformational ensembles or return calibrated populations. It tells you where to look and when to be suspicious, so you can decide whether to run the state-sampling and validation steps above before spending in the lab.

Ready to see what your model might be hiding? Upload your sequence to Orbion and check the Characterization module's structural insights before your next construct or docking run.

References

  1. Jumper J, et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596:583-589. DOI: 10.1038/s41586-021-03819-2

  2. Del Alamo D, Sala D, Mchaourab HS, Meiler J. (2022). Sampling alternative conformational states of transporters and receptors with AlphaFold2. eLife, 11:e75751. DOI: 10.7554/eLife.75751

  3. Wayment-Steele HK, et al. (2024). Predicting multiple conformations via sequence clustering and AlphaFold2. Nature, 625:832-839. DOI: 10.1038/s41586-023-06832-9

  4. Stein RA, Mchaourab HS. (2022). SPEACH_AF: Sampling protein ensembles and conformational heterogeneity with Alphafold2. PLOS Computational Biology, 18(8):e1010483. DOI: 10.1371/journal.pcbi.1010483

  5. Jing B, Berger B, Jaakkola T. (2024). AlphaFold Meets Flow Matching for Generating Protein Ensembles. arXiv preprint. DOI: 10.48550/arXiv.2402.04845

  6. Saldaño T, et al. (2022). Impact of protein conformational diversity on AlphaFold predictions. Bioinformatics, 38(10):2742-2748. DOI: 10.1093/bioinformatics/btac202

  7. Chakravarty D, Porter LL. (2022). AlphaFold2 fails to predict protein fold switching. Protein Science, 31(6):e4353. DOI: 10.1002/pro.4353

  8. Heo L, Feig M. (2022). Multi-state modeling of G-protein coupled receptors at experimental accuracy. Proteins: Structure, Function, and Bioinformatics, 90(11):1873-1885. DOI: 10.1002/prot.26382

  9. Ruff KM, Pappu RV. (2021). AlphaFold and Implications for Intrinsically Disordered Proteins. Journal of Molecular Biology, 433(20):167208. DOI: 10.1016/j.jmb.2021.167208

Investors: [[../entities/insight|Insight]]