Blog
Expression & PurificationAI & Computation

Predicting Aggregation-Prone Regions From Sequence (Before They Wreck Your Prep)

Sep 30, 2026 · 8 min read

Your protein is soluble after lysis, survives affinity capture, and turns cloudy during concentration. The sequence contained the warning all along—but only if you distinguish a buried hydrophobic core from an aggregation-prone segment that actually becomes exposed.

Sequence predictors can identify regions with chemical patterns associated with self-association. They are most useful as early risk maps, not binary declarations that a protein “will aggregate.” The same segment may be harmless in a stable fold, dangerous in a partially unfolded intermediate, or functional in an amyloid assembly.

Key Takeaways

  • Aggregation is a pathway, not a sequence label. Expression, folding, concentration, temperature, agitation, interfaces, and buffer determine whether a hotspot becomes accessible.
  • Most predictors score intrinsic segment propensity. They do not fully model the native structure, chaperones, purification history, or kinetic traps.
  • Exposure changes the interpretation. A strong hotspot buried in a stable core is different from the same chemistry in a disordered tail or unstable edge strand.
  • Use multiple evidence tracks. Combine aggregation propensity with disorder, secondary structure, solvent exposure, stability, charge, and topology.
  • Design the smallest intervention. Boundary changes, surface mutations, host changes, fusion tags, and process conditions solve different failure modes.
  • Validate under the intended workflow. Soluble yield at low concentration does not certify behavior during SEC, concentration, freezing, or assay incubation.

What Makes a Sequence Segment Aggregation-Prone?

Aggregation-prone regions often combine several features:

  • hydrophobic side chains that prefer burial;
  • aromatic residues capable of ordered stacking;
  • low net charge or poor charge patterning;
  • β-sheet-forming tendency;
  • limited gatekeeper residues that interrupt self-complementary structure;
  • repeated or low-complexity patterns;
  • local instability that makes the segment accessible.

Algorithms emphasize different signals. TANGO models competition among β-aggregation, secondary structure, and other conformational states (Fernandez-Escamilla et al., 2004). AGGRESCAN derives aggregation propensities from experimentally informed amino-acid scales (Conchillo-Solé et al., 2007). CamSol estimates intrinsic solubility and can incorporate structural context (Sormanni et al., 2015).

No method has a complete view of the route your protein takes from translation to concentrated sample.

Native Folding and Aggregation Compete

Many folded proteins contain segments that would aggregate if isolated. They are safe because native folding buries those surfaces quickly and persistently. Risk rises when:

  • translation exposes the segment before the rest of the domain emerges;
  • the host lacks compatible folding machinery;
  • a construct boundary destabilizes the domain;
  • purification removes a stabilizing ligand or partner;
  • pH or salt weakens native contacts;
  • concentration increases collision frequency;
  • temperature or freeze–thaw populates partially unfolded states;
  • a tag is removed and exposes a hydrophobic terminus.

This explains why “the predictor says aggregation” is not a sufficient conclusion. The useful output is a map of sequences that deserve structural and process context.

Classify Each Hotspot Before Acting

Buried core hotspot

A strong segment is deeply buried in a high-confidence, well-packed domain. Mutating it may destabilize the fold more than it improves solubility. First protect the native state: ligand, cofactor, temperature, buffer, construct boundary, or host.

Exposed surface hotspot

A hydrophobic patch remains solvent-accessible in credible structural models. Carefully chosen polar or charged substitutions may reduce self-association, provided they avoid interfaces and functional surfaces.

Disordered hotspot

A low-complexity or disordered region contains self-association chemistry. Truncation, boundary redesign, or a solubility tag may help if the region is not functionally required.

Conditional hotspot

The segment is buried in one conformation and exposed in another. Stability and conformational-state control may matter more than direct mutation.

Membrane segment

A transmembrane helix will look extremely hydrophobic to generic aggregation tools. It is not a soluble-protein hotspot; it is a topology feature that requires an appropriate membrane or mimic.

A Sequence-to-Structure Triage Workflow

1. Run more than one predictor

Agreement across conceptually different methods raises priority. Disagreement is informative: one method may be responding to hydrophobicity, another to β-structure, and another to global solubility.

2. Overlay topology and disorder

Exclude signal peptides and transmembrane helices from naive interpretation. Flag hotspots within disordered tails, linkers, or low-complexity segments.

3. Map hotspots onto the best structural evidence

Ask whether each region is buried, exposed, interfacial, or uncertain. Review alternative states and predicted aligned error; a high-confidence single structure can still hide transient exposure.

4. Add functional constraints

Annotate active sites, binding sites, oligomer interfaces, PTMs, disulfides, and conserved motifs. A hotspot overlapping a required interface cannot simply be “fixed” by making it polar.

5. Trace the experimental stage of failure

Aggregation timing narrows the mechanism:

Failure pointMore likely causeFirst intervention
During expressioncotranslational folding, overload, host mismatchlower expression rate, change host, chaperone or tag
After lysisexposed hydrophobic surfaces, missing partnerstabilize buffer, add ligand, shorten handling
After tag cleavagenewly exposed terminus or destabilizationredesign junction, leave removable spacer, change protease conditions
During SECdilution or buffer incompatibilityscreen pH, salt, additives, detergent
During concentrationreversible self-association or nucleationreduce target concentration, mutate exposed patch, optimize excipients
Freeze–thawinterfacial stress or partial unfoldingcryoprotectant, aliquots, controlled freezing

Choosing the Intervention

Construct redesign

Use when risk localizes to a dispensable disordered tail, unstable domain boundary, or exposed fragment created by the construct. Make a boundary series rather than one large deletion.

Surface mutation

Use when a credible exposed hotspot lies outside functional interfaces. Favor minimal changes. Review charge distribution, local hydrogen bonding, conservation, and structural recovery.

Solubility tag

Use when the problem occurs during expression or early handling. MBP, SUMO, Fh8, and other tags act through different mechanisms and do not guarantee post-cleavage behavior. See our solubility-tag decision guide.

Host or expression-rate change

Use when cotranslational folding or processing is the suspected bottleneck. Lower temperature and slower expression can reduce the rate at which exposed aggregation-prone chains accumulate.

Formulation and handling

Use when a folded protein aggregates only at high concentration, during agitation, or after freeze–thaw. The sequence is unchanged, but the pathway is altered.

How Not to Design an “Anti-Aggregation” Mutant

Avoid these shortcuts:

  1. Replace every hydrophobic residue in the hotspot. You may destroy the core or binding surface.
  2. Add charge without checking local electrostatics. A new salt bridge or repulsive patch can change folding and function.
  3. Optimize one structural model. Transient exposure may arise in an alternate state.
  4. Select only by predicted solubility. The best-scoring sequence may lose activity or stability.
  5. Test at analytical dilution only. Many proteins fail only near the concentration needed for assays or structures.

A good design changes the minimum number of residues needed to test a mechanism.

Worked Example: Aggregation During Concentration

A 42 kDa enzyme expresses solubly and elutes as a main monomer peak with a small shoulder. Above 8 mg/mL it develops visible particles and loses 30% activity overnight.

Sequence analysis finds two hotspots:

  • Region A is buried in the hydrophobic core with high structural confidence.
  • Region B is a five-residue exposed patch on a peripheral helix, distant from the active site and oligomer interface.

The rational plan is not to mutate both.

  1. Keep Region A unchanged and screen conditions that protect the native fold.
  2. Design three conservative variants around Region B: one aromatic-to-polar change, one charge introduction, and one two-residue combination.
  3. Confirm predicted structural recovery and unchanged active-site geometry.
  4. Express wild type and variants under identical conditions.
  5. Compare yield, SEC profile, dynamic light scattering, activity, and recovery after concentration to the actual target concentration.
  6. Include a buffer-only formulation arm so sequence and process effects are separable.

If the surface variants improve concentration behavior without harming activity, the mechanism is supported. If formulation alone rescues the sample, mutation may be unnecessary.

Prediction Output Should Drive an Experiment

A useful aggregation report includes:

  • residue-level hotspot map;
  • consensus and disagreement across predictors;
  • structural exposure and confidence;
  • overlap with disorder, topology, and function;
  • likely stage of exposure;
  • candidate interventions ranked by reversibility;
  • explicit validation conditions.

Orbion’s AstraUNFOLD provides amyloidogenicity and related sequence-risk signals, while Design can compare construct and mutation candidates with solubility and aggregation-oriented scoring. These outputs are strongest when used together: one track locates the liability, another asks whether the proposed fix preserves the rest of the protein.

For broader troubleshooting, see The Hidden Link Between Disorder, Aggregation, and Failed Purifications. For formulation-stage diagnosis, see Why Your Protein Loses Activity After Purification.

Bottom Line

Sequence-based aggregation prediction is an early-warning system. It can identify segments likely to self-associate, but the experimental danger depends on whether those segments become exposed and under what conditions.

Map hotspots onto topology, disorder, structure, function, and workflow. Then choose the smallest intervention that tests a mechanism. The objective is not a sequence with the lowest theoretical aggregation score. It is a protein that remains folded, functional, and usable at the concentration and timescale your experiment requires.

References

  1. Fernandez-Escamilla AM, Rousseau F, Schymkowitz J, Serrano L. Prediction of sequence-dependent and mutational effects on the aggregation of peptides and proteins. Nature Biotechnology. 2004. doi:10.1038/nbt1012
  2. Conchillo-Solé O, et al. AGGRESCAN: a server for the prediction and evaluation of “hot spots” of aggregation in polypeptides. BMC Bioinformatics. 2007. doi:10.1186/1471-2105-8-65
  3. Sormanni P, Aprile FA, Vendruscolo M. The CamSol method of rational design of protein mutants with enhanced solubility. Journal of Molecular Biology. 2015. doi:10.1016/j.jmb.2014.09.026
  4. Chiti F, Dobson CM. Protein misfolding, amyloid formation, and human disease. Annual Review of Biochemistry. 2017. doi:10.1146/annurev-biochem-061516-045115