Blog
Protein Engineering

ΔTm or ΔΔG? Which Stability Metric to Trust When You're Ranking Mutations

Aug 3, 2026 · 25 min read

You ran a stability predictor on 40 candidate mutations, and now every row has two numbers. One mutation shows a predicted ΔTm of +6.2°C but a ΔΔG of only +0.3 kcal/mol. The next shows +2.8°C and +1.6 kcal/mol. They rank in opposite orders. You have three synthesis slots this quarter and a PI who wants results, and the spreadsheet is quietly asking you to pick a lane. Rank by the melting shift, and you chase the +6.2°C mutation. Rank by free energy, and you skip it. One of those choices sends you down a three-month dead end.

ΔTm and ΔΔG are the two most common outputs of any stability workflow, computational or experimental. They measure related things, but they are not interchangeable — and when they disagree, the disagreement is information, not noise. This guide explains what each number actually means, why the same mutation can look great on one axis and mediocre on the other, and how to build a ranking that survives contact with the bench.


Key Takeaways

  • ΔΔG (kcal/mol) is a free-energy change at a fixed reference temperature; ΔTm (°C) is the shift in the melting midpoint. They are linked by the approximation ΔTm ≈ ΔΔG / ΔS_m, so the same +1 kcal/mol stabilization can read as ~3°C in one protein and ~6°C in another — purely because of the protein's unfolding enthalpy.
  • Rank primarily by ΔΔG; use ΔTm as corroboration. A large raw ΔTm at a single position is the most easily inflated number in your table. It is not a reason to promote a mutation on its own.
  • Both metrics are directional, limited-resolution predictions — a ranked shortlist to test, not ground truth. Confidence drops out-of-distribution: buried positions, large or charged substitutions, and proteins unlike the training data.
  • Lead with single mutants. Predictors largely assume additivity and miss epistasis, so stacking "good" singles routinely underperforms the naive sum. Treat every combination as a hypothesis, not a conclusion.
  • Check the functional cost before the stability gain. A mutation that raises Tm but deletes a PTM site, fills the active-site pocket, or flips membrane topology is a bad trade — no matter how good the ΔΔG looks.
  • Most mutations destabilize. Stabilizing hits are a minority, so expect to test several candidates per win, and use early computational triage to spend scarce bench slots on the ones most likely to be real.

Two Numbers, Two Different Questions

Before you can decide which metric to trust, you need to be precise about what each one asks. They are not two measurements of the same quantity. They are answers to two different questions about the same free-energy landscape.

First, a convention, because this trips up more rankings than any other single thing. Throughout this guide, a stabilizing mutation is one that makes the folded state harder to unfold: it raises the free energy of unfolding (positive ΔΔG, unfolding convention) and raises the melting temperature (positive ΔTm). Many tools and papers instead report ΔΔG of folding, where negative means stabilizing. Before you sort a single column, confirm which sign your tool uses. A flipped convention silently inverts your entire shortlist, and it will look completely reasonable while doing so.

What ΔΔG Measures

ΔΔG is the change in the free energy of unfolding caused by your mutation, evaluated at a fixed reference temperature (commonly 25°C):

ΔΔG = ΔG_unf(variant) − ΔG_unf(wild-type)

ΔG_unf is the free-energy difference between the folded and unfolded states — how much thermodynamic "cost" separates your protein from a denatured chain at that temperature. A positive ΔΔG means the variant sits in a deeper well: it takes more free energy to unfold it. This is the quantity most directly tied to how much of your protein stays folded at a given working temperature, and it is measured in kcal/mol.

The key property: ΔΔG is defined at a temperature you choose. That makes it comparable across mutations in the same protein under the same conditions, and it maps cleanly onto the folded/unfolded population you care about at 37°C, 4°C, or wherever your assay lives.

What ΔTm Measures

ΔTm is the shift in the melting temperature — the temperature at which the folded and unfolded states are equally populated and ΔG_unf = 0:

ΔTm = Tm(variant) − Tm(wild-type)

A positive ΔTm means the variant melts at a higher temperature. It is measured in °C (a shift in °C equals a shift in K), and it is what a differential scanning fluorimetry (DSF/thermal shift) or DSC experiment reports most directly. ΔTm is intuitive, cheap to measure, and easy to communicate — "we bought ourselves eight degrees" is a sentence a whole team understands.

But Tm is a single point on a curve. It tells you where stability crosses zero, not how deep the well is at your working temperature. And as the next section shows, the conversion between a free-energy change and a melting-point shift depends on a property of the protein that has nothing to do with your mutation's quality.


Here is the relationship that ties the two metrics together, and it is the single most useful piece of thermodynamics for anyone ranking mutations.

A protein's stability as a function of temperature — its stability curve — is described by the Gibbs–Helmholtz relation (Becktel & Schellman, 1987):

ΔG_unf(T) = ΔH_m · (1 − T/Tm) − ΔC_p · [ (Tm − T) + T · ln(T/Tm) ]

where ΔH_m is the enthalpy of unfolding at the melting temperature (the van 't Hoff enthalpy — a measure of how cooperative and "all-or-nothing" the transition is), and ΔC_p is the heat-capacity change on unfolding (Robertson & Murphy, 1997). You do not need to solve this to use it. You need one consequence of it.

For a small stabilization near the wild-type Tm, the shift in melting temperature relates to the free-energy change by:

ΔTm ≈ ΔΔG / ΔS_m  =  ΔΔG · Tm / ΔH_m

because at the midpoint ΔS_m = ΔH_m / Tm (entropy and enthalpy exactly balance when ΔG = 0). This approximation is the workhorse behind the classic result that stabilizing mutations raise Tm by lowering the entropy of unfolding (Matthews, Nicholson & Becktel, 1987).

Read that formula slowly, because it contains the whole problem: ΔTm depends on ΔH_m, and ΔH_m is a property of the protein, not of your mutation.

A Worked Example

Take two proteins, both melting at 57°C (Tm ≈ 330 K), and give each the identical stabilizing mutation worth ΔΔG = +1.0 kcal/mol.

ProteinΔH_m (kcal/mol)ΔS_m = ΔH_m/Tm (kcal/mol·K)ΔTm ≈ ΔΔG/ΔS_m
A — large, highly cooperative1200.36+2.8°C
B — small, weakly cooperative600.18+5.5°C

Same mutation. Same thermodynamic stabilization. Twice the ΔTm in protein B, for no reason other than that its unfolding transition is less enthalpically cooperative. A small, weakly cooperative domain will always show inflated melting-point shifts relative to the free energy actually gained. A large, cooperative multidomain protein will show muted ones.

This is why you cannot compare raw ΔTm values across different proteins, and why even within one protein a mutation that happens to perturb ΔH_m (by changing the folding cooperativity) can post a ΔTm that overstates or understates its real free-energy contribution. ΔΔG is the more transferable currency because it is a free energy at a fixed temperature; ΔTm is that free energy filtered through the protein's enthalpy.

The caveat, stated honestly: this linear relationship holds for small shifts near the wild-type Tm, and it assumes the mutation does not substantially change ΔH_m or ΔC_p. For large ΔTm values, or mutations that reshape the transition itself, the approximation degrades. It is a lens for interpretation, not a conversion factor to report to three decimal places.


ΔTm vs ΔΔG: A Side-by-Side Comparison

ΔTm (melting-temperature shift)ΔΔG (folding free-energy change)
What it measuresShift in the temperature where folded = unfoldedChange in free energy of unfolding at a fixed reference T
Units°C (equivalently, K)kcal/mol
Reference conditionThe melting midpoint (a moving target per variant)A temperature you fix (e.g., 25°C)
Directly measured byDSF / thermal shift, DSCChemical denaturation (urea/GuHCl), ΔG fits
Greatest strengthIntuitive, cheap, high-throughput; great for screeningTransferable; comparable across mutations; maps to folded population at working T
Primary failure modeMagnitude depends on protein ΔH_m; inflated at weakly cooperative positions; assumes clean two-state meltingSign-convention confusion; conservative magnitudes; still limited-resolution out-of-distribution
Best used forCorroboration; screening readout; when your target is thermal tolerancePrimary ranking key for prioritization
Watch out forBig raw values at single positions; irreversible/aggregating melts distort TmWhich sign means "stabilizing"; over-trusting small differences

The short version: ΔΔG is your ranking key, ΔTm is your corroboration. The rest of this guide is about the exceptions and the failure modes that make that rule non-trivial.


Why You Should Rank by ΔΔG (and Use ΔTm to Corroborate)

Three reasons make ΔΔG the better primary sort key for most campaigns:

  1. It is temperature-anchored. ΔΔG is defined at a reference temperature you control, so two mutations' ΔΔG values are directly comparable. Two ΔTm values are each measured at their own melting midpoint, which shifts from variant to variant — you are comparing readings taken under slightly different conditions.

  2. It is less sensitive to cooperativity artifacts. As the worked example showed, ΔTm scales inversely with ΔH_m. A mutation that raises Tm partly by making the transition sharper is inflating its own melting-point shift without contributing proportional free energy at your working temperature.

  3. It maps to what you usually want. For shelf life, resistance to unfolding at 37°C, or tolerance of a purification step, the operative quantity is how much of your protein stays folded at that temperature — which is ΔG at that temperature, i.e., ΔΔG at your reference. A variant can raise Tm yet barely change ΔG at 37°C if the stabilization is concentrated near the melting transition.

The honest exception: if your application is the melting point — a thermostable enzyme for a high-temperature process, a diagnostic reagent that must survive shipping excursions, a screen whose readout is thermal tolerance — then ΔTm at the relevant temperature is your operational target, and you should weight it accordingly. "Rank by ΔΔG" is a strong default, not a universal law. Know which quantity your product actually cares about.

When a Big ΔTm Misleads You

The most dangerous row in your spreadsheet is a large positive ΔTm sitting next to a small or negative ΔΔG. Treat it as a flag, not a find. Common causes:

  • Low-cooperativity inflation. A position that couples strongly to a weakly cooperative part of the unfolding transition can post an outsized ΔTm for a modest free-energy change.
  • Predictor magnitude noise. Learned predictors can assign large ΔTm magnitudes at particular positions that do not reflect proportional free-energy gains, especially where training data is thin. The rank order among mutations tends to be more reliable than the absolute magnitude of any single one.
  • Assay artifacts on the experimental side. DSF and DSC assume a clean, reversible, two-state melt. Aggregation during heating, irreversible unfolding, or multi-state transitions distort the apparent Tm — sometimes producing a "shift" that is really a change in aggregation behavior, not stability.

When ΔTm is large but ΔΔG is not, do not promote the mutation on the melting shift alone. Corroborate first.


Both Metrics Are Directional, Limited-Resolution Predictors

Whether your numbers come from a physics-based calculation, a machine-learning model, or a preliminary assay, treat them the same way: a ranked shortlist of hypotheses, not measurements of truth. This is not false modesty. It is what the benchmark literature shows.

Predictors are trained and calibrated on databases like ProThermDB (Nikam et al., 2021) and FireProtDB (Stourac et al., 2021). Those datasets are invaluable, but they are biased in ways that propagate straight into predictions: heavily skewed toward destabilizing mutations, concentrated in a handful of well-studied proteins, and unevenly sampled across residue types and burial. A model learns the distribution it is shown.

Two findings should shape how much you trust any given number:

  • Systematic bias. A self-consistency test showed that many ΔΔG predictors violate a basic physical requirement — the antisymmetry that the forward mutation A→B and the reverse B→A should have equal and opposite ΔΔG (Usmanova et al., 2018). Programs that look accurate on standard benchmarks can carry a directional bias that inflates apparent performance.
  • Modest out-of-distribution accuracy. On a fresh, non-redundant test set (S669), a thorough comparison of available tools found correlations that are useful for enrichment but far from oracle-grade, and that degrade on proteins unlike the training data (Pancotti et al., 2022).

Confidence Guidance

Trust the numbers more when:

  • The mutation is a conservative substitution at a solvent-exposed position.
  • ΔΔG and ΔTm agree in sign and rough magnitude.
  • Your protein resembles the kinds of proteins in public stability datasets (soluble, single-domain, well-represented folds).
  • You are using the ranking to pick a shortlist, not to predict an exact degree count.

Trust the numbers less when:

  • The position is buried, or the substitution is large, charged, or introduces/removes a proline or glycine.
  • The two metrics disagree sharply (see the combined-reading table below).
  • Your protein is membrane-embedded, heavily disordered, or otherwise underrepresented in training data.
  • You are tempted to believe an absolute magnitude ("this one is worth 9.4°C") rather than a rank.

Reading the Two Numbers Together

ΔΔG (stabilizing?)ΔTmInterpretationAction
Strongly stabilizing (+)Large +Corroborated best casePrioritize
Strongly stabilizing (+)Small + / flatReal free-energy gain, likely a high-ΔH_m or cooperative proteinPrioritize; expect a modest Tm shift
Weak / neutralLarge +Suspicious — possible ΔTm inflation or an enthalpy-driven artifactCorroborate before trusting; do not promote on ΔTm alone
Destabilizing (−)Large +ContradictionRe-check sign convention and assay; likely artifact
Destabilizing (−)DownConsistentDeprioritize

The Multi-Mutant Trap: Why Good Singles Don't Add Up

You found four single mutations, each with a clean +1.5 kcal/mol. Stack all four and you should get +6 kcal/mol and a beautifully thermostable protein. You build it. It is barely better than the best single, or worse.

This is the most expensive mistake in stability engineering, and it has a name: epistasis — the non-additivity of mutational effects, where the impact of one mutation depends on the presence of others.

The additivity of mutational effects holds approximately when sites are distant and structurally independent, and breaks down when they interact (Wells, 1990). Most stability predictors score single mutants against the wild-type background and, implicitly or explicitly, assume you can sum them. They cannot see the coupling. Epistasis is precisely the term they omit, and it is a first-order driver of predictability failure in both protein evolution and design (Miton & Tokuriki, 2016).

Massively parallel stability measurements make the point at scale. From high-throughput assays of designed and natural proteins (Rocklin et al., 2017) to mega-scale datasets of hundreds of thousands of variants with directly measured folding free energy (Tsuboyama et al., 2023), the same pattern holds: additive models capture much of the single-mutant signal but systematically miss the coupling between mutations that share structural context. Combinations are their own experiments.

Practical Rules for Combinations

  1. Lead with singles. Rank, shortlist, and validate single mutants first. Your single-mutant predictions are your most reliable currency.
  2. Combine distant, independent sites. Two stabilizing singles are most likely to add cleanly when they are spatially separated and functionally unrelated. Be most suspicious when candidates are within a few residues of each other or share contacts — that is where sign epistasis and steric clashes live.
  3. Expect diminishing returns. Even non-interacting stabilizing mutations tend to show sub-additive stacking as you approach the protein's tolerance ceiling. A predicted sum is an upper bound, not an estimate.
  4. Treat every combination as a hypothesis. Predict it, but validate it before you build on it. A combination that underperforms its parts is telling you those positions are coupled — useful information for the next round.

Don't Stabilize Away the Function

A mutation can be thermodynamically excellent and biologically useless. The two most common ways to win on stability and lose on the protein:

  • Active-site and interface residues are often stability-suboptimal by design. Catalytic and binding residues are conserved for function, not folding, and many of them are locally destabilizing. "Fixing" them frequently means breaking the chemistry. A stabilizing mutation that fills or rigidifies a pocket can quietly abolish binding or catalysis.
  • PTM sites and functional motifs are one substitution away from deletion. A mutation that removes a phosphosite, disrupts an N-glycosylation sequon, or eliminates a functionally important disordered segment changes the biology even when it improves the melt. Domain-wide mutagenesis studies show that stabilizing substitutions are real but a minority of all mutations, and that stability and function pull in different directions across a surprising fraction of positions (Nisthal et al., 2019).

Before you promote any candidate, check the functional deltas alongside the stability deltas:

  • Binding-site preservation: does the mutation sit in or near a predicted ligand/interface pocket?
  • PTM changes: does it add or remove a modification site?
  • Disorder and topology shifts: for membrane proteins, does it alter predicted transmembrane topology? For any protein, does it collapse a functionally disordered region?

Then, for your top few, validate the structure. Predict the variant's fold and check that the pocket geometry near the active site is preserved — pocket RMSD versus wild-type, and local confidence (pocket pLDDT/PAE) high enough to trust the comparison. A stabilizing mutation that distorts the binding site is not a win you get to keep.


A Decision Framework for Ranking Mutations

Put it together into a repeatable workflow. Rank by free energy, corroborate with the melt, guard the function, then validate the survivors.

STABILITY-RANKING DECISION TREE
│
├─ 0. Confirm sign convention
│     └─ Which sign of ΔΔG means "stabilizing" in YOUR tool?
│        Get this wrong and every step below inverts.
│
├─ 1. Rank primarily by ΔΔG (most stabilizing first)
│     └─ Free energy at a fixed reference T is your comparable currency.
│
├─ 2. Corroborate with ΔTm
│     ├─ ΔΔG and ΔTm agree in sign?  ── no ──▶ demote / re-check assay + convention
│     └─ ΔTm huge but ΔΔG weak?      ── yes ─▶ FLAG: likely inflation, do not promote on ΔTm alone
│
├─ 3. Apply functional guardrails (per candidate)
│     ├─ In/near a binding site or interface? ──▶ high risk to function
│     ├─ Adds/removes a PTM site?             ──▶ biology may change
│     └─ Shifts disorder or membrane topology?──▶ investigate before trusting
│
├─ 4. Prefer singles; treat combinations as hypotheses
│     └─ Combine only distant, independent sites; predicted sums are upper bounds.
│
├─ 5. Validate structure for the top candidates
│     └─ Predict variant fold; check pocket RMSD vs WT + pocket pLDDT/PAE.
│
└─ 6. Wet-lab confirm the shortlist BEFORE combining
      └─ Measure singles; only stack what you have measured.

And the checklist to run before you commit a synthesis order:

PRE-SYNTHESIS CHECKLIST
[ ] Sign convention confirmed (stabilizing = positive ΔΔG unfolding, or your tool's equivalent)
[ ] Ranked by ΔΔG, not by raw ΔTm magnitude
[ ] ΔTm agrees in sign with ΔΔG for every promoted candidate
[ ] Large-ΔTm / small-ΔΔG rows flagged and NOT auto-promoted
[ ] Functional deltas checked: binding site, PTM sites, disorder, topology
[ ] Combinations restricted to spatially distant, functionally independent sites
[ ] Structure validated for top candidates (pocket RMSD + local confidence)
[ ] Plan tests singles first; combinations gated on single-mutant results

Case Study: The +6°C Mutation That Wasn't

The following is an illustrative scenario, not a specific result — the pattern is common enough to be worth walking through.

Problem. A team is engineering a ~300-residue bacterial esterase for a diagnostic reagent that has to survive shipping without a cold chain. They generate 40 single-mutant candidates and score each for ΔTm and ΔΔG. Two candidates lead on different axes:

  • Candidate M1: ΔTm = +6.2°C, ΔΔG = +0.3 kcal/mol. Sits two residues from the catalytic serine.
  • Candidate M2: ΔTm = +2.8°C, ΔΔG = +1.6 kcal/mol. On a solvent-exposed loop, far from the active site.

Sorted by melting shift, M1 wins by a wide margin and goes to the top of the list.

Solution. The team ranks by ΔΔG instead, dropping M1 well down the order. The mismatch — big ΔTm, tiny ΔΔG — trips the flag: this is exactly the low-cooperativity/artifact signature, and M1's position abutting the catalytic serine raises a functional alarm. The binding-site delta confirms M1 sits inside the predicted catalytic pocket. Structure validation of M1 shows a measurable pocket-RMSD shift versus wild-type. M2 and three other distant, ΔΔG-led candidates pass the functional guardrails cleanly and go forward as singles. The team explicitly declines to build the "stack all five" combination up front.

Outcome. In the assay, M1's melting shift is smaller than predicted and the esterase has lost most of its activity — the stabilizing substitution rigidified the catalytic pocket, precisely the risk the guardrails flagged. M2 and two of the distant singles confirm as modest, real stabilizers. When the team combines the two most spatially separated singles, the effect is close to additive and holds up. A naive triple combination that included a third, nearby candidate underperforms the sum of its parts — textbook epistasis. Net result: the ΔΔG-led ranking plus functional guardrails steered three synthesis slots onto real, function-preserving wins, and the one mutation that looked best by raw ΔTm would have burned a slot and cost a month.

The lesson is not "ΔTm is bad." It is that the biggest melting shift in the table is the number most likely to mislead you, and the discipline of ranking by free energy, corroborating, and checking function is what separates a shortlist from a wish list.


The Economics of Chasing a Mirage

Why does this ranking discipline matter enough to formalize? Because every candidate you promote spends real resources, and the mirage mutations cost exactly as much as the real ones.

A single variant is not free to test. Gene synthesis for a ~1 kb construct runs on the order of low hundreds of euros; expression, purification, and a stability readout add consumables plus days to a couple of weeks of hands-on time per variant. A 40-variant round is a month or more of a scientist's bench time and a meaningful reagent budget before you learn anything.

Now weigh two strategies:

  • Rank by raw ΔTm, no guardrails. You promote the +6.2°C candidate. Three to four weeks later you have a destabilized-in-practice, function-dead protein and a spent slot. Do that across a campaign and a predictable fraction of your bench budget evaporates on artifacts.
  • Rank by ΔΔG, corroborate, guard function, triage early. You spend a few hours of computational triage to cut 40 candidates to a shortlist of 8–10 that are stabilizing on the transferable metric and function-preserving. You test those. Your hit rate per bench slot goes up, and the slots you do spend are on candidates that can actually ship.

The asymmetry is the whole argument. Computational triage is cheap and fast; a wasted synthesis-express-purify-assay cycle is neither. You are not trying to replace the wet lab — you are trying to make sure the wet lab spends its slots on the candidates most likely to be real. Given that stabilizing mutations are a minority of all mutations to begin with, front-loaded triage is how you avoid testing your way through a haystack one straw at a time.


The Bottom Line

QuestionAnswerWhich metric
"Which number do I sort by?"Free energy is the transferable currencyRank by ΔΔG
"Do I trust this big melting shift?"Not on its own — corroborateΔTm as a check, flag large-ΔTm/small-ΔΔG rows
"How exact are these numbers?"Directional, limited-resolutionTreat both as a ranked shortlist, not truth
"Can I just add up my best singles?"No — epistasis breaks additivityValidate combinations; lead with singles
"Did I break the protein?"Check before you celebrate stabilityFunctional deltas + structure validation
"When is ΔTm the primary metric?"When your product target is thermal toleranceWeight ΔTm to your application

The single most important habit: never rank a mutation list by raw ΔTm magnitude alone. Rank by ΔΔG, read ΔTm as corroboration, guard the function, and treat every number as a hypothesis you are about to test — not a result you already have.


How Orbion Puts Both Metrics In Front of You

This is the workflow the Stabilize module is built around. For every variant — single (A123G) or multi (A123G/K456R) — Orbion reports ΔTm via AstraDTM and ΔΔG via AstraDDG side by side, so you never have to choose which number to compute; you get both and can rank by the transferable one while corroborating with the melt. Import a whole panel at once with batch CSV upload (up to 100 variants), or enter candidates by hand.

The stability numbers arrive with the guardrails this guide argues for. Alongside ΔTm and ΔΔG, Stabilize shows the biophysical deltas — Δ disorder and Δ amyloidogenicity (via AstraUNFOLD), Δ pLDDT, and Δ membrane topology — and the functional-preservation signals that tell you what a mutation costs: Δ binding-site preservation and PTM sites gained or lost. The large-ΔTm mutation sitting inside your active-site pocket is flagged before you order it, not after. For your top candidates, structure validation predicts the variant fold and reports TM-bundle RMSD and pocket RMSD versus wild-type, plus pocket pLDDT and PAE, so you can confirm you have not distorted the site you need.

Running a larger campaign? The standalone Mutation Engine scans and combinatorially optimizes stabilizing candidates in genetic-algorithm campaigns, then exports a ranked shortlist straight back into Stabilize via CSV import — so the singles-first, validate-the-combinations discipline is built into the handoff.

None of this replaces the bench, and Orbion does not pretend otherwise — the predictions are a ranked shortlist to test, which is exactly how you should use any stability metric. What Orbion gives you is both numbers, the functional cost, and the structural check in one place, so the shortlist you take to the lab is the one most likely to survive it.

Ready to rank a real mutation panel? Upload your sequence to Orbion, drop in your variants, and see ΔΔG, ΔTm, and the functional guardrails side by side — before you spend a single synthesis slot.


References

  1. Becktel WJ, Schellman JA. (1987). Protein stability curves. Biopolymers, 26(11):1859–1877. https://doi.org/10.1002/bip.360261104

  2. Matthews BW, Nicholson H, Becktel WJ. (1987). Enhanced protein thermostability from site-directed mutations that decrease the entropy of unfolding. Proceedings of the National Academy of Sciences, 84(19):6663–6667. https://doi.org/10.1073/pnas.84.19.6663

  3. Robertson AD, Murphy KP. (1997). Protein structure and the energetics of protein stability. Chemical Reviews, 97(5):1251–1268. https://doi.org/10.1021/cr960383c

  4. Wells JA. (1990). Additivity of mutational effects in proteins. Biochemistry, 29(37):8509–8517. https://doi.org/10.1021/bi00489a001

  5. Miton CM, Tokuriki N. (2016). How mutational epistasis impairs predictability in protein evolution and design. Protein Science, 25(7):1260–1272. https://doi.org/10.1002/pro.2876

  6. Usmanova DR, Bogatyreva NS, Ariño Bernad J, Eremina AA, Gorshkova AA, Kanevskiy GM, Lonishin LR, Meister AV, Yakupova AG, Kondrashov FA, Ivankov DN. (2018). Self-consistency test reveals systematic bias in programs for prediction of change of stability upon mutation. Bioinformatics, 34(21):3653–3658. https://doi.org/10.1093/bioinformatics/bty340

  7. Nisthal A, Wang CY, Ary ML, Mayo SL. (2019). Protein stability engineering insights revealed by domain-wide comprehensive mutagenesis. Proceedings of the National Academy of Sciences, 116(33):16367–16377. https://doi.org/10.1073/pnas.1903888116

  8. Nikam R, Kulandaisamy A, Harini K, Sharma D, Gromiha MM. (2021). ProThermDB: thermodynamic database for proteins and mutants revisited after 15 years. Nucleic Acids Research, 49(D1):D420–D424. https://doi.org/10.1093/nar/gkaa1035

  9. Stourac J, Dubrava J, Musil M, Horackova J, Damborsky J, Mazurenko S, Bednar D. (2021). FireProtDB: database of manually curated protein stability data. Nucleic Acids Research, 49(D1):D319–D324. https://doi.org/10.1093/nar/gkaa981

  10. Pancotti C, Benevenuta S, Birolo G, Alberini V, Repetto V, Sanavia T, Capriotti E, Fariselli P. (2022). Predicting protein stability changes upon single-point mutation: a thorough comparison of the available tools on a new dataset. Briefings in Bioinformatics, 23(2):bbab555. https://doi.org/10.1093/bib/bbab555

  11. Rocklin GJ, Chidyausiku TM, Goreshnik I, Ford A, Houliston S, Lemak A, Carter L, Ravichandran R, Mulligan VK, Chevalier A, Arrowsmith CH, Baker D. (2017). Global analysis of protein folding using massively parallel design, synthesis, and testing. Science, 357(6347):168–175. https://doi.org/10.1126/science.aan0693

  12. Tsuboyama K, Dauparas J, Chen J, Laine E, Behbahani YM, Weinstein JJ, Mangan NM, Ovchinnikov S, Rocklin GJ. (2023). Mega-scale experimental analysis of protein folding stability in biology and design. Nature, 620:434–444. https://doi.org/10.1038/s41586-023-06328-6