Blog
Expression & PurificationAI & Computation

Which PTMs Actually Survive in Your Expression Host? How to Read a PTM Prediction Before You Pick E. coli or HEK293

Aug 7, 2026 · 26 min read

You ran your sequence through a PTM predictor and got back a clean list: two N-glycosylation sequons, a cluster of phosphosites in a cytoplasmic loop, three disulfides, a tyrosine that scores as a sulfation site. It looks like a spec sheet. So you open the E. coli vector that's already on your bench — it's Friday, E. coli is fast — and you clone. Three weeks later you have inclusion bodies, a binding assay that reads a tenth of what the literature says, and a protein whose intact mass doesn't match anything you designed. The prediction wasn't wrong. You just read it as a promise instead of a question.

A predicted PTM site is not a modification. It's a candidate — a place on the sequence where a modification could go, if the cell that makes the protein carries the machinery to install it. Your expression host decides that. Pick E. coli for a protein whose function rides on a complex sialylated glycan and every glycosite on that tidy list evaluates to nothing. Pick yeast and you'll get a glycan — the wrong one, high-mannose and potentially immunogenic. The gap between "the predictor flagged a site" and "my protein carries that modification" is the entire subject of this article, and it's where a surprising number of expression campaigns quietly die.

This is the reverse of the question we answered in How to Choose Expression Systems for Your Protein's PTM Requirements. That piece starts from a requirement — "my protein needs glycosylation" — and walks you up the host ladder. This one starts from the artifact you actually have in front of you: a prediction, full of sites, most of which you can't yet tell apart. How do you read it? And how do you know, before you clone, which of those sites your candidate host will actually produce?


Key Takeaways

  • A predicted site clears three gates, not one. The predictor answers only the first: is the sequence motif there? It says nothing about whether the site is occupied in vivo, or whether your host can install it. This article is about that third gate — host competence — because it's the one you control at cloning.
  • E. coli installs zero N- or O-glycosylation, no sialylation, no tyrosine sulfation, and no γ-carboxylation. For a glycoprotein, every predicted glycosite evaluates to "none" — no amount of codon optimization or IPTG titration changes that.
  • Yeast and insect cells do glycosylate — in non-human forms. Yeast adds high-mannose/hypermannose N-glycans; insect cells add truncated paucimannose glycans, largely unsialylated. "Installed" is not "correct": these forms can be immunogenic or cleared rapidly in vivo.
  • Only mammalian cells (HEK293/CHO) install the full stack — complex sialylated glycans, γ-carboxylation, tyrosine sulfation, and furin-type pro-domain processing. That's why the majority of approved biologics are produced in CHO.
  • One missing modification can zero out function. An aglycosylated IgG Fc loses effector function; an uncarboxylated clotting factor is inactive; an unsulfated receptor N-terminus can lose one to two orders of magnitude of binding affinity.
  • Predictions flag candidates, not occupancy. They can't tell you stoichiometry, and they can't see host-cell state. Confirm the modifications that actually landed by intact mass spectrometry and released-glycan analysis before you trust any downstream number.

Prediction ≠ Occupancy ≠ Host-Competence

The single most useful habit when you read a PTM prediction is to split one question into three. They fail independently, and conflating them is what puts a glycoprotein into E. coli.

Gate 1 — Is the motif there? (What the predictor answers.) Most PTMs have a sequence and/or structural context that a model can recognize: the N-X-S/T sequon (X ≠ Pro) for N-glycosylation, Ser/Thr/Tyr in a kinase-recognizable context for phosphorylation, paired cysteines with compatible geometry for a disulfide, a Glu-rich Gla domain for γ-carboxylation, an acidic-flanked tyrosine for sulfation. A predictor scores your sequence for these. A high-confidence hit means "this looks like a real site." It does not mean the site is used.

Gate 2 — Would it be occupied in the native cell? (Biology, not sequence.) Even in the organism the protein comes from, a motif is not automatically modified. N-glycosylation sequons are occupied variably — some essentially always, some only fractionally, depending on local structure and how accessible the site is to oligosaccharyltransferase as the chain threads through the ER. Phosphosites are frequently conditional and substoichiometric: present only in a signaling state, on a fraction of molecules. So "predicted" sites include a mix of constitutive modifications, situational ones, and false positives. Occupancy is a measurement, not a prediction.

Gate 3 — Can your host install it? (The one you choose.) This is the gate this article is about, because it's the one you set the moment you pick an expression system. A site that is real (Gate 1) and normally occupied (Gate 2) still comes out bare if the host lacks the enzymes, the compartment, or the cofactor to make it. E. coli has no oligosaccharyltransferase complex for eukaryotic N-glycosylation and a reducing cytoplasm that won't form disulfides; no reading of the prediction changes that biochemistry.

The predictor clears the first gate. Your read of the protein's biology handles the second. Your host choice — the subject of the rest of this piece — determines the third.


How to Read a PTM Prediction, Line by Line

Before you filter by host, understand what each predicted line is actually claiming. The types behave very differently, and they don't all matter equally to your host decision. Predictors return a long, uneven list because the modification landscape itself is uneven: across the curated Swiss-Prot proteome, a few types dominate — phosphorylation most of all, with glycosylation and acetylation among the next most frequent — atop a long tail of rarer modifications (Khoury et al., 2011). The common types fill most of your prediction; the rare ones, when they show up, are often the load-bearing ones.

  • N-glycosylation sequon. The strongest sequence signal of the lot — N-X-S/T is a crisp motif — which is exactly why it's over-called. A flagged sequon is necessary but not sufficient: buried or fast-folding sequons often go unoccupied. Treat it as "watch this site," and know that which host you pick changes not just whether a glycan appears but what glycan appears.
  • O-glycosylation site. Weaker sequence determinants (no strict consensus for mucin-type GalNAc O-glycosylation), so predictions are noisier and context-dependent. Host form matters enormously here too — yeast performs O-mannosylation, not the human mucin-type pathway.
  • Phosphorylation (Ser/Thr/Tyr). Usually the longest part of the list and the most conditional. Many are regulatory and substoichiometric. Ask whether the phospho-state matters for what you're making the protein to do — a structural study of the folded core may not care; a kinase-substrate assay absolutely does.
  • Disulfide bonds. Predicted from paired cysteines and geometry. The host question collapses to one thing: oxidizing compartment or not. (The geometry question — whether a predicted pair should form at all — is its own trap; see Disulfide Bond Engineering: When to Add, Remove, or Redesign.)
  • Lipidation (myristoylation, palmitoylation, prenylation, GPI anchor). Sequence-motif-driven and eukaryote-specific. Relevant when membrane association or trafficking is part of the function.
  • γ-carboxylation, tyrosine sulfation, C-terminal amidation. The specialist modifications. Rare in the proteome, but when your protein has them they're usually load-bearing — and only some hosts can do them at all.
  • Signal-peptide cleavage and pro-domain processing. Not decorations but maturation events — the difference between a pro-form and the active protein. Host proteases decide whether this happens correctly.

A note on what to ignore for host selection: dynamic intracellular regulatory PTMs — ubiquitination, SUMOylation, most transient phosphorylation — are rarely something you reproduce on a purified recombinant product. They mark proteins for degradation or relocalization in the living cell; they're not a stable feature of the material coming off your column. Note them, but don't let them drive host choice — decide the host on modifications that both persist on the purified product and matter for its structure, activity, safety, or mass.


The Host-Competence Filter

Here is the filter to lay over your prediction. Each row is a modification your predictor might flag; each column is what that host actually does with it. Read down the column of the host you're considering and mark every predicted site that comes back Partial or None — those are the sites that won't be there, or won't be right.

Predicted PTME. coliYeast (P. pastoris / S. cerevisiae)Insect (Sf9/Hi5, BEVS)Mammalian (HEK293/CHO)
N-glycan installed at allNoneFullFullFull
Complex, sialylated, human-type N-glycanNoneNonePartialFull
O-glycosylation (mucin-type, GalNAc)NonePartial*PartialFull
Sialylation (terminal)NoneNonePartialFull
Disulfide bondsPartial†FullFullFull
Phosphorylation (eukaryotic S/T/Y)None‡FullFullFull
N-myristoylationNoneFullFullFull
S-palmitoylationNoneFullFullFull
Prenylation (farnesyl/geranylgeranyl)NoneFullFullFull
GPI anchorNoneFull§FullFull
N-terminal acetylationRareFullFullFull
Lysine acetylation / methylationPartialFullFullFull
γ-carboxylation (Gla domains)NoneNoneNoneFull¶
Tyrosine sulfationNoneNonePartialFull
C-terminal amidationNoneNoneFullFull
Signal-peptide cleavageFullFullFullFull
Pro-domain / furin-type processingNonePartial**PartialFull

Legend. Full = installed in a native-like form. Partial = installed but in a non-human form, at low/variable levels, or only with special measures. None = not installed.

* Yeast performs O-mannosylation, not the human mucin-type GalNAc pathway — a predicted mammalian O-glycosite will not be reproduced faithfully. † E. coli cannot form disulfides in its reducing cytoplasm; you need periplasmic export (a signal peptide plus the Dsb machinery) or engineered strains (SHuffle, Origami). See the disulfide engineering guide. ‡ Bacterial phospho-signaling exists but does not reproduce the eukaryotic Ser/Thr/Tyr kinome — treat predicted human phosphosites as absent in E. coli. § Yeast makes GPI anchors with a different terminal structure than mammalian cells. ¶ γ-carboxylation depends on the vitamin-K-dependent carboxylase and can be incomplete even in mammalian cells when the protein is overexpressed and the enzyme is saturated (Berkner, 2005). ** Yeast Kex2 processes some dibasic pro-sites, but this is not a general substitute for the mammalian proprotein-convertase (furin/PC) family.

The two most important lessons live in the first four rows. First: "installed" is not "correct." Yeast and insect cells both return Full for "N-glycan installed at all" — and both return None or Partial for "complex, sialylated, human-type." That single distinction is why a glycoprotein can express beautifully in P. pastoris, pass every in-house assay, and still fail as a therapeutic. Second: the specialist modifications collapse to one column. γ-carboxylation, tyrosine sulfation, and full complex glycosylation are effectively mammalian-only. If your prediction flags any of them and your protein's function depends on them, the host decision is already made.

This filter is exactly the operation Orbion's Characterization module automates — mapping each predicted PTM onto a per-host full / partial / none rating across E. coli, yeast, insect, and mammalian — but you should understand the biochemistry underneath it, because the consequences of a mismatch are what actually cost you.


Glycosylation: The Modification Most Likely to Break on the Wrong Host

Glycosylation is where host choice bites hardest, because it's common, it's often required for folding and function, and — uniquely among PTMs — the host doesn't just decide whether it happens but what structure you get. Four hosts, four different answers to the same predicted sequon:

  • E. coli — nothing. No N- or O-glycosylation machinery. A secreted eukaryotic glycoprotein that needs its glycans to fold will often go straight to inclusion bodies here. (When that happens, see Your Protein Expressed as Inclusion Bodies — but the more useful move is to not pick E. coli for a glycoprotein in the first place.)
  • Yeast — high-mannose, and a lot of it. Yeast installs N-glycans but extends them into high-mannose and hypermannose structures; S. cerevisiae can add long, immunogenic outer chains with terminal α-1,3-mannose. Engineering glycosylation in yeast toward human-like structures is an active field precisely because the default output is wrong for human use (De Pourcq et al., 2010).
  • Insect — paucimannose. The baculovirus-insect system trims N-glycans to short paucimannose structures (roughly Man3GlcNAc2, often core-fucosylated) and installs little to no terminal sialic acid without metabolic engineering (Shi & Jarvis, 2007). Fine for many structural-biology and functional applications; wrong when you need a sialylated, human-type glycan.
  • Mammalian — complex and sialylated. HEK293 and CHO build complex, branched, sialylated N-glycans close to the human repertoire, which is why the majority of approved recombinant glycoprotein therapeutics come from these hosts (Lalonde & Durocher, 2017; Dumont et al., 2016; Kim et al., 2012).

The therapeutic case that makes this concrete is the IgG Fc N-glycan at Asn297. That single conserved glycan modulates Fc-receptor binding and, through it, antibody effector functions — ADCC and CDC. Its fine structure tunes potency: reducing core fucosylation, for example, markedly increases ADCC (Jefferis, 2009). Express that antibody or Fc-fusion in E. coli and the Fc comes out aglycosylated — folded, perhaps, but effector-dead. Express it in a system that adds high-mannose glycans and you've changed its pharmacokinetics and potentially its immunogenicity. The predictor flags "N-glycosylation site at Asn297" identically in every case; the host decides whether that flag becomes a functional glycan, a wrong glycan, or nothing. For the deeper biophysics of why glycans matter beyond this one site, see Glycosylation Matters More Than You Think.

The choice of expression system, in short, is the choice of glycan — a point made forcefully two decades ago and still routinely underappreciated (Brooks, 2004; Walsh & Jefferis, 2006).


Disulfides, Phosphorylation, and the Reducing-Cytoplasm Problem

Two more predicted-PTM classes hinge on host biology in ways worth reading carefully.

Disulfides. If your prediction pairs cysteines into disulfides, the host question is whether the folding compartment is oxidizing. The E. coli cytoplasm is reducing — disulfides won't form there — so a disulfide-bonded protein needs periplasmic export (an N-terminal signal peptide such as pelB, OmpA, or DsbA, which routes the protein to the oxidizing periplasm and its Dsb catalysts) or a specialized cytoplasmic strain (SHuffle, Origami). Yeast, insect, and mammalian secretory pathways all form disulfides natively. Two cautions: a predicted disulfide is not always a disulfide that should form — geometry and dynamics decide, and forcing a bad one creates cross-linked aggregates — and even the right host only gives the bonds a chance to form correctly, not a guarantee. The full decision logic is in the disulfide engineering guide.

Phosphorylation. Predicted eukaryotic Ser/Thr/Tyr phosphosites are usually the longest and most conditional part of a prediction. E. coli does not carry the eukaryotic kinome and will not install them; yeast, insect, and mammalian cells can, though which specific sites get phosphorylated depends on which kinases are active. Before you let a phosphosite drive host choice, ask whether the phospho-state matters for your purpose. Making the protein to crystallize its folded domain? The absence of a flexible-loop phosphosite is likely irrelevant — and may even help you get a homogeneous, crystallizable sample. Making it to run a phospho-dependent binding or activity assay? Then a host that can't install the site will hand you a protein that reads as inactive for reasons that have nothing to do with your actual question.


The Modifications Only Mammalian Cells Install

Three predicted PTMs act as near-automatic host selectors. If your protein genuinely needs them, everything below mammalian is off the table.

  • γ-carboxylation. Vitamin-K-dependent γ-carboxylation converts specific glutamate residues in a Gla domain to γ-carboxyglutamate, creating the calcium-binding sites that let clotting factors (Factor VII, IX, X, prothrombin, protein C) dock onto membranes. It requires the vitamin-K-dependent carboxylase, which is present in mammalian cells and absent from bacteria, yeast, and insect cells. Miss it and the protein expresses fine and clots nothing. Even in mammalian hosts the modification can be incomplete when the carboxylase is saturated by overexpression (Berkner, 2005) — a textbook case of a modification you must confirm, not assume.
  • Tyrosine sulfation. Sulfation of tyrosine residues, installed by tyrosylprotein sulfotransferases in the trans-Golgi, controls high-affinity interactions in many secreted and membrane proteins — chemokine-receptor N-termini, adhesion molecules, some coagulation factors (Stone et al., 2009). Mammalian cells do it well; insect cells do it weakly and variably; yeast and E. coli not at all. An unsulfated receptor N-terminus can lose one to two orders of magnitude of binding affinity — enough to make a perfectly good binder look dead.
  • Complex proteolytic maturation. Furin and the proprotein-convertase family cleave pro-domains at specific basic sites to activate many growth factors, hormones, and receptors. E. coli can't do it; yeast Kex2 covers some dibasic sites but isn't a general substitute; mammalian cells do it natively. A prediction that flags a pro-domain cleavage is flagging the difference between an inactive pro-form and the mature protein.

Honesty check: mammalian is the most capable host, not a guarantee of correctness. CHO can add non-human sialic acid (Neu5Gc) and α-Gal epitopes that raise immunogenicity concerns; HEK293 and CHO glycosylate somewhat differently; γ-carboxylation can run incomplete. "Use mammalian" narrows the problem — it doesn't close it. You still confirm.


Reading a Mismatch: What an Absent or Wrong PTM Does to Your Protein

When your host can't install a predicted, functionally relevant PTM, the failure rarely announces itself as "missing PTM." It shows up downstream, wearing a disguise. Learn the disguises.

  • Loss of activity or effector function. Aglycosylated Fc → no ADCC/CDC. Uncarboxylated clotting factor → no membrane binding, no coagulation. Unsulfated receptor N-terminus → collapsed affinity. The protein looks fine on a gel and fails at its job.
  • Misfolding and aggregation. Many secreted eukaryotic proteins need their N-glycans to fold. Strip them (E. coli) and you get inclusion bodies or aggregation-prone material — a folding failure that's really a PTM failure.
  • Altered pharmacokinetics. Glycan structure drives clearance. High-mannose glycans (yeast, and the terminal mannose exposed on some insect/mammalian species) are recognized by the mannose receptor and cleared faster; missing sialylation shortens half-life. Same polypeptide, different residence time in the body.
  • Immunogenicity. Non-human glycans — yeast hypermannose with terminal α-1,3-mannose, insect core structures, CHO-derived Neu5Gc/α-Gal — are potential immunogens in a therapeutic context (Brooks, 2004; De Pourcq et al., 2010).
  • The wrong mass — and a confusing assay. Absent or heterogeneous glycans mean your intact mass doesn't match the sequence you designed, complicating QC and identity confirmation. Worse, a missing phospho- or sulfo-site can make a functional assay read "negative" when the molecule is simply the wrong version of itself — you conclude the candidate is weak and move on, when the host was the problem.

The through-line: a host×PTM mismatch converts into a different-looking problem — a solubility problem, a potency problem, a PK problem, a QC problem — and teams burn weeks chasing the symptom. Reading the prediction against the host up front is how you skip that detour.


The Workflow: Predict → Filter → Decide → Confirm

Four steps turn a PTM prediction into a host decision you can defend.

  1. Predict the PTM sites on your sequence — the complete candidate landscape, not just the two sites UniProt happens to list.
  2. Filter each predicted site by your candidate host's competence (the table above): mark every site that comes back Partial or None.
  3. Decide, site by site, among three moves:
    • Accept — the host can't install it, but your application doesn't depend on it (e.g., a flexible-loop phosphosite for a structural study).
    • Switch host — a functionally critical modification is Partial/None; move up to a host that can install it.
    • Engineer the site out — an unwanted predicted site (e.g., a surface N-glycosylation sequon near a binding interface causing heterogeneity) can be removed by mutation (N→Q knocks out an N-X-S/T sequon) so it never complicates the product.
  4. Confirm what actually landed — because Gates 1 and 2 are still just predictions. Intact mass spectrometry, peptide mapping, and released-glycan analysis are how you verify occupancy and structure (Beck et al., 2013).
READ EACH PREDICTED PTM SITE
        │
        ▼
Is this modification required for what you are
making the protein to DO (activity, binding, PK,
safety, correct mass)?
        │
   ┌────┴─────────────────────────────┐
   │ NO                               │ YES
   ▼                                  ▼
Is the site UNWANTED               Can your CANDIDATE host
(heterogeneity, wrong-site         install it in the CORRECT
mod, interface clash)?             FORM?  (check the filter)
   │                                  │
 ┌─┴──┐                        ┌──────┴───────┐
 │NO  │YES                     │ YES          │ NO / PARTIAL
 ▼    ▼                        ▼              ▼
Leave  Engineer the      Host is viable   SWITCH HOST
it     site out          for this site.   up the ladder
       (e.g. N→Q)        Keep it.         until it's FULL
        │                     │              │
        └──────────┬──────────┴──────────────┘
                   ▼
   Repeat for every predicted site, then choose the
   SIMPLEST host that returns FULL on all the
   modifications you decided you need.
                   ▼
   CONFIRM on the expressed protein: intact MS +
   released-glycan / peptide mapping. Occupancy and
   stoichiometry are measured, never assumed.

Pre-clone checklist:

  • Predicted the full PTM landscape from sequence — not just the experimentally annotated subset.
  • Classified each site: structural/decorative vs. functionally load-bearing for your application.
  • Filtered every load-bearing site against the candidate host (full / partial / none).
  • For each Partial/None on a load-bearing site, made an explicit call: accept, switch host, or engineer out — none left unresolved.
  • Flagged unwanted predicted sites (interface-adjacent sequons, heterogeneity sources) for removal.
  • Chose the simplest host that returns Full on everything you need — no more machinery than the protein requires.
  • Planned the confirmation assay (intact MS + glycan/peptide analysis) before expressing, so occupancy gets measured, not assumed.

Case Study: An Fc-Fusion That Was "Weak" — Until It Changed Hosts

An illustrative scenario — a composite of common Fc-fusion failure modes, not a specific customer case.

Problem. A team was developing an Fc-fusion: a receptor ectodomain fused to a human IgG1 Fc. The PTM prediction was busy — the conserved Fc N-glycosylation sequon at Asn297, three additional N-glycosylation sequons on the ectodomain, several disulfides, and a cluster of tyrosines in the ectodomain scoring as sulfation sites. Under deadline to get material for a binding assay, they expressed it in E. coli using a vector already validated in the lab. The protein came off largely as aggregate; the fraction they could purify bound its ligand roughly an order of magnitude weaker than the literature suggested and showed no measurable effector activity. The read in the team meeting: "the molecule is weak, deprioritize."

Analysis. Read against a host-competence filter, the E. coli result was not a verdict on the molecule — it was a verdict on the host. Every functionally critical predicted PTM was None or Partial in E. coli: the Asn297 complex glycan that drives Fc effector function (None), the tyrosine sulfation the ectodomain uses for high-affinity ligand binding (None), and the multiple disulfides that need an oxidizing compartment (Partial at best, cytoplasmic). The aggregation, the collapsed affinity, and the dead effector arm were all predictable from the prediction. The molecule had never been given a chance to be itself. The filter also flagged one ectodomain N-glycosylation sequon sitting near the binding interface as a likely source of heterogeneity worth removing.

Solution. Re-host to CHO, which returns Full on all three critical modifications — complex sialylated Asn297 glycan, tyrosine sulfation, and correct disulfide formation through the secretory pathway. Use the predicted sites to design the QC up front: intact mass to confirm the expected glycan mass, released-glycan analysis on Asn297, and a sulfation-specific check on the ectodomain. Engineer out the interface-adjacent N-sequon (N→Q) to cut heterogeneity, keeping the two sequons that don't interfere.

Outcome. In CHO, the Fc-fusion recovered high-affinity binding and measurable effector function, and intact MS confirmed the Asn297 glycan and the predicted occupied sites. The "weak molecule" had been an artifact of host choice the whole time. The cost of learning that the hard way was several weeks and a candidate that came within a meeting of being shelved — all avoidable by reading the prediction against the host before the first clone, not after the first assay.


The Economics of Getting Host × PTM Wrong

The compute to predict and filter is trivial. The wet-lab work built on a bad host choice is not — and the bill usually lands late, after you've invested in the wrong material.

MismatchWhat you see at the benchReal cost
Glycoprotein in E. coliInclusion bodies or inactive, aggregation-prone proteinWeeks of refolding that can't restore glycans anyway; restart in a new host
Human therapeutic in yeast/insectActive in vitro; immunogenic or fast-cleared in vivoA late, expensive failure in animal or clinical studies
Clotting factor without γ-carboxylationExpresses well, zero coagulation activityA whole program chasing a "bad construct" that was a bad host
Unsulfated receptor / ligandBinding affinity off by 10–100×Good candidates mis-ranked or killed on artifact data
Uninstalled phospho-site in an assayAssay reads "inactive" / no regulationMisleading SAR, false negatives, a molecule wrongly deprioritized
Absent or heterogeneous glycan massIntact MS doesn't match the designQC failure, batch rejection, investigation time

The pattern is the one that recurs across expression work: minutes of reading a prediction against a host determine months of downstream experiments. The expensive mistake is never the prediction — it's committing cells, reagents, and calendar to a host the prediction already told you was wrong.


Bottom Line

A predicted PTM site is a candidate, not a modification. Whether it becomes a real modification — and whether that modification is the right one — is decided by the host you clone into. Read every predicted site through three gates: is the motif there (the predictor's job), would it be occupied in vivo (biology), and can your host install it in the correct form (your choice at cloning). E. coli installs none of the eukaryotic glycosylation, sulfation, or γ-carboxylation your prediction may flag; yeast and insect cells glycosylate but in non-human forms; only mammalian cells install the full human stack — which is why the hardest-to-make biologics are made in CHO. Filter your prediction against the host before you clone, decide each load-bearing site (accept, switch, or engineer out), pick the simplest host that returns Full on everything you need, and then confirm what actually landed by mass spec. The prediction points; the host delivers; the MS proves it.


How Orbion Helps

Everything above is a filtering operation: predict the sites, then cross them against what each host can build. Orbion's Characterization module runs both halves in one place so you're reading a decision, not assembling one from scattered servers.

Paste a sequence and the Astra models map the full modification landscape onto it:

  • AstraPTM predicts 39 PTM types at residue level — the site-by-site view you filter — and AstraPTM-Mini gives a 40-type protein-level summary so you can see the modification profile at a glance before drilling in.
  • PTM expression-system filtering is the differentiator for this exact problem: it takes those predicted PTMs and rates them by expression-system compatibility — E. coli, yeast, insect, mammalian — using a compatibility matrix across 30+ PTM types with full / partial / none ratings per host. That's the host-competence filter from this article, applied automatically to your sequence: you see, before you clone, which predicted modifications your chosen host will actually produce and which come back partial or absent.
  • AstraUNFOLD reads transmembrane topology, disorder, and aggregation propensity, and AstraSUIT predicts experimental suitability and host association — the surrounding context (a membrane protein or a disulfide-rich secreted protein) that also pushes your host decision.

From there, the call feeds straight into the Bench module: choose an expression system (E. coli, insect, mammalian, cell-free, or auto-recommend) and generate host-specific, construct-aware protocols — with codon optimization tuned to the organism you picked (E. coli, HEK293, CHO, Sf9, Pichia, or S. cerevisiae). Predict the sites, filter by host, and hand the winning host straight to protocol generation.

The honest framing is the one this article argues for: the filter tells you which predicted PTMs a host can install — it does not measure occupancy or stoichiometry. Predictions flag candidate sites; confirm what actually landed on the expressed protein by intact mass spectrometry and glycan analysis. Used that way — triage first, confirmation second — you stop discovering your host was wrong three weeks and one confusing assay too late. Start with your sequence at orbion.life.


References

  1. Walsh, G. & Jefferis, R. Post-translational modifications in the context of therapeutic proteins. Nature Biotechnology 24(10):1241–1252 (2006). doi:10.1038/nbt1252
  2. Khoury, G. A., Baliban, R. C. & Floudas, C. A. Proteome-wide post-translational modification statistics: frequency analysis and curation of the Swiss-Prot database. Scientific Reports 1:90 (2011). doi:10.1038/srep00090
  3. Brooks, S. A. Appropriate glycosylation of recombinant proteins for human use: implications of choice of expression system. Molecular Biotechnology 28(3):241–256 (2004). doi:10.1385/MB:28:3:241
  4. Jefferis, R. Glycosylation as a strategy to improve antibody-based therapeutics. Nature Reviews Drug Discovery 8(3):226–234 (2009). doi:10.1038/nrd2804
  5. De Pourcq, K., De Schutter, K. & Callewaert, N. Engineering of glycosylation in yeast and other fungi: current state and perspectives. Applied Microbiology and Biotechnology 87(5):1617–1631 (2010). doi:10.1007/s00253-010-2721-1
  6. Shi, X. & Jarvis, D. L. Protein N-glycosylation in the baculovirus-insect cell system. Current Drug Targets 8(10):1116–1125 (2007). doi:10.2174/138945007782151360
  7. Lalonde, M.-E. & Durocher, Y. Therapeutic glycoprotein production in mammalian cells. Journal of Biotechnology 251:128–140 (2017). doi:10.1016/j.jbiotec.2017.04.028
  8. Dumont, J., Euwart, D., Mei, B., Estes, S. & Kshirsagar, R. Human cell lines for biopharmaceutical manufacturing: history, status, and future perspectives. Critical Reviews in Biotechnology 36(6):1110–1122 (2016). doi:10.3109/07388551.2015.1084266
  9. Kim, J. Y., Kim, Y.-G. & Lee, G. M. CHO cells in biotechnology for production of recombinant proteins: current state and further potential. Applied Microbiology and Biotechnology 93(3):917–930 (2012). doi:10.1007/s00253-011-3758-5
  10. Stone, M. J., Chuang, S., Hou, X., Shoham, M. & Zhu, J. Z. Tyrosine sulfation: an increasingly recognised post-translational modification of secreted proteins. New Biotechnology 25(5):299–317 (2009). doi:10.1016/j.nbt.2009.03.011
  11. Berkner, K. L. The vitamin K-dependent carboxylase. Annual Review of Nutrition 25:127–149 (2005). doi:10.1146/annurev.nutr.25.050304.092713
  12. Beck, A., Wagner-Rousset, E., Ayoub, D., Van Dorsselaer, A. & Sanglier-Cianférani, S. Characterization of therapeutic antibodies and related products. Analytical Chemistry 85(2):715–736 (2013). doi:10.1021/ac3032355