Proteases are one of the most clinically validated enzyme classes in medicine — from the ACE and neprilysin inhibitors in cardiovascular disease to the HIV, HCV, and SARS-CoV-2 protease inhibitors in antiviral therapy and the DPP-4 inhibitors in metabolic disease. Yet the hard targets still fail early, and for concrete reasons: unstable constructs, ambiguous membrane context, catalytic metal sites and exosites that demand different chemistry, and paralogue families where one lead engages a dozen relatives. The class isn't unexplored — the difficulty is routing a hard target from sequence to a tractable first experiment.
So the practical question for a protease team is concrete: can AI say enough about a difficult protease — its family, its construct risks, and its pocket — to decide what structural biology to buy first?
To answer it, we ran the entire Astra AI Suite across 545 eligible reviewed human proteases from UniProt Swiss-Prot and scored every output against the strongest publicly available experimental reference. This post is the short version. The full per-model breakdown is the Proteases volume of the Orbion Model Performance Series, linked at the end.
Key takeaways
- 545 reviewed human proteases, with every model output scored against the strongest public experimental reference for its task.
- Function and family routing is reliable: 91.6% recognized as hydrolases (EC 3) and 90.1% carrying protease/peptidase GO terms — enough to route a sequence into protease-specific chemistry rather than a generic workflow.
- Membrane and cofactor context comes from sequence alone: membrane-class routing matches UniProt on 81.7% of the cohort, with a zinc cofactor called on 124 targets — the catalytic-metal signal a metalloprotease program needs early.
- Residue-level topology is strong: per-residue transmembrane agreement at AUROC 0.951 (median 0.995) across 109 annotated membrane proteases — a construct-boundary hypothesis from sequence.
- The construct-design PTMs land: N-linked glycosylation at F1 0.924 and disulfide bonds at F1 0.829 — the structural modifications that decide expression host and domain boundaries.
- Binding is a ranked hypothesis generator, reported honestly: 99 of 242 co-crystal-scored proteases reach a review-worthy panel (set-F1 ≥ 0.5), while exact top-1 ligand-ID recall is 0.053 — reported openly as the strictest possible governance check, not the product promise.
The benchmark
- Cohort. 545 eligible reviewed (Swiss-Prot) human proteases, from a 546-protease intake (one sequence failed a length eligibility check).
- References. Swiss-Prot annotation for sequence-level features (function, membrane class, topology, PTM sites); PDB co-crystal heavy-atom contacts at 4 Å for ligand binding, with buffer, cryoprotectant, and common-additive ligands excluded before scoring.
- Reporting. Each capability is shown with its headline metric and its failure modes — we don't average the weaknesses away.
One distinction matters for reading the numbers. The classification, topology, and PTM metrics describe how the deployed models behave on the human protease class in production, including proteins drawn from their training corpora; held-out generalization is reported in the per-model preprints. And unlike the Enzymes and GPCRs volumes, this volume makes no experimental thermostability claim — the protease evidence doesn't support a ΔTm statement, so topology and disorder outputs are used for construct and region triage only.
Function and family: 91.6%
The first job on a difficult protease is knowing it belongs in a protease workflow at all. AstraROLE2 recognizes 91.6% of the cohort as hydrolases (EC 3) and predicts the protease/peptidase GO term on 90.1% — a mechanism-routing signal strong enough to send a sequence into protease-specific structure modeling, construct design, and binding-panel triage. This is enrichment within a curated protease cohort rather than a test against non-proteases, and MEROPS-style clan assignment isn't claimed here — but as an intake filter, it reliably puts the target on the right track.
Membrane, cofactor, and topology: 81.7% and AUROC 0.951
Construct risk is where proteases quietly waste program budget — a membrane-anchored or zinc-dependent protease expressed as a naive soluble construct loses the very context the program cares about.
- Membrane class. AstraSUIT2's top membrane-class call matches the UniProt TM-feature-derived class on 445 of 545 proteases (81.7%), telling a program whether to build a soluble construct, a membrane-retaining one, or both.
- Cofactor. A metal-ion cofactor is predicted on 156 proteins, with zinc specifically called on 124 — the catalytic-metal context that separates a metalloprotease cleft from a serine or cysteine active site.
- Residue-level topology. AstraUNFOLD's per-residue membrane probability reaches AUROC 0.951 (median 0.995) against UniProt transmembrane annotations on 109 annotated membrane proteases — a construct-boundary hypothesis precise enough to decide where an anchor or ectodomain cut should fall.
PTM sites: 0.924 and 0.829
For secreted and membrane proteases, the modifications that matter most are the structural ones — they decide expression host, domain boundaries, and which cysteines to leave alone during mutagenesis. AstraPTM2 reaches F1 0.924 on N-linked glycosylation (high-recall) and F1 0.829 on disulfide bonds (high-precision), across 39 modification classes reported at two operating points. Phosphorylation is weaker as a thresholded exact-site task (high-precision F1 0.149) and should be treated as hypothesis-generating rather than a definitive regulatory map — a limit we state rather than smooth over.
Binding: a ranked panel, scored honestly
AstraBIND returns a ranked ligand panel and candidate contact residues from sequence alone — which ligand or cofactor hypotheses to review first for a co-crystal-scored protease. The honest framing matters here more than a single headline:
- On the 242 PDB-contact-scored proteases (2,396 ground-truth protein–ligand pairs), micro set-F1 is 0.367, and 99 of 242 reach set-F1 ≥ 0.5 — a higher-confidence subset worth panel review.
- Exact top-1 ligand-ID recall is 0.053 (Recall@3 0.137, Recall@5 0.200). This is a single best-guess three-letter PDB component matched against the full ligand vocabulary — deliberately the strictest test, reported as a governance check on panel identity, not the product promise. It reads low by construction.
- Residue-level pocket localization is the buyer signal: median ligand-residue F1 0.389 and median per-residue AUROC 0.660 across the scored subset — evidence of whether predicted contacts cluster around the observed pocket before choosing a co-crystal or competition experiment.
The way to use it: freeze the ranked panel at the start of a decision cycle, separate cofactors from inhibitor-like and context ligands, and use exact-ID recovery as governance — not as a drug-like chemotype claim.
One protease, end to end: neprilysin
Aggregate metrics are for governance; buying decisions happen one target at a time. The whitepaper walks the whole suite through a single, transparent case — neprilysin (MME, UniProt P08473), a 750-aa type II single-pass zinc metalloendopeptidase of the M13 gluzincin family, inhibited by sacubitrilat in sacubitril/valsartan and a degrader of amyloid-β:
- Route the sequence. The function head calls hydrolase (0.988) and protease GO (0.950), sending it to protease chemistry.
- Call the anchor state. Single-pass membrane (0.992) with a zinc cofactor (0.946) and a signal-anchor at residues 29–51 — so the ectodomain and membrane-retaining construct routes get designed deliberately.
- Flag the construct-design risks. N-glycosylation sites recovered at the annotated positions plus hypothesis-grade additions, and a high-precision disulfide map — informing expression host and mutation-avoidance before construct ordering.
- Generate the binding panel. 11 of 14 strict co-crystal ligand IDs recovered (set-F1 0.85), with the zinc cofactor ranked at the catalytic cleft — pocket-localization evidence a structural biologist can inspect before the first co-crystal.
Neprilysin is an upper-tail positive control, not the typical case: realistic per-target expectation sits at the catalytic-class medians (set-F1 0.29–0.50). But it shows what the integrated output looks like — a sequence-only decision package a program team can act on.
Why this matters for protease programs
Much of the human degradome stays hard for practical reasons, not biological disinterest: unstable constructs, ambiguous membrane context, and paralogue families that demand early counter-screening. From one sequence, Astra returns the evidence a protease team needs for a first decision review — family and chemistry, membrane and cofactor routing, residue-level topology, glycosylation and disulfide risk, and a ranked binding-panel hypothesis. That means fewer blind construct variants, fewer undirected co-crystal attempts, and a defensible first active-site panel before expensive wet-lab routing begins. This isn't a claim that the model replaces structural biology — it's a claim that a protease program can decide what structural biology to buy first.
Read the full benchmark
The Proteases volume reports every model's headline performance, its failure modes, the operating envelope, and a decision framework mapping common protease program questions to the model combinations that answer them.
→ Read the full Proteases benchmark (free): https://www.orbion.life/research/protease-performance
Part of the Orbion Model Performance Series — alongside GPCRs, Transporters, Enzymes, and Ion Channels.
References & sources
- Turner, A. J., Isaac, R. E. & Coates, D. The neprilysin (NEP) family of zinc metalloendopeptidases: genomics and function. BioEssays 23(3):261–269 (2001). PMID:11223883 — the M13 gluzincin protease family.
- McMurray, J. J. V. et al. Angiotensin-neprilysin inhibition versus enalapril in heart failure. New England Journal of Medicine 371:993–1004 (2014). doi:10.1056/NEJMoa1409077 — neprilysin as a validated therapeutic target.
- Iwata, N. et al. Metabolic regulation of brain Aβ by neprilysin. Science 292(5521):1550–1552 (2001). doi:10.1126/science.1059946 — neprilysin amyloid-β degradation.
- Rawlings, N. D. et al. The MEROPS database of proteolytic enzymes. Nucleic Acids Research 46(D1):D624–D632 (2018). doi:10.1093/nar/gkx1134 — protease family and catalytic-class classification.
- The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Research 51(D1):D523–D531 (2023). doi:10.1093/nar/gkac1052 — Swiss-Prot reference annotations.
- Burley, S. K. et al. RCSB Protein Data Bank. Nucleic Acids Research 47(D1):D464–D474 (2019). doi:10.1093/nar/gky1004 — co-crystal ligand-contact references.
- Davis, J. & Goadrich, M. The relationship between Precision-Recall and ROC curves. Proceedings of the 23rd ICML, 233–240 (2006). doi:10.1145/1143844.1143874 — AUROC/AUPRC interpretation.
- Full per-model methodology and metrics: Orbion, Astra AI on Proteases — Model Performance Series (2026) — https://www.orbion.life/research/protease-performance



