Research

Model Performance Series

June 2026

Astra AI on Proteases

Proteases are one of medicine's most clinically validated enzyme classes — and among the hardest to move from sequence to a first experiment: unstable constructs, ambiguous membrane anchors, catalytic metal sites, and paralogue selectivity risk. This report shows how the Astra suite characterizes them from sequence alone — function and family, membrane and cofactor routing, residue-level topology, PTM sites, and a ranked binding panel — against public experimental ground truth across 545 reviewed human proteases.

Çağlar Bozkurt, Aniruddh Goteti · Orbion GmbH · Benchmarked on 545 Human Proteases

0.951
AUROC · Residue-Level Topology
91.6%
Recognized as Hydrolases
F1 0.924
N-Linked Glycosylation Sites

What We Found

Three Results That Matter

Each section of the whitepaper reports the headline performance for one prediction area, its known weaknesses, and where we recommend it for production use.

01

0.951 AUROC on Residue-Level Topology

Most proteases are soluble — but a membrane-anchored protease expressed as a naive soluble construct loses the context the program cares about. On the annotated membrane-anchored fraction, per-residue transmembrane prediction agrees with UniProt segments at AUROC 0.951 (median 0.995), and a zinc cofactor is called on 124 targets — giving construct designers an anchor state, an ectodomain boundary, and the catalytic-metal context before the first experiment.

02

F1 Up to 0.924 on Construct-Design PTMs

For secreted and membrane proteases, the modifications that decide expression host and domain boundaries are the structural ones. Across all 39 PTM classes the suite flags per-residue sites at two operating points — strongest on N-linked glycosylation (F1 0.924) and disulfide bonds (F1 0.829), the marks that govern protease folding and surface delivery.

03

A Ranked Binding Panel, Scored Honestly

From sequence alone the suite returns a ranked ligand panel and candidate contact residues — which cofactor and inhibitor-like hypotheses to review first at the catalytic cleft. On the co-crystal-scored proteases, 99 of 242 reach a review-worthy panel and predicted contacts localize around the observed pocket (median residue AUROC 0.660). Exact top-1 ligand-ID recovery is low by construction (0.053), and we report it plainly as the strictest test.

Every number above is reproduced in the whitepaper against public references — Swiss-Prot annotation, PDB co-crystal contacts, and curated experimental thermal-shift data. We report where the models are weak as plainly as where they are strong.

Read the Whitepaper