This week, we’re sharing a new paper, Overcoming the accuracy–generalization tradeoff in docking and scoring for prospective virtual screening, where we present DODock and DOScore — our proprietary docking and scoring software framework for novel drug targets.
Structure-based virtual screening (SBVS) has long been a promise for discovery success with access to billions of synthetically accessible compounds, searchable directly against a target’s 3D structure. Yet, in practice, it’s only used as a primary hit-discovery strategy in fewer than 1% of drug discovery campaigns. This is the problem we set out to close.
Why traditional screening methods have underdelivered and why we’ve set out to rebuild it
Classical docking methods are generalizable, but face a major impasse. Their additive, pairwise, scoring functions can’t capture the many-body, conformer-dependent physics of real molecular recognition, resulting in accuracy levels of 40-55% self-docking success.
Recent diffusion-based and cofolding models (AlphaFold3, Boltz-1, Chai-1, and others) have attempted to overcome the ceiling by learning flexibility directly from data. And on headline benchmarks, they do look better. The catch: much of that performance doesn’t apply to truly novel targets. Under strict protein and ligand splits, rather than time-based splits commonly found in literature, success rates for these models fall steeply as similarity to the training set decreases. In some cases the drop is drastic, from roughly 80% success on familiar targets to 25% on novel ones. These models are great at fitting their training distribution beautifully but they’re doing so by recognizing, not generalizing.

Our approach: learn only where it helps
DODock combines a diffusion model with a physics-based search and refinement stage built on an expanded, interpretable Vina-style energy function (DOFast) and an enhanced global optimizer. The goal is to put learned components exactly where they add value (proposing and ranking poses) and keep an interpretable, low-parameter physical model wherever generalization is what matters. Once machine learning proposes candidate poses, a compact physics model refines them, reasoning from how molecules physically interact with each other rather than from resemblance to training data. These refined poses are then ranked, giving our resulting structures.
The middle step is what sets our approach apart: because the physics reasons from first principles rather than prior data, it holds up on novel targets, which is exactly where AI-only approaches lose accuracy.
DOScore follows the same principle for affinity prediction. Structurally-labeled, affinity-annotated complexes are rare so we used DODock to generate large-scale structural training data from assay data, turning docking accuracy into scoring supervision.
How DODock and DOScore outperform on docking and scoring
We tested our models under strict simultaneous protein- and ligand-similarity splits, removing every protein and ligand that resembles the training set. This is the true indication of how a screen will do on a target it has never seen, and is most relevant for actual drug discovery. We outperform other models under the following benchmarks:
- CASF-2016 (redocking): DODock hits 81.1% success at 2 Å and 58.6% at 1 Å— more than 25 points above widely used commercial and academic docking programs.
- Runs N’ Poses: while cofolding methods drop from ~75–88% success on similar complexes down to ≤25% on dissimilar ones, DODock stays flat, exceeding 50% even in the least-similar bin.
- OpenBind (EV-A71 2A protease, a genuinely novel target): DODock reaches 80% top-1 success, versus 4–28% for current cofolding models and 2% for DiffDock.
- PoseBusters: consistently strong performance across kinases, oxidoreductases, hydrolases, transporters, and nucleic-acid-binding proteins alike. Results showed a mean of 82% at 2 Å, not concentrated in a handful of data-rich classes.
- DUD-E and DEKOIS 2.0: DOScore enriches actives strongly across the large majority of targets.
Our docking predicts the correct binding pose about 81% of the time on standard redocking (CASF-2016), more than 25 points better than classical docking models, whose additive scoring caps out near 55%.
Our scoring reliably ranks true drug candidates above look-alikes, even on targets unlike the model’s training data.
We tested it in the lab, not only on benchmarks
Benchmarks are only half the story. Most claims in the field rest solely on benchmarks which can flatter a model by testing it on molecules too close to the ones it was trained on. We wanted to show that our model overcomes this challenge and is able to predict unfamiliar targets so we pushed DODock and DOScore into real prospective campaigns across historically difficult targets:
- Prospective screens across four targets spanning compact enzyme pockets (CD73, IRAK4), an extended-substrate protease (Factor XIa), and an allosterically addressed protein–protein interface (IL-17A) all yielded chemically novel, biochemically and cellularly active hits. For CD73, a conformationally dynamic target that has resisted non-nucleotide chemotypes for years, we yielded a ~30% hit rate (54 of 183 compounds).
- A blind pose prediction for PCSK9 bound to AstraZeneca’s AZD0780, made before any experimental structure existed, matched the subsequently determined crystal structure to about 1.2 Å RMSD at an atypical binding site outside the well-characterized catalytic and LDLR-binding surfaces.
Why we’re publishing all of it
We’re publishing all of it – the full method and an extensive supplement – open for anyone to reproduce and to test on targets we haven’t seen yet. Science moves forward when results can be checked, and when wrong, proven wrong. We believe that the marginal role of SBVS in drug discovery reflects a tractable accuracy problem, not a fundamental limit. What should draw our attention with such models isn’t peak benchmarking performance, rather that accuracy tends to degrade with distance from the training set. This is the key property that actually matters once you direct a model towards a target it’s never seen before, which is, after all, the entire point of prospective drug discovery. Our paper dives into how our team at Deep Origin is closing the gap to bring the promise of vast virtual chemical space into reality – take a look and let us know what you think.
Read the research
- The paper: Overcoming the accuracy–generalization tradeoff in docking and scoring for prospective virtual screening
- The preprint on bioRxiv: supplemental materials are there — more than 80 pages covering every algorithm, hyper-parameter and data split.
- Live Q&A: Useful Generalization is the Benchmark That Matters — join us on September 2 to hear the results presented and put questions to the team.
