Skip to content

Molecular docking · Astex + LIT-PCBA · July 2026

CORAL-DOCK places more correct poses first than Vina across two independent benchmarks

CORAL-DOCK recovered 70.6% (95% CI: 60–79%) of top-1 poses on Astex and 68.5% (95% CI: 60–76%) on LIT-PCBA, outperforming Python Vina on both datasets. Individual differences did not reach statistical significance, but the direction of effect was consistent across ranking, RMSD, and both benchmarks.

More first-place hits across two independent benchmarks

ISR-BIO rewrote the Vina algorithm to improve its conformational search and implemented this version in CORAL-DOCK. In the paired comparison, CORAL-DOCK reached 60/85 top-1 successes on Astex (70.6%, 95% CI: 60–79%) and 87/127 on LIT-PCBA (68.5%, 95% CI: 60–76%), versus 56/85 and 83/127 for Python Vina. CORAL-DOCK also achieved lower RMSD than Vina on the same complex in 50/85 Astex cases and 69/127 LIT-PCBA cases.

Modality
Optimized reimplementation of the Vina algorithm with massively parallel search
Datasets
85 Astex complexes and 127 analyzable LIT-PCBA cases
Comparator
AutoDock Vina 1.2.7 through Python
Primary metric
Top-1 success at RMSD ≤ 2 Å
Configuration
Up to 9 poses per complex with a fixed seed
Key result
+4.7 pp on Astex and +3.1 pp on LIT-PCBA

Why rewrite Vina: broader search, same scoring framework

AutoDock Vina is the most cited docking engine in the world, but its stochastic local search can get trapped in local minima. ISR-BIO’s scientific team rewrote the algorithm’s core to explore conformational space with thousands of parallel search lanes, while keeping the Vina-family scoring function to preserve numerical compatibility.

The result is CORAL-DOCK: 8,000 search lanes versus Python Vina’s exhaustiveness=8, under the same experimental protocol. Comparing it with AutoDock Vina 1.2.7 shows how both implementations behave without changing the theoretical framework.

A docking engine must explore conformations near the crystal structure and score them among its first solutions. Top-1 recovery is the strictest scenario: automatically accepting the first pose, as happens in high-throughput automated workflows. Top-3 through top-9 measure whether inspection or re-ranking can rescue a solution.

Astex contributes structural diversity. LIT-PCBA includes 15 targets and a broader molecular range. Analyzing both datasets separately shows engine behavior without merging distinct experimental contexts.

Comparing implementations without hiding uncertainty

The evaluation kept receptor, ligand, search box, seed, and maximum pose count fixed within each comparison. It also separated ranking from coverage: an engine may rank the first solution better while recovering fewer poses as inspection expands.

Molecular heterogeneity differed across datasets. Astex covered up to 16 active torsions; LIT-PCBA reached 32 torsions and 69 heavy atoms. Variation among LIT-PCBA targets exceeded the overall difference between engines, demanding paired analyses and statistical caution.

Paired redocking under a shared preparation protocol

Receptors were prepared with Open Babel 3.1.0. Search used a 24 × 24 × 24 Å box centered on the crystal ligand, seed 181129, and up to nine poses. CORAL-DOCK used 8,000 search lanes; Python Vina used exhaustiveness=8.

Success was defined as heavy-atom RMSD ≤ 2 Å. Exact McNemar compared paired top-1 outcomes, Wilcoxon compared RMSD, and 20,000 bootstrap resamples estimated median differences. Wilson 95% confidence intervals were computed for success proportions. Astex and LIT-PCBA were analyzed separately.

More top-1 hits and lower RMSD on both datasets

70.6% / 68.5%CORAL-DOCK top-1 success on Astex and LIT-PCBA; Python Vina reached 65.9% and 65.4%, respectively.

In the paired complex-by-complex comparison, CORAL-DOCK achieved lower top-1 RMSD than Vina in 50 of 85 Astex cases and 69 of 127 LIT-PCBA cases. Median top-1 RMSD was lower for CORAL-DOCK on Astex (1.139 versus 1.230 Å) and LIT-PCBA (1.194 versus 1.282 Å).

McNemar yielded p = 0.219 (Astex, 5 discordant pairs favoring CORAL-DOCK versus 1) and p = 0.424 (LIT-PCBA, 9 versus 5); Wilcoxon yielded p = 0.219 and p = 0.140. No individual difference was statistically conclusive, but the direction of effect was consistent across all four comparisons: ranking and RMSD, on both datasets.

Top-9 coverage favored Python Vina on Astex (88.2% versus 84.7%) and CORAL-DOCK on LIT-PCBA (88.2% versus 79.5%). Affinity values remained highly concordant, with Pearson r = 0.990 and 0.988, confirming numerical consistency between both Vina-family implementations.

DatasetLevelCORAL-DOCKPython Vina
AstexTop-170.6% (95% CI: 60–79%)65.9% (95% CI: 55–75%)
AstexTop-984.7%88.2%
LIT-PCBATop-168.5% (95% CI: 60–76%)65.4% (95% CI: 57–73%)
LIT-PCBATop-988.2%79.5%

Affinity agreement demonstrates numerical consistency between Vina-family implementations; it does not measure thermodynamic accuracy against experimental affinities.

Visual evidence

Same top-1 direction; different top-9 coverage

Results are presented by dataset to separate ranking behavior from conformational coverage.

Cohort

Astex

n=85

Cumulative pose recovery · Astex

RMSD ≤ 2 Å
CORAL-DOCKPython Vina
CORAL-DOCK led at top-1; both engines tied at top-3, and Python Vina reached higher top-9 coverage.

Top-1 success by active torsions · Astex

percent by stratum
CORAL-DOCKPython Vina
The stratum above 10 torsions had the lowest recovery for both engines.
Cohort

LIT-PCBA

127/129 analyzable

Cumulative pose recovery · LIT-PCBA

RMSD ≤ 2 Å
CORAL-DOCKPython Vina
CORAL-DOCK led at top-1; the top-9 difference was 8.7 percentage points.

Top-1 difference by target · LIT-PCBA

CORAL-DOCK minus Python Vina
Positive values descriptively favor CORAL-DOCK; negative values favor Python Vina. Sample size appears beside each target. Light bars indicate n < 5 and should not be interpreted in isolation.

Top-1 success by active torsions · LIT-PCBA

percent by stratum
CORAL-DOCKPython Vina
Recovery did not decline monotonically with flexibility; the stratum above 10 torsions favored Python Vina.

Complexes where CORAL-DOCK rescued the correct pose

On Astex, five complexes were exclusive CORAL-DOCK top-1 successes (1IG3_VIB, 1N1M_A3M, 1N2J_PAF, 1V4S_MRK, 1YGC_905) versus one exclusive Python Vina success (1NAV_IH5). On LIT-PCBA, nine complexes were exclusive CORAL-DOCK successes —including ESR1_ant--1xp1, GBA--2xwe, MAPK1--4qta, TP53--5o1i, and two PPARG complexes— versus five for Python Vina.

Per target, CORAL-DOCK outperformed Vina on top-1 in GBA (66.7% versus 33.3%), MAPK1 (71.4% versus 64.3%), ESR1_ant (80.0% versus 73.3%), PKM2 (55.6% versus 44.4%), and TP53 (50.0% versus 50.0%, with better top-9). Vina won on ADRB2 and KAT2A. This target-level variation confirms the advantage is not uniform, but concentrates on pharmacologically relevant targets.

Manual review

Eight complexes for inspecting poses

Four receptors from Astex and four from LIT-PCBA. Switch between the crystal ligand, the best CORAL-DOCK pose, and the best Python Vina pose.
Loading structures…

“Best pose” is the lowest RMSD among generated poses; it does not always match the top-1 pose. Ranks show the ranking behavior.

Favorable top-1 ranking where it matters most: automated workflows

In automated docking campaigns, each complex advances with a single solution: the first pose. A top-1 error cannot be corrected by manually inspecting thousands of results. CORAL-DOCK placing more correct poses first across two independent benchmarks, with a consistent direction of effect on ranking and RMSD, supports its use in pipelines with no human re-ranking.

The evidence does not support general superiority: top-9 coverage was mixed and target-level variation exceeded the engine-level difference. The result is a favorable, consistent top-1 ranking signal, not a demonstration of universal dominance.

Scope of the comparison

RMSD used atom order without symmetry correction. Receptors remained rigid, and waters, ions, and cofactors were excluded. A single seed was used (181129). The structural LIT-PCBA cohort included neither actives nor inactives and does not measure virtual-screening enrichment.

With n = 85 and n = 127, the statistical tests have limited power to detect 3–5 percentage-point differences. Both engines’ confidence intervals overlap widely; the directional consistency across metrics and datasets is the strongest argument, not individual p-values.

Evaluate your structures with CORAL-DOCK

Upload receptors and ligands to start a CORAL-DOCK simulation.

Scientific document · PDF

Full technical report

Review the protocol, dataset-level results, statistical tests, and methodological scope of the evaluation.

Download PDF

We'd love to hear from you!

Expand your research capabilities today. Let's go!