CORAL-AI V.01 · Molecular prioritization benchmark · October 2025
Combining 3D structure and sequence achieved EF₁%=38.2 in molecular screening
CORAL-AI V.01 combined structural and sequence representations to rank protein–ligand interactions. In the evaluated benchmark, it reached EF₁%=38.2 and AUC-ROC=0.823; the compared methods achieved EF₁% values between 15.5 and 33.3.
Quick facts
Protein–ligand prioritization benchmark
- Modality
- Virtual screening of protein–ligand interactions
- Data
- Complexes and decoys from the DUD-E benchmark
- Inputs
- 3D molecular structure and protein sequence
- Primary metric
- EF₁% · enrichment within the top-ranked 1%
- Comparators
- OnionNet-SFCT, GNINA, and GenScore
- Primary result
- CORAL-AI V.01 · EF₁%=38.2
- Secondary metric
- AUC-ROC=0.823
- Scope
- Computational benchmark; no prospective experimental validation
01 · Context
The value of screening lies in deciding what to test first
Molecular libraries may contain millions of compounds, while experimental capacity remains limited. Virtual screening narrows that space by ranking candidates before they reach the laboratory. Its value depends less on classifying an entire library correctly than on concentrating active compounds near the top of the list.
In aquaculture, where discovery programs often operate with less infrastructure and data than human pharmaceutical programs, a better initial ranking can reduce uninformative assays and focus resources on candidates with stronger computational support.
02 · Challenge
Combining complementary signals under severe class imbalance
Protein–ligand affinity depends on geometry, chemical environment, and receptor context. A model based only on 3D structure may miss evolutionary information from the protein; a sequence-only model cannot directly observe the spatial arrangement of atoms. The evaluated data also contained approximately 1.6% positive instances, so high overall accuracy could conceal poor ranking of the few active compounds.
The objective was to improve early recovery without treating a well-ranked prediction as therapeutic activity. EF₁% was therefore defined as the primary metric, while AUC-ROC remained a complementary measure of global performance.
03 · Approach
A hybrid architecture for structure, sequence, and uncertainty
CORAL-AI V.01 represented molecular complexes with graph neural networks and contextualized each receptor with a protein language model. A cross-attention layer integrated both sources before producing an interaction score. Class-balanced sampling prevented decoys from dominating training.
The architecture included uncertainty estimation through ensembles and Monte Carlo dropout. This information does not replace the primary prediction, but it can distinguish candidates with similar scores and different confidence before compounds are selected for later experimental work.
04 · Results
More active compounds concentrated within the top-ranked 1%
CORAL-AI V.01 achieved EF₁%=38.2. GenScore reached 33.3; GNINA, 18.8; and OnionNet-SFCT, 15.5. The relative difference was 14.7% against the closest comparator and 146.5% against the lowest value included in the review.
The model also achieved AUC-ROC=0.823. The two metrics answer different questions: AUC-ROC summarizes separation across the full dataset, whereas EF₁% measures the concentration of active compounds where experimental selection typically begins.
| Method | EF₁% | Difference relative to CORAL-AI V.01 |
|---|---|---|
| CORAL-AI V.01 | 38.2 | Reference |
| GenScore | 33.3 | CORAL-AI V.01 +14.7% |
| GNINA | 18.8 | CORAL-AI V.01 +103.2% |
| OnionNet-SFCT | 15.5 | CORAL-AI V.01 +146.5% |
These results compare prioritization performance within this benchmark. They do not measure experimental affinity, therapeutic efficacy, or universal performance across arbitrary libraries.
Comparative benchmark
Enrichment factor within the top 1%
05 · What it means
A more concentrated ranking reduces the experimental front
Within this evaluation, combining structure and sequence recovered a larger proportion of active compounds near the top of the ranking. For a discovery program, this means a smaller fraction of the library can advance to physical assays before relevant signals are found.
The result validates the architecture's prioritization capacity in the studied benchmark. It does not demonstrate efficacy in animals or replace experimental confirmation. Application in aquaculture requires prospective evaluation for each target, chemical domain, and library under a protocol with defined controls.
06 · Methodological note
Scope of the evaluation
The review used DUD-E as its set of active complexes and decoys and compared EF₁% with OnionNet-SFCT, GNINA, and GenScore. CORAL-AI V.01 integrated molecular graphs, sequence representations, cross-attention, class balancing, and uncertainty estimation. EF₁% measured early recovery and AUC-ROC measured global discrimination. No prospective assays, in vivo validation, campaign costs, or regulatory outcomes were included.
References: Mysinger et al., Journal of Medicinal Chemistry 55, 6582–6594 (2012); Kim et al., Chemical Science 14, 3245–3258 (2023); Lin et al., Science 379, 1123–1130 (2023).
Do you have a molecular library to prioritize?
Define the target, chemical domain, and validation protocol before starting a screening campaign.
