Machine Learning–Ready Dataset

Curated Repository Of Well-resolved Non-covalent Interactions

C · R · O · W · N

141,261 protein–ligand complexes, built directly from PDB crystal structures. Quality-filtered, structurally repaired, protonated, and energy-minimized for structure-based ML.

141k
Curated complexes
63,146
Unique PDB IDs
14,184
Unique proteins
16,028
Murcko scaffolds
Motivation

Existing databases force a trade-off between quality and scale

Machine learning models for protein–ligand interactions need both structural reliability and broad chemical coverage. Current resources offer one or the other — never both.

01

Curated but narrow

Databases like PDBBind and HiQBind offer carefully checked structures but cover only a fraction of the PDB. Their scale limits protein, ligand, and scaffold diversity, making generalization harder for data-hungry structure-based models.

02

Broad but noisy

Large-scale resources such as PLINDER catalogue nearly 650,000 systems, but retain many deposited-structure artifacts. Unresolved atoms, steric clashes, incorrect connectivity, and incomplete quality annotations can become systematic noise in training data.

03

Preprocessed but inherited

Resources such as SPINDR move toward automated preprocessing, but still build on upstream annotations. CROWN starts directly from deposited mmCIF structures, allowing each complex to be independently defined and validated against experimental electron density.

solution

CROWN reconciles scale and rigor

A fully automated pipeline combines broad PDB coverage with stringent quality control, structural repair, physiological protonation, and constrained energy minimization — producing a uniform, geometry-centric dataset without requiring affinity labels.

Pipeline

From 189,265 crystal structures to 141,261 curated complexes

CROWN applies five sequential curation, quality-control, and optimization steps, targeting crystallographic resolution, ligand identity, pocket completeness, interaction quality, structural repair, protonation, and post-minimization stability.

CROWN preprocessing pipeline attrition funnel diagram

Preprocessing pipeline for CROWN, illustrated as a staged attrition funnel. Starting from 189,265 PDB X-ray crystal structures, the pipeline applies curation, quality-control, and optimization steps to define 225,627 PLI systems and retain 141,261 final systems with protonated and energy-minimized receptor and ligand files.

Key Features

What distinguishes CROWN

Every complex has been quality-filtered, structurally repaired, protonated at physiological pH, energy-minimized with custom restraints, and annotated for leakage-controlled machine-learning splits.

QUALITY

Complete density validation

Every entry has valid electron-density quality metrics, with mean RSR below 0.3 and RSCC above 0.8 for both ligand and pocket residues. Entries without verifiable density support are excluded.

QUALITY

Chemically focused systems

CROWN removes ions, crystallization artifacts, covalent ligands, and unsupported chemistries while retaining binding sites with organic and selected inorganic cofactors. Ligands must be drug-like and meaningfully buried in the pocket.

PROCESSING

Automated structural repair

Alternate conformers are collapsed, missing connectivity is repaired, symmetry mates are recovered, and defects outside the pocket are rebuilt or standardized while pocket defects are filtered out.

PROCESSING

Constrained energy minimization

A custom flat-bottomed tethering potential lets pocket atoms relax within crystallographic uncertainty (0.25 Å) while preserving experimental geometry. This reduces clashes and strained poses without erasing the crystal signal.

DESIGN

Leakage-aware metadata

Every entry includes cluster labels for protein sequence similarity, ligand ECFP4 similarity, and PLEC interaction similarity at 50%, 70%, and 90% cutoffs, supporting stricter train-test splits.

DESIGN

Broad chemical coverage

CROWN contains 26,576 unique CCD ligands and 16,028 Murcko scaffolds, preserving broad ligand-property distributions including larger contemporary modalities such as PROTACs, macrocycles, and molecular glues.

Comparison

CROWN in the landscape of protein–ligand databases

A side-by-side comparison of dataset scope, ligand diversity, structural issues, and correction steps across five widely used resources.

PropertyPDBBindHiQBindBioLiP2PLINDERCROWN
Dataset scope
Total entries19,44931,57386,458649,915141,261
Unique PDB-CCD pairs17,75817,24722,720201,83675,549
Unique PDB IDs17,75817,08817,701111,86763,146
Unique UniProt IDs3,3542,64214,93322,24314,184
Unique CATH IDs5834431,0171,5652,041
Unique species8617153,4664,8821,473
Affinity annotations
Ligand diversity
Unique CCD IDs13,95612,4286,51947,30026,576
Unique Murcko scaffolds8,2517,6603,21322,74316,028
Ion ligands001,80522,7280
Covalent ligands870241,70832,2760
Artifact ligands153452318,6260
Structure issues
Missing bonds9901142,10437,4220
Steric overlaps3233136565,6700
Unresolved ligand atoms46311,56418,8150
Unresolved pocket atoms1,1029551,57912,4570
Non-standard pocket residues3072459724,4730
Structure corrections
Protonation
Energy minimization

Comparison of dataset scope, ligand diversity, and ligand quality filtering across five structural protein–ligand interaction datasets. CROWN uniquely combines complete removal of listed structure issues with both protonation and energy minimization. Values for CROWN are reported for the corrected structures.

Access

Freely available under CC BY 4.0

Browse, search, compare crystal and minimized coordinates, and download individual entries or the complete dataset. The full preprocessing pipeline is open-source.