Curated Repository Of Well-resolved Non-covalent Interactions
C · R · O · W · N
141,261 protein–ligand complexes, built directly from PDB crystal structures. Quality-filtered, structurally repaired, protonated, and energy-minimized for structure-based ML.
Existing databases force a trade-off between quality and scale
Machine learning models for protein–ligand interactions need both structural reliability and broad chemical coverage. Current resources offer one or the other — never both.
Curated but narrow
Databases like PDBBind and HiQBind offer carefully checked structures but cover only a fraction of the PDB. Their scale limits protein, ligand, and scaffold diversity, making generalization harder for data-hungry structure-based models.
Broad but noisy
Large-scale resources such as PLINDER catalogue nearly 650,000 systems, but retain many deposited-structure artifacts. Unresolved atoms, steric clashes, incorrect connectivity, and incomplete quality annotations can become systematic noise in training data.
Preprocessed but inherited
Resources such as SPINDR move toward automated preprocessing, but still build on upstream annotations. CROWN starts directly from deposited mmCIF structures, allowing each complex to be independently defined and validated against experimental electron density.
CROWN reconciles scale and rigor
A fully automated pipeline combines broad PDB coverage with stringent quality control, structural repair, physiological protonation, and constrained energy minimization — producing a uniform, geometry-centric dataset without requiring affinity labels.
From 189,265 crystal structures to 141,261 curated complexes
CROWN applies five sequential curation, quality-control, and optimization steps, targeting crystallographic resolution, ligand identity, pocket completeness, interaction quality, structural repair, protonation, and post-minimization stability.

Preprocessing pipeline for CROWN, illustrated as a staged attrition funnel. Starting from 189,265 PDB X-ray crystal structures, the pipeline applies curation, quality-control, and optimization steps to define 225,627 PLI systems and retain 141,261 final systems with protonated and energy-minimized receptor and ligand files.
What distinguishes CROWN
Every complex has been quality-filtered, structurally repaired, protonated at physiological pH, energy-minimized with custom restraints, and annotated for leakage-controlled machine-learning splits.
Complete density validation
Every entry has valid electron-density quality metrics, with mean RSR below 0.3 and RSCC above 0.8 for both ligand and pocket residues. Entries without verifiable density support are excluded.
Chemically focused systems
CROWN removes ions, crystallization artifacts, covalent ligands, and unsupported chemistries while retaining binding sites with organic and selected inorganic cofactors. Ligands must be drug-like and meaningfully buried in the pocket.
Automated structural repair
Alternate conformers are collapsed, missing connectivity is repaired, symmetry mates are recovered, and defects outside the pocket are rebuilt or standardized while pocket defects are filtered out.
Constrained energy minimization
A custom flat-bottomed tethering potential lets pocket atoms relax within crystallographic uncertainty (0.25 Å) while preserving experimental geometry. This reduces clashes and strained poses without erasing the crystal signal.
Leakage-aware metadata
Every entry includes cluster labels for protein sequence similarity, ligand ECFP4 similarity, and PLEC interaction similarity at 50%, 70%, and 90% cutoffs, supporting stricter train-test splits.
Broad chemical coverage
CROWN contains 26,576 unique CCD ligands and 16,028 Murcko scaffolds, preserving broad ligand-property distributions including larger contemporary modalities such as PROTACs, macrocycles, and molecular glues.
CROWN in the landscape of protein–ligand databases
A side-by-side comparison of dataset scope, ligand diversity, structural issues, and correction steps across five widely used resources.
| Property | PDBBind | HiQBind | BioLiP2 | PLINDER | CROWN |
|---|---|---|---|---|---|
| Dataset scope | |||||
| Total entries | 19,449 | 31,573 | 86,458 | 649,915 | 141,261 |
| Unique PDB-CCD pairs | 17,758 | 17,247 | 22,720 | 201,836 | 75,549 |
| Unique PDB IDs | 17,758 | 17,088 | 17,701 | 111,867 | 63,146 |
| Unique UniProt IDs | 3,354 | 2,642 | 14,933 | 22,243 | 14,184 |
| Unique CATH IDs | 583 | 443 | 1,017 | 1,565 | 2,041 |
| Unique species | 861 | 715 | 3,466 | 4,882 | 1,473 |
| Affinity annotations | ✓ | ✓ | ✗ | ✗ | ✗ |
| Ligand diversity | |||||
| Unique CCD IDs | 13,956 | 12,428 | 6,519 | 47,300 | 26,576 |
| Unique Murcko scaffolds | 8,251 | 7,660 | 3,213 | 22,743 | 16,028 |
| Ion ligands | 0 | 0 | 1,805 | 22,728 | 0 |
| Covalent ligands | 870 | 24 | 1,708 | 32,276 | 0 |
| Artifact ligands | 15 | 34 | 523 | 18,626 | 0 |
| Structure issues | |||||
| Missing bonds | 990 | 114 | 2,104 | 37,422 | 0 |
| Steric overlaps | 323 | 313 | 656 | 5,670 | 0 |
| Unresolved ligand atoms | 463 | 1 | 1,564 | 18,815 | 0 |
| Unresolved pocket atoms | 1,102 | 955 | 1,579 | 12,457 | 0 |
| Non-standard pocket residues | 307 | 245 | 972 | 4,473 | 0 |
| Structure corrections | |||||
| Protonation | ✓ | ✓ | ✗ | ✗ | ✓ |
| Energy minimization | ✗ | ✗ | ✗ | ✗ | ✓ |
Comparison of dataset scope, ligand diversity, and ligand quality filtering across five structural protein–ligand interaction datasets. CROWN uniquely combines complete removal of listed structure issues with both protonation and energy minimization. Values for CROWN are reported for the corrected structures.
Freely available under CC BY 4.0
Browse, search, compare crystal and minimized coordinates, and download individual entries or the complete dataset. The full preprocessing pipeline is open-source.
