Data Science exposed thousands of exact duplicates in high-profile structural databases
Presenter
July 8, 2026
Abstract
The talk discusses several databases of solid crystalline materials and protein structures, such as Google's GNoME AI and the Protein Data Bank (PDB), where every entry is given by dozens or hundreds of atomic positions in Euclidean or fractional coordinates with respect to a unit cell. Unfortunately, all such cells discontinuously change under almost any noise. This discontinuity of cell-based representations has been resolved by a hierarchy of geometric invariants starting with the ultra-fast vectorial and matrix descriptors, which distinguish all non-duplicate structures in the world’s largest databases of experimental materials, by using only atomic centers without chemical elements, and finishing with the complete isoset invariant and Lipschitz continuous metrics that separate all periodic sets of points under isometry. Though independent experiments and non-trivial simulations are highly unlikely to produce identical numerical outputs, the Geometric Data Science approach in [1-3] revealed thousands of exact duplicates with all x,y,z coordinates identical. Detecting exact and near-duplicates is essential to avoid biases in AI and ML models that use this data as an input.
[1] O. Anosova, V. Kurlin, M. Senechal. The importance of definitions in crystallography. IUCrJ, v.11(4), p.453-463 (2024).
[2] O. Anosova, A. Gorelov, W. Jeffcott, Z. Jiang, V. Kurlin. Complete and bi-continuous invariant of protein backbones under rigid motion. MATCH, v.94(1), p.97-134 (2025).
[3] A. Wlodawer, Z. Dauter, P. Rubach, W. Minor, M. Jaskolski, Z. Jiang, W. Jeffcott, O. Anosova, V. Kurlin. Duplicate entries in the Protein Data Bank: how to detect and handle them. Acta Cryst D, v.81, 170-180 (2025).