What is it about?
The Protein Data Bank is an archive of three-dimensional structures of macromolecules (e.g., proteins and nucleic acids) which has been growing at an increasing rate over more than 50 years. In order to accommodate this tremendous growth as well as improvements to data formats and standards, it has become necessary to extend the identifiers of structures from 4 to 12 characters and adjust how data is organized in the archive. This article presents this updated version of the archive, which has been staged as a "Beta Archive," to prepare researchers for the planned transition occurring in mid-2027.
Featured Image
Photo by National Institute of Allergy and Infectious Diseases on Unsplash
Why is it important?
The current structure of the PDB archive and data identifiers will be officially changing on or around July 21st, 2027, which will impact any researchers currently using PDB data in their work. The unveiling of this public "Beta Archive" provides researchers with an early window into the organization and file-naming of the new archive ahead of the official transition date, so that they can adapt any existing software and infrastructure as needed.
Read the Original
This page is a summary of: Protein Data Bank (PDB) Archive: a new architecture (beta) for scalable, PDBx/mmCIF-based data distribution, Acta Crystallographica Section D Structural Biology, July 2026, International Union of Crystallography,
DOI: 10.1107/s2059798326006194.
You can read the full text:
Contributors
The following have contributed to this page







