Large-scale analysis of AlphaFold structures reveals organism-specific physicochemical signatures
Large-scale AlphaFold structure analysis reveals organism-specific physicochemical signatures reconstructed via DE-STRESS metrics across 48 proteomes and PDB structures.
Published Oct 4, 2026Paper ↗
Only vote on papers you've read. Sign in with GitHub to vote.
The dataset delivers massive AlphaFold and PDB feature dumps plus DE-STRESS tree comparisons, yet the Zenodo record omits critical baselines, raw AFDB clusters, and rigorous null checks, leaving its organism-specific physicochemical signatures vulnerable to sequence redundancy…
Abstract
Data for: Large-scale analysis of AlphaFold structures reveals organism-specific physicochemical signatures This record contains the input data needed to reproduce the analysis in the paper above. The analysis code is at https://github.com/wells-wood-research/proteinPropertiesReconstructTreeOfLife, and the repository README explains how to use these files. To set up the data, run `bash scripts/setup_data.sh RECORD_ID` from the repository root. It downloads these files and places each one where the scripts expect it. Files Tree of Life analysis destress_data.zip - DE-STRESS metrics [1] (about 70 structural features per model) for 564,446 AlphaFold2 [2] models of the proteomes of 48 organisms (AlphaFold DB v4) [3] and for 216,681 experimentally determined PDB structures [4] (downloaded June 2024). Contains `destress_data_af2.csv` and `destress_data_pdb_082024.csv`.af2_plddt_scores.csv - Mean pLDDT for each AlphaFold2 model, calculated from the C-alpha atom B-factor values in the AlphaFold DB PDB files. Models with mean pLDDT < 70 are removed in the analysis.filtered_af2db_clusters_data.csv - AlphaFold DB FoldSeek cluster assignments (cluster representative, UniProt ID, cluster flag, taxon ID) for the proteins in the data set. Filtered from the AFDB cluster membership table [5].uniprot_results_org_subcellloc_uniprotkb.csv - Organism, subcellular location, GO terms and gene encoding type for each protein, from the UniProt REST API (snapshot downloaded November 2024) [6].af2db_org_uniprot_desc_results.csv - Organism name and UniProt description for each protein, from the AlphaFold DB REST API (snapshot downloaded October 2024) [3].af2_structures_non_redundant.csv - The non-redundant list of 251,376 structures (one randomly selected structure per organism and FoldSeek cluster representative, random seed 42) [5].ncbi_phylo_tree.phy and ncbi_phylo_tree_euk.phy - Reference phylogenies in Newick format for all 48 organisms and for the 31 eukaryotes, created with the Common Tree tool of the NCBI Taxonomy Browser [7].randomTreeDistances.rda - Pre-computed distances between random pairs of bifurcating trees, used as the null model for the Clustering Information Distance (CID). This is the `randomTreeDistances` object from the R data package TreeDistData (Smith, https://github.com/ms609/TreeDistData). It is an array of 24 tree-distance methods x 13 summary statistics (min, percentiles, max, mean, sd) x 197 leaf counts (4 to 200); the analysis uses the "cid" entries. It is used with the R package TreeDist (10.32614/CRAN.package.TreeDist) [8], which was also used to calculate the CID between the DE-STRESS trees and the NCBI trees. The file is included for convenience and is not our work; please cite TreeDist/TreeDistData if you use it. scFv antibody analysis fleishman_antibody_expression_data.csv - Yeast display protein production measurements for the designed scFvs [9].fleishman_destress_data.csv - DE-STRESS metrics [1] for AlphaFold2 models of the designed scFvs (five models per design).pdb_destress_data.csv - DE-STRESS metrics for 35 experimentally determined scFv structures from the Structural Antibody Database (SAbDab) [10, 11].scfv_pdb_list.txt - PDB IDs of the SAbDab scFv structures [10, 11]. Not included The AlphaFold DB and PDB structure files, the AlphaFold2 models of the designed scFvs, and the full AFDB cluster membership table are not included. The first two are public (AlphaFold DB: https://alphafold.ebi.ac.uk/download; PDB: https://www.rcsb.org), and the cluster table is available from https://afdb-cluster.steineggerlab.workers.dev/. The list of AlphaFold DB proteome archives used is in the GitHub repository (`src/data_download/afdb_proteome_index.html`). References [1] Stam, Michael J., and Christopher W. Wood. 2021. ‘DE-STRESS: A User-Friendly Web Application for the Evaluation of Protein Designs’. Protein Engineering, Design and Selection 34 (February): gzab029. https://doi.org/10.1093/protein/gzab029. [2] Jumper, John, Richard Evans, Alexander Pritzel, et al. 2021. ‘Highly Accurate Protein Structure Prediction with AlphaFold’. Nature 596 (7873): 7873. https://doi.org/10.1038/s41586-021-03819-2. [3] Varadi, Mihaly, Stephen Anyango, Mandar Deshpande, et al. 2022. ‘AlphaFold Protein Structure Database: Massively Expanding the Structural Coverage of Protein-Sequence Space with High-Accuracy Models’. Nucleic Acids Research 50 (D1): D439–44. https://doi.org/10.1093/nar/gkab1061. [4] Berman, Helen M., John Westbrook, Zukang Feng, et al. 2000. ‘The Protein Data Bank’. Nucleic Acids Research 28 (1): 235–42. https://doi.org/10.1093/nar/28.1.235. [5] Barrio-Hernandez, Inigo, Jingi Yeo, Jürgen Jänes, et al. 2023. ‘Clustering Predicted Structures at the Scale of the Known Protein Universe’. Nature 622 (7983): 637–45. https://doi.org/10.1038/s41586-023-06510-w. [6] The UniProt Consortium. 2023. ‘UniProt: The Universal Protein Knowledgebase in 2023’. Nucleic Acids Research 51 (D1): D523–31. https://doi.org/10.1093/nar/gkac1052. [7] Schoch, Conrad L., Stacy Ciufo, Mikhail Domrachev, et al. 2020. ‘NCBI Taxonomy: A Comprehensive Update on Curation, Resources and Tools’. Database 2020 (January): baaa062. https://doi.org/10.1093/database/baaa062. [8] Smith, Martin R. 2020. ‘Information Theoretic Generalized Robinson–Foulds Metrics for Comparing Phylogenetic Trees’. Bioinformatics 36 (20): 5007–13. https://doi.org/10.1093/bioinformatics/btaa614. [9] Baran, Dror, M. Gabriele Pszolla, Gideon D. Lapidoth, et al. 2017. ‘Principles for Computational Design of Binding Antibodies’. Proceedings of the National Academy of Sciences 114 (41): 10900–10905. https://doi.org/10.1073/pnas.1707171114. [10] Dunbar, James, Konrad Krawczyk, Jinwoo Leem, et al. 2014. ‘SAbDab: The Structural Antibody Database’. Nucleic Acids Research 42 (Database issue): D1140–46. https://doi.org/10.1093/nar/gkt1043. [11] Schneider, Constantin, Matthew I. J. Raybould, and Charlotte M. Deane. 2022. ‘SAbDab in the Age of Biotherapeutics: Updates Including SAbDab-Nano, the Nanobody Structure Tracker’. Nucleic Acids Research 50 (D1): D1368–72. https://doi.org/10.1093/nar/gkab1050.