Genome metadata
Complete sampling metadata and harmonized taxonomy for 542 SAR11 genomes and 20 outgroups, including assignment evidence and confidence fields, together with the CheckM2 v1.0.2 quality report for the 542 released SAR11 genomes.
Data access
Access the genome metadata and analysis resources underlying the SAR11 Genome Atlas.
Compact release files are available directly from the atlas. Larger sequence, annotation, analysis, and HMM archives are hosted on Zenodo; Zenodo links are marked Embargoed.
Release files
Availability is shown for each resource in the updated atlas release.
Complete sampling metadata and harmonized taxonomy for 542 SAR11 genomes and 20 outgroups, including assignment evidence and confidence fields, together with the CheckM2 v1.0.2 quality report for the 542 released SAR11 genomes.
Genome assemblies, predicted protein sequences, and GFF annotations for all 542 SAR11 genomes, distributed as versioned external archives.
Complete protein-level annotations for the 675,669-gene collection, distributed through an external data repository.
Core OrthoFinder 3 results for the 542-genome dataset, including orthogroup membership and gene-count tables, unassigned genes, overall and per-genome statistics, species-overlap counts, root-level hierarchical orthogroups, the labeled species tree, and run metadata.
Representative annotations for 4,577 orthogroups and the count tables underlying the KO and COG pie charts and Pfam bar chart in the OG Information Viewer.
One observed representative amino-acid sequence for each of the 4,577 orthogroups, provided for rapid exploratory OG assignment. Representatives were selected as the sequence with the lowest EMBOSS infoalign percent change from the OG multiple-alignment consensus. These files are the database source for the browser-based OG Search.
The 4,577 orthogroup profile HMMs are distributed as a combined HMM library. Individual OG profiles can be retrieved from the combined file with hmmfetch.
Resolved gene trees for the 3,411 orthogroups containing at least four protein sequences, generated with the OrthoFinder 3 default workflow. Amino-acid sequences were aligned with FAMSA, approximate maximum-likelihood trees were inferred with FastTree using its -fastest option, and the trees were rooted and resolved by OrthoFinder using its hybrid species-overlap/duplication-loss coalescent model.
The default species phylogeny was inferred with IQ-TREE 2 from SAR11_165 HMM profiles obtained from the Meren Lab workflow. SAR11_165 is a SAR11-focused subset of an earlier 200-gene Alphaproteobacterial SCG collection. Topology-based taxonomy assignments use only this SAR11_165 tree and require crown support of at least 0.95. The rooted, outgroup-pruned SAR11_165 tree is the default atlas phylogeny; bac120 and FastTree results are retained only as comparison trees.
Pairwise nucleotide- and amino-acid identity results for all 542 SAR11 genomes, calculated with FastANI v1.34 and CompareM v0.1.2 aai_wf. Both the original directional FastANI output and a symmetric ANI matrix are provided.
The edge table used by the Neighboring Network viewer and the complete gene-coordinate table underlying the neighborhood and operon analyses.
The network edge table used by the CORGIAS Network viewer was calculated with the rooted SAR11_165 IQ-TREE 2 phylogeny. Significant associations calculated with the rooted bac120 IQ-TREE 2 phylogeny and the complete SAR11_165 CORGIAS analysis results are also available.
High-similarity protein matches for all 542 genomes, generated against UniProtKB release 2026_01 with DIAMOND v2.1.10.164 using a minimum identity of 85% and --max-target-seqs 1. The AlphaFoldDB-linked subset additionally requires at least 80% query and subject coverage. Gene-level and orthogroup-level summary tables are provided.
A broader search against UniProtKB release 2026_01 was generated with DIAMOND using --max-target-seqs 10. In the combined structure-reference table, close matches require identity ≥85% with query and subject coverage ≥80%; homologous references require identity ≥30%, query and subject coverage ≥80%, and E-value ≤1e-5. The table contains the 8,057 AlphaFoldDB structure references used by the Web interface across 2,994 orthogroups.
Gene-level quantification results for 509 Tara Oceans metatranscriptomic runs mapped to the SAR11 protein-coding gene collection. Join the protein identifiers to the orthogroup assignments and sum TPM within each sample and orthogroup to reproduce the Expression Scores used by the Metatranscriptome Viewer.
The publication metadata table displayed on the Literature page, compiled for exploring the history and research themes of the SAR11 field.