BGC Atlas - database tables (bgc_atlas_v2), snapshot 2026-09-19 Three forms of the same 42 tables: .tsv.gz tab-separated, header row, gzip-compressed. The default form. pandas.read_csv(path, sep='\t') | DuckDB read_csv_auto(path) Bundled together as bgc_atlas_v2_tables_tsv.tar.gz bgc_atlas_v2_tables_parquet.tar.gz the same tables as Parquet (zstd). Smaller and typed; column types come from the PostgreSQL catalog, not from inference. bgc_atlas_v2_core.dump PostgreSQL custom-format dump of the core relational subset (31 tables), schema + data + constraints: createdb mydb pg_restore --no-owner -d mydb bgc_atlas_v2_core.dump HOW TO JOIN ----------- Families: bgc_gcf(bgc_id, gcf_id, distance, is_primary, rank) A BGC may belong to more than one family. is_primary marks the one the site displays, rank orders the rest, distance is the clustering distance. Filter on is_primary unless you specifically want the multi-family view. 14,606,494 of 16,491,574 BGCs have a family (88.6%). The remaining 1,885,080 cluster with nothing. That is not an error and not a novelty signal: they are predominantly short contig-edge fragments, too small to place in a family. Never report a family count without this coverage figure. Taxonomy: bgc_taxonomy(bgc_id, taxonomy_id, rank, support, method) -> taxonomies Sequences: bgcs.id is the join key to the region GenBank and FASTA downloads. It is uuid5 of the region FILENAME, not of the sequence, so an identical BGC in two genomes gets two different ids. REPRODUCING THE PUBLISHED BGC COUNT ----------------------------------- The site reports 16,449,773 BGCs, not the 16,491,574 rows in bgcs. The difference is redundancy: some OMD-v2 genomes are also held by GTDB. 16,491,574 rows in bgcs - 41,801 rows in redundant_bgc (one bgc_id per redundant copy) = 16,449,773 the published corpus sample_redundancy carries the same information at sample level, with the canonical sample each duplicate maps to. CONSISTENCY ----------- Each table was copied in its own transaction, so this is a point-in-time snapshot per table rather than one globally consistent read. Row counts were reconciled against the live database at build time and matched exactly. TABLES ------ analyses 928,326 rows antismash_runs 883,860 rows assemblies 882,105 rows assembly_analyses 908,851 rows assembly_taxonomy 752,799 rows bgc_features 16,491,488 rows bgc_gcf 19,254,436 rows bgc_gene_domains 573 rows bgc_genes 1,009 rows bgc_taxonomy 16,386,489 rows bgcs 16,491,574 rows biomes 550 rows data_sources 17 rows endemism_distance_decay 123 rows gcf_cooccurrence_stats 1 row gcf_hierarchy 371,251 rows gcf_mibig_anchor 2,067 rows gcf_study_count 673,956 rows gcf_taxonomy_distribution_mv 764,286 rows gcfs 764,320 rows mgnify_genome_catalogues 19 rows mgnify_genomes 56,782 rows mgnify_publications 1,785 rows mgnify_super_studies 10 rows redundant_bgc 41,801 rows run_analyses 25,022 rows run_assemblies 60,377 rows run_studies 56,414 rows runs 67,284 rows sample_analyses 908,115 rows sample_assemblies 881,942 rows sample_biomes 99,995 rows sample_measurements 386,209 rows sample_redundancy 5,855 rows sample_runs 67,279 rows samples 863,831 rows studies 53,240 rows study_bgc_summary 45,181 rows study_biomes 0 rows study_gcf 2,662,435 rows study_samples 353,699 rows taxonomies 266,571 rows VERIFY ------ sha256sum -c SHA256SUMS