BGC Atlas - database tables (bgc_atlas_v2), snapshot 2026-09-19
Three forms of the same 42 tables:
.tsv.gz tab-separated, header row, gzip-compressed. The default form.
pandas.read_csv(path, sep='\t') | DuckDB read_csv_auto(path)
Bundled together as bgc_atlas_v2_tables_tsv.tar.gz
bgc_atlas_v2_tables_parquet.tar.gz
the same tables as Parquet (zstd). Smaller and typed; column
types come from the PostgreSQL catalog, not from inference.
bgc_atlas_v2_core.dump
PostgreSQL custom-format dump of the core relational subset
(31 tables), schema + data + constraints:
createdb mydb
pg_restore --no-owner -d mydb bgc_atlas_v2_core.dump
HOW TO JOIN
-----------
Families: bgc_gcf(bgc_id, gcf_id, distance, is_primary, rank)
A BGC may belong to more than one family. is_primary marks the one
the site displays, rank orders the rest, distance is the clustering
distance. Filter on is_primary unless you specifically want the
multi-family view.
14,606,494 of 16,491,574 BGCs have a family (88.6%).
The remaining 1,885,080 cluster with nothing. That is
not an error and not a novelty signal: they are predominantly
short contig-edge fragments, too small to place in a family.
Never report a family count without this coverage figure.
Taxonomy: bgc_taxonomy(bgc_id, taxonomy_id, rank, support, method) -> taxonomies
Sequences: bgcs.id is the join key to the region GenBank and FASTA downloads.
It is uuid5 of the region FILENAME, not of the sequence, so an
identical BGC in two genomes gets two different ids.
REPRODUCING THE PUBLISHED BGC COUNT
-----------------------------------
The site reports 16,449,773 BGCs, not the 16,491,574 rows in bgcs.
The difference is redundancy: some OMD-v2 genomes are also held by GTDB.
16,491,574 rows in bgcs
- 41,801 rows in redundant_bgc (one bgc_id per redundant copy)
= 16,449,773 the published corpus
sample_redundancy carries the same information at sample level, with the
canonical sample each duplicate maps to.
CONSISTENCY
-----------
Each table was copied in its own transaction, so this is a point-in-time
snapshot per table rather than one globally consistent read. Row counts were
reconciled against the live database at build time and matched exactly.
TABLES
------
analyses 928,326 rows
antismash_runs 883,860 rows
assemblies 882,105 rows
assembly_analyses 908,851 rows
assembly_taxonomy 752,799 rows
bgc_features 16,491,488 rows
bgc_gcf 19,254,436 rows
bgc_gene_domains 573 rows
bgc_genes 1,009 rows
bgc_taxonomy 16,386,489 rows
bgcs 16,491,574 rows
biomes 550 rows
data_sources 17 rows
endemism_distance_decay 123 rows
gcf_cooccurrence_stats 1 row
gcf_hierarchy 371,251 rows
gcf_mibig_anchor 2,067 rows
gcf_study_count 673,956 rows
gcf_taxonomy_distribution_mv 764,286 rows
gcfs 764,320 rows
mgnify_genome_catalogues 19 rows
mgnify_genomes 56,782 rows
mgnify_publications 1,785 rows
mgnify_super_studies 10 rows
redundant_bgc 41,801 rows
run_analyses 25,022 rows
run_assemblies 60,377 rows
run_studies 56,414 rows
runs 67,284 rows
sample_analyses 908,115 rows
sample_assemblies 881,942 rows
sample_biomes 99,995 rows
sample_measurements 386,209 rows
sample_redundancy 5,855 rows
sample_runs 67,279 rows
samples 863,831 rows
studies 53,240 rows
study_bgc_summary 45,181 rows
study_biomes 0 rows
study_gcf 2,662,435 rows
study_samples 353,699 rows
taxonomies 266,571 rows
VERIFY
------
sha256sum -c SHA256SUMS