About BGC Atlas
BGC Atlas is a searchable catalogue of biosynthetic gene clusters (BGCs) from publicly available metagenomes and genomes. It holds — BGCs, grouped into gene cluster families and annotated with the taxonomy and the environment each one came from.
Browse by cluster, family, sample or study; search by antiSMASH property or protein sequence; or take the whole database from downloads.
Caner Bağcı, Alec Talamas-Tanner, Nadine Ziemert
BGC Atlas v2: biosynthetic gene clusters with taxonomic and environmental context at scale
bioRxiv, 18 September 2026 · Preprint
Caner Bağcı, Matin Nuhamunada, Hemant Goyat, Casimir Ladanyi, Ludek Sehnal, Kai Blin, Satria A Kautsar, Azat Tagirdzhanov, Alexey Gurevich, Shrikant Mantri, Christian von Mering, Daniel Udwary, Marnix H Medema, Tilmann Weber, Nadine Ziemert
BGC Atlas: a web resource for exploring the global chemical diversity encoded in bacterial genomes
Nucleic Acids Research, Volume 53, Issue D1, 6 January 2025, Pages D618–D624
Every assembly comes from a public repository or a published dataset. Counts below are a breakdown by source; the assembly and BGC totals are on the home page.
Metagenome assemblies
Communities assembled directly from environmental reads.
| Source | Assemblies | BGCs |
|---|---|---|
| Loading… | ||
Genome catalogs
Metagenome-assembled, single-cell and isolate genomes.
| Source | Assemblies | BGCs |
|---|---|---|
| Loading… | ||
Detection
BGCs were identified with metaSMASH 1.0.3, a re-engineered fork of antiSMASH built for metagenome-scale assemblies. It preserves antiSMASH's detection and annotation logic, so each region carries the same products, product categories and domain content. A BGC is counted as complete when it does not run off the end of its contig: — of — qualify. The rest are truncated by fragmentary assembly rather than by anything biological, so completeness says nothing about whether a cluster is functional.
Clustering
BGCs are grouped into gene cluster families with BiG-SLiCE 2.0.2 at a single distance threshold of 0.4. Clustering runs in two passes: a seed pass over BGCs that are complete or at least median length, then a second pass over what the seed could not place, which produces the fragmented families (the F in GCF-F-…). 371,217 seed and 246,709 fragmented families are listed.
88.6% of BGCs fall into a family. The rest are mostly regions truncated at a contig edge — short fragments with too little to match on, rather than novel chemistry.
A BGC can sit in more than one family: 2,499,475 do. Each one has a primary family, and the others are listed beside it.
1,348 listed families contain a MIBiG reference cluster and carry the name of the compound it makes.
The data BGC Atlas produces — cluster annotations, family assignments, taxonomy and the harmonised environmental metadata — is released under CC BY 4.0. Use it for anything, including commercially, provided you cite the resource.
The underlying assemblies and genomes are not ours to relicense. They remain subject to the terms of the repository or publication they came from — see the sources listed above, each of which links to its own database or paper.
Bulk files are on the downloads page. If you have BGC data you would like included, see contribute.
BGC Atlas is built and maintained by Caner Bağcı in the Translational Genome Mining for Natural Products group at the University of Tübingen. The full list of contributors is in the citations above.
Funded by the German Center for Infection Research (DZIF) and the Volkswagen Foundation.