BGC Atlas — BiG-SLiCE clustering runs (clustering v3), published 2026-09-21 The two BiG-SLiCE 2.0.2 runs the site's gene cluster families were built from, each published whole as a single `result/data.db`. bigslice-v3-seed.tar.gz 13.6 GB extracts to bigslice-output-v3/result/data.db (31.3 GB) sha256 5e75c403be1cb3a6863a25ac225f2925a4eb93301bdf6dad03189ce04291becf the main atlas families, clustered from the dedup-95 seed clustering 1: 371,651 families at threshold 0.4 bigslice-v3-residual.tar.gz 6.1 GB extracts to bigslice-output-residual-v3/result/data.db (14.5 GB) sha256 874757a92311362b2fb01ff5c051502546f857615e35bb6a48779fbf94cb7226 de-novo families of what the seed run could not place clustering 1: 394,186 families at threshold 0.4 clustering 2: 107,451 families at threshold 0.7 FAMILY IDS ---------- `gcf.id` in each file maps to a family on the site: seed run GCF-{id:06d} e.g. gcf.id 336 -> GCF-000336 residual run GCF-F-{400000+id:06d} e.g. gcf.id 42 -> GCF-F-400042 The two are INDEPENDENT id spaces. gcf.id 42 in one is unrelated to gcf.id 42 in the other. Residual clustering 2 (@0.7) was never loaded and has no pages. `bgc.name` is `genomes/.regionNNN`; strip the prefix and suffix to get `bgcs.id`, the key the other downloads use. VERIFY ------ sha256sum -c SHA256SUMS