# BGC UUID → source/sample table

Generated from `bgc_atlas_v2` (Postgres). One row per BGC = one row per
`bgcs.id`. Files:

- `bgc_uuid_source.parquet` — zstd-compressed, columnar (preferred).
- `bgc_uuid_source.tsv.gz` — gzipped TSV, identical content (fallback).

Row count: **16,491,574** (every BGC in `bgc_atlas_v2`).

## The join key

`bgc_id` **is** the bare UUID that appears as the gbk member name (minus `.gbk`)
and the FASTA header (minus `>`). Join your gbk/FASTA UUIDs directly on this
column — no translation needed. Verified: e.g.
`ef1e0ff2-a653-52b8-a80a-6026c43c8374` = region
`GB_GCA_000003955.1_CM000741.1.region001`.

### Important: these UUIDs are NOT sequence-content hashes

`bgc_id = uuidv5(namespace, "<genome>_<contig>.region<NNN>")` — derived from the
**region filename**, not the nucleotide sequence. Consequences:

- An identical BGC in two genomes gets **two different `bgc_id`s**. UUIDs do
  **not** collapse across genomes.
- Therefore **one row per `bgc_id` already equals one row per occurrence** — the
  occurrence (genome + contig + region) is baked into the UUID. There is no
  fan-out to reconstruct; `bgc_id` is unique in this table.
- "This GCF occurs in N samples across M sources" = group your GCF's member
  `bgc_id`s by `sample_id` / `source` in this table.

## Columns

| column | source | notes |
|---|---|---|
| `bgc_id` | `bgcs.id` | **PK / join key.** Bare UUID = gbk member / FASTA header. Never null. |
| `source` | `samples.data_source_ids[1]` → `data_sources.name` | **Authoritative provenance** (metalog, MGnify, GTDB, custom_*, SMAG, GEM, TPMC, OMD v2, SPIRE MAGs, Microflora Danica ...). This is the website-facing, per-sample source that powers "N samples across M sources". Null for ~6,973 orphan BGCs whose assembly has no sample. |
| `assembly_source` | `assemblies.data_source_ids[1]` → `data_sources.name` | Data provider of the **assembly**. Usually equals `source`, but **not always**: the 5.03M `source=metalog` BGCs have `assembly_source=SPIRE` (metalog samples, SPIRE-assembled). Null for MGnify/metalog assemblies that carry no assembly-level source. |
| `sample_id` | `samples.id` | Internal sample UUID. The per-sample occurrence key. Null for orphans. |
| `sample_accession` | `samples.accession` | External sample accession (SAMN…/ERS…/GCA…/MFD… etc.), if present. |
| `genome_accession` | `assemblies.accession` | Original genome/assembly accession the region was called on. |
| `contig` | `bgcs.contig` | Contig/scaffold within the genome. |
| `environment_biome` | `samples.environment_biome` | Biome label (metagenomes); often null for isolates. |
| `gtdb_lineage` | `taxonomies.gtdb_lineage` via `bgcs.taxonomy_id` | Full GTDB lineage string; mainly populated for GTDB / isolate sources. |

### Source breakdown (`source`)

| source | rows | | source | rows |
|---|--:|---|---|--:|
| MGnify | 6,959,864 | | Microflora Danica (short-read) | 132,251 |
| metalog | 5,036,087 | | Microflora Danica (long-read) | 95,284 |
| GTDB | 1,927,971 | | custom_deep_sea_hydrothermal | 35,554 |
| Ocean Microbiomics Database v2 | 863,100 | | custom_ocean_surface | 20,222 |
| SPIRE MAGs | 435,994 | | custom_ziemert_lab | 15,413 |
| custom_microflora_danica | 370,910 | | OWC | 10,499 |
| SMAG | 249,200 | | custom_fram_ras | 1,321 |
| GEM | 186,972 | | *(null — orphan, no sample)* | 6,973 |
| TPMC | 143,959 | | | |

Note: for the 17,907 samples with two data sources, `source` takes the first
(`data_source_ids[1]`); `assembly_source` may differ.
