Update indexing constraints and CLI options

This commit introduces new constraints for kmer size and minimizer selection, defines super-kmers, and adds extensive new command-line options for filtering, conversion, merging, and packing indices. Documentation across the codebase has also been updated.
Eric Coissac committed 2026-09-12 14:00:23 +02:00
1 parent c1e139c597
commit 5313788f7c
22 files changed
+69 -76

No files matched your search

+3 -3
@@ -15,7 +15,7 @@ All functionality is exposed through a single binary, `obikmer`, organized as su
## Commands
| Command | Purpose |
|---|---|
| --- | --- |
| [`superkmer`](usage-superkmer) | Extract super-kmers from a sequence file and write them to stdout |
| [`index`](usage-index_command) | Build a genome index |
| [`merge`](usage-merge) | Merge multiple indexes into one |
@@ -48,7 +48,7 @@ See [Genome predicates and taxonomy paths](usage-predicates) for the selection l
These constraints are checked at startup; an invalid value exits immediately with an error.
| Parameter | Constraint | Reason |
|---|---|---|
| --- | --- | --- |
| $k$ (`--kmer-size`) | odd, $k \in [11, 31]$ | odd length guarantees the canonical form is always well defined; the range keeps a kmer within a 64-bit word while retaining specificity |
| $m$ (`--minimizer-size`) | odd, $3 \le m \le k-1$ | same palindrome argument as $k$; must be strictly shorter than the kmer |
| $z$ (`-z`, approximate evidence only) | $z \le k-1$ | the effective indexed kmer size is $k-z+1$ |
@@ -58,7 +58,7 @@ These constraints are checked at startup; an invalid value exits immediately wit
Genome labels are arbitrary Unicode strings, with the following restrictions:
| Character | Forbidden | Reason |
|---|---|---|
| --- | --- | --- |
| `/` | yes | filesystem path separator |
| `=` | yes | separator used by `--new-label` |
| `\0` | yes | null byte |
+1 -3
@@ -26,8 +26,7 @@ Two verification modes are available, selected at build time (`index --approx`)
## On-disk layout
```
<index_root>/
<index_root>/
index.meta global configuration (k, minimizer size, partition count,
evidence mode, whether counts are stored) and genome list/metadata
scatter.done / count.done / index.done build-progress sentinels
@@ -45,7 +44,6 @@ Two verification modes are available, selected at build time (`index --approx`)
counts/ per-genome kmer counts (if counts were requested)
presence/ per-genome presence/absence bits
layer_1/, layer_2/, ... added by later merges, same internal structure
```
`unitigs.bin` is the only file from which the indexed kmer content can be fully recovered; it is always retained. Every other file (MPHF, evidence, counts) is derived from it.
+1 -1
@@ -5,7 +5,7 @@
Every nucleotide is encoded on 2 bits, most-significant-bit first within each word:
| Base | Encoding |
|------|----------|
| --- | --- |
| A | `00` |
| C | `01` |
| G | `10` |
+1 -1
@@ -22,7 +22,7 @@ A value near 0 indicates low complexity (e.g. a homopolymer run); near 1 indicat
## Final score
The filter evaluates $\hat{H}(ws)$ for every word size from 1 to ws_max and keeps the minimum:
The filter evaluates $\hat{H}(ws)$ for every word size from 1 to ws\_max and keeps the minimum:
$$\text{entropy}(kmer) = \min_{ws=1}^{ws_{\max}} \hat{H}(ws)$$
+4 -6
@@ -6,9 +6,7 @@ An index is split into a fixed number of **partitions**, each handling an indepe
The canonical minimizer of a super-kmer (see [Minimizer selection](theory-minimizer_selection)) is hashed to produce a $p$-bit routing value that selects the destination partition:
```
canonical minimizer → hash(minimizer) → p-bit value → partition index
```
canonical minimizer → hash(minimizer) → p-bit value → partition index
Within a partition, kmers are indexed as plain values via a minimal perfect hash function (see [On-disk storage](formats-index_layout)); the minimizer plays no further role once a super-kmer has reached its partition.
@@ -17,10 +15,10 @@ Within a partition, kmers are indexed as plain values via a minimal perfect hash
Even though $H$ already makes minimizer values well-distributed (see [Minimizer selection](theory-minimizer_selection)), choosing $p$ well below the number of bits available in the minimizer ($2m$) leaves a comfortable entropy margin, provided the number of distinct minimizers actually observed is much larger than the number of partitions.
| Minimizer size $m$ | Minimizer bits ($2m$) | Typical partition-index bits $p$ | Partitions |
|----|-----------|-----------|------------|
| --- | --- | --- | --- |
| 9 | 18 | 6–8 | 64–256 |
| 11 | 22 | 8–10 | 256–1 024 |
| 13 | 26 | 10–12 | 1 024–4 096|
| 15 | 30 | 10–14 | 1 024–16 384|
| 13 | 26 | 10–12 | 1 024–4 096 |
| 15 | 30 | 10–14 | 1 024–16 384 |
The number of partitions must satisfy $p \le 2m$, and in practice $p$ is chosen well below that bound to leave a comfortable entropy margin. For $k=31$, $m=13$, $p=10$ (1024 partitions), partition load is well balanced on real genomic data.
+4 -4
@@ -11,7 +11,7 @@ A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately ([Golan & Shur 2025](#ref-Golan2025-xf); [Zheng *et al.* 2020](#ref-Zheng2020-ji)):
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately ([Golan & Shur 2025](#ref-Golan2025-xf); [Zheng et al. 2020](#ref-Zheng2020-ji)):
$$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$
@@ -25,17 +25,17 @@ Super-kmers are the unit of work used throughout construction and querying: sequ
## Bibliography
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
<div id="refs" class="references csl-bib-body hanging-indent">
<div id="ref-Golan2025-xf" class="csl-entry">
Golan, S. & Shur, A.M. (2025). [Expected density of random minimizers](https://doi.org/10.1007/978-3-031-82670-2\_25). In: *Lecture notes in computer science*, Lecture notes in computer science. Springer Nature Switzerland, Cham, pp. 347–360.
Golan, S. & Shur, A.M. (2025). <a href="https://doi.org/10.1007/978-3-031-82670-2\_25">Expected density of random minimizers</a>. In: <span style="font-style: italic;">Lecture Notes in Computer Science</span>, Lecture Notes in Computer Science. Springer Nature Switzerland, Cham, pp. 347–360.
</div>
<div id="ref-Zheng2020-ji" class="csl-entry">
Zheng, H., Kingsford, C. & Marçais, G. (2020). [Improved design and analysis of practical minimizers](https://doi.org/10.1093/bioinformatics/btaa472). *Bioinformatics (Oxford, England)*, 36, i119–i127.
Zheng, H., Kingsford, C. & Marçais, G. (2020). <a href="https://doi.org/10.1093/bioinformatics/btaa472">Improved design and analysis of practical minimizers</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 36, i119–i127.
</div>
+9 -9
@@ -8,7 +8,7 @@ The minimizer partitions a sequence into super-kmers: maximal runs of overlappin
## Hash-based ("random") minimizer
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions ([Golan & Shur 2025](#ref-Golan2025-xf); [Kille *et al.* 2023](#ref-Kille2023-px); [Pan & Reinert 2024](#ref-Pan2024-hb); [Zheng *et al.* 2020](#ref-Zheng2020-ji), [2021](#ref-Zheng2021-cc)).
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions ([Golan & Shur 2025](#ref-Golan2025-xf); [Kille et al. 2023](#ref-Kille2023-px); [Pan & Reinert 2024](#ref-Pan2024-hb); [Zheng et al. 2020](#ref-Zheng2020-ji); [2021](#ref-Zheng2021-cc)).
Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition.
@@ -24,7 +24,7 @@ The hash function is a 64-bit mixing function (splitmix64-style finalizer) appli
$$H(x) = \text{mix64}(x \oplus s)$$
``` text
```text
H(x):
x ← x ⊕ s
x ← x ⊕ (x >> 30)
@@ -37,7 +37,7 @@ H(x):
The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes; if one of them happened to be the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$, since $\text{mix64}(0) = 0$ is a fixed point — it would win far more windows than the composition-uniform behavior established above predicts, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. Exhaustive checks confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$:
| $m$ | argmin (canonical) | decoded sequence | minimal period |
|-----|--------------------|-------------------|----------------|
| --- | --- | --- | --- |
| 3 | 16 | `CAA` | 3 |
| 5 | 78 | `ACATG` | 5 |
| 7 | 5512 | `CCCGAGA` | 7 |
@@ -58,35 +58,35 @@ See [Partitioning and indexing architecture](theory-indexing_architecture) for m
## Bibliography
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
<div id="refs" class="references csl-bib-body hanging-indent">
<div id="ref-Golan2025-xf" class="csl-entry">
Golan, S. & Shur, A.M. (2025). [Expected density of random minimizers](https://doi.org/10.1007/978-3-031-82670-2\_25). In: *Lecture notes in computer science*, Lecture notes in computer science. Springer Nature Switzerland, Cham, pp. 347–360.
Golan, S. & Shur, A.M. (2025). <a href="https://doi.org/10.1007/978-3-031-82670-2\_25">Expected density of random minimizers</a>. In: <span style="font-style: italic;">Lecture Notes in Computer Science</span>, Lecture Notes in Computer Science. Springer Nature Switzerland, Cham, pp. 347–360.
</div>
<div id="ref-Kille2023-px" class="csl-entry">
Kille, B., Garrison, E., Treangen, T.J. & Phillippy, A.M. (2023). [Minmers are a generalization of minimizers that enable unbiased local jaccard estimation](https://doi.org/10.1093/bioinformatics/btad512). *Bioinformatics (Oxford, England)*, 39.
Kille, B., Garrison, E., Treangen, T.J. & Phillippy, A.M. (2023). <a href="https://doi.org/10.1093/bioinformatics/btad512">Minmers are a generalization of minimizers that enable unbiased local Jaccard estimation</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 39.
</div>
<div id="ref-Pan2024-hb" class="csl-entry">
Pan, C. & Reinert, K. (2024). [A simple refined DNA minimizer operator enables 2-fold faster computation](https://doi.org/10.1093/bioinformatics/btae045). *Bioinformatics (Oxford, England)*, 40.
Pan, C. & Reinert, K. (2024). <a href="https://doi.org/10.1093/bioinformatics/btae045">A simple refined DNA minimizer operator enables 2-fold faster computation</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 40.
</div>
<div id="ref-Zheng2020-ji" class="csl-entry">
Zheng, H., Kingsford, C. & Marçais, G. (2020). [Improved design and analysis of practical minimizers](https://doi.org/10.1093/bioinformatics/btaa472). *Bioinformatics (Oxford, England)*, 36, i119–i127.
Zheng, H., Kingsford, C. & Marçais, G. (2020). <a href="https://doi.org/10.1093/bioinformatics/btaa472">Improved design and analysis of practical minimizers</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 36, i119–i127.
</div>
<div id="ref-Zheng2021-cc" class="csl-entry">
Zheng, H., Kingsford, C. & Marçais, G. (2021). [Sequence-specific minimizers via polar sets](https://doi.org/10.1093/bioinformatics/btab313). *Bioinformatics (Oxford, England)*, 37, i187–i195.
Zheng, H., Kingsford, C. & Marçais, G. (2021). <a href="https://doi.org/10.1093/bioinformatics/btab313">Sequence-specific minimizers via polar sets</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 37, i187–i195.
</div>
+2 -2
@@ -10,13 +10,13 @@ obikmer annotate INDEX --dump
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INDEX` | Index directory to annotate (modified in place) |
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--csv` | — | CSV file of metadata to apply (must contain an id column); required unless `--dump` is used |
| `--sep` | `,` | CSV field separator |
| `--id-col` | `id` | Name of the column containing genome labels |
+2 -2
@@ -9,7 +9,7 @@ obikmer convert INDEX (--exact-evidence | --approx-evidence BITS | --hybrid-evid
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INDEX` | Index directory to convert (modified in place) |
## Options
@@ -17,7 +17,7 @@ obikmer convert INDEX (--exact-evidence | --approx-evidence BITS | --hybrid-evid
Exactly one of the first three is required:
| Option | Description |
|---|---|
| --- | --- |
| `--exact-evidence` | Convert to exact evidence (zero false positives) |
| `--approx-evidence BITS` | Convert to approximate (fingerprint-only) evidence; `BITS` = fingerprint bits per slot (b) |
| `--hybrid-evidence` | Convert to hybrid evidence (both exact and approximate bundles kept) |
+2 -2
@@ -9,13 +9,13 @@ obikmer dump INDEX [OPTIONS]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INDEX` | Index directory to dump |
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--force-presence` | off | Output presence/absence (0/1) even if the index stores counts |
| `--debug` | off | Prefix each row with the partition and layer columns |
| `--head N` | none | Limit output to the first N kmers |
+1 -1
@@ -9,7 +9,7 @@ obikmer estimate [OPTIONS]
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `-k, --kmer-size` | `31` | Kmer size used at query time (matches `index`'s `--kmer-size`) |
| `-z, --findere-z` | none | Findere z parameter |
| `--evidence-bits` | none | Fingerprint bits per slot (b) |
+3 -3
@@ -9,13 +9,13 @@ obikmer filter SOURCE -o OUTPUT [OPTIONS]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `SOURCE` | Source index directory |
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `-o, --output` | — (required) | Output index directory |
| `-f, --force` | off | Overwrite an existing output directory |
| `--presence` | off | Output presence/absence instead of counts |
@@ -27,7 +27,7 @@ obikmer filter SOURCE -o OUTPUT [OPTIONS]
## Predicate options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--ingroup` | none | Ingroup predicate (repeatable; each occurrence is ANDed) |
| `--outgroup` | none | Outgroup predicate (repeatable; each occurrence is ORed) |
| `--min-count` | 0, or group size + N if negative | Minimum number of ingroup genomes carrying the kmer |
+3 -3
@@ -9,18 +9,18 @@ obikmer index -o OUTPUT [OPTIONS] [INPUTS...]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INPUTS...` | Input sequence files or directories (FASTA/FASTQ/GenBank, gzip optional). If omitted, reads from stdin. |
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `-o, --output` | — (required) | Output index directory |
| `--force` | off | Overwrite an existing output directory |
| `--label` | input file name without extension | Genome label stored in the index |
| `--meta KEY=VALUE` | none | Attach a categorical metadata field to the genome (repeatable) |
| `-k, --kmer-size` | `31` | Kmer size (odd, in [11, 31]) |
| `-k, --kmer-size` | `31` | Kmer size (odd, in \[11, 31\]) |
| `-m, --minimizer-size` | `11` | Minimizer size (odd, in $[3, k-1]$) |
| `--theta` | `0.7` | Entropy threshold for the low-complexity filter |
| `--level-max` | `6` | Maximum sub-word size for the entropy score |
+2 -2
@@ -9,13 +9,13 @@ obikmer merge -o OUTPUT SOURCE... [OPTIONS]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `SOURCE...` | Index directories to merge (at least one required) |
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `-o, --output` | — (required) | Output index directory |
| `--force` | off | Overwrite an existing output directory |
| `--force-presence` | off | Store the merged index as presence/absence even if all sources have counts |
+2 -2
@@ -9,13 +9,13 @@ obikmer pack INDEX [--sparse]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INDEX` | Index directory to pack (modified in place) |
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--sparse` | off | Pack presence/absence and count matrices into a sparse, deduplicated format instead of the dense one |
The index directory is locked for exclusive access while packing.
+16 -17
@@ -9,13 +9,13 @@ obikmer phylo INDEX [OPTIONS]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INDEX` | Index directory |
## Distance matrix (`--distance`)
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--distance` | `jaccard` | See the two tables below for the full list of accepted values |
| `--gamma-shape ALPHA\|auto` | none | Rate-heterogeneity correction, for `snp-*` values that support it (see below). Either a fixed $\alpha$ or `auto`/`estimate` to fit it from the data (see "Automatic $\alpha$ estimation" below). No effect on the other values; rejected if given together with a value that doesn't support it |
| `--presence-threshold` | `1` | Minimum count for a kmer to be considered present, for `jaccard`/`mash` on a count index |
@@ -30,7 +30,7 @@ Every value routes to one of two independent computations:
### Whole-index metrics
| Value | Definition |
|---|---|
| --- | --- |
| `jaccard` | $D = 1 - \dfrac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}$ over the sets of kmers present in each genome |
| `mash` | derived from the Jaccard distance via $D = -\dfrac{1}{k} \ln\!\left(\dfrac{2J}{1+J}\right)$ where $J = 1 - D_{\text{jaccard}}$ and $k$ is the index's kmer size; clamped to 1.0 when $J \le 0$ |
| `hamming` | number of kmer positions where presence differs between the two genomes (presence index only, not normalized): $D = \sum_i \mathbb{1}[a_i \ne b_i]$ |
@@ -112,7 +112,7 @@ Without `-o`, the matrix goes to stdout in relaxed-PHYLIP format (`n` on the fir
## `--exclude-genome`, `--min-shared-family`
| Option | Description |
|---|---|
| --- | --- |
| `--exclude-genome LABEL` | Exclude a genome (repeatable). Drops its row/column from the distance/shared-kmer matrix output, and removes it from the sampling used by `--pseudo-alignment`/`--sankoff`/a `snp-*` `--distance` value. Does not change the value computed for any remaining pair |
| `--min-shared-family N` | Auto-exclude, on top of `--exclude-genome`, any genome whose mean shared-family count against every other genome (see "Family Overlap" below) falls below `N`. Applies only to `--pseudo-alignment`/`--sankoff`/`snp-*` `--distance` — never to the whole-index metrics or their matrix/NJ/UPGMA output |
@@ -123,7 +123,7 @@ Neighbor-Joining and UPGMA trees (`--nj`/`--upgma`) are always built from every
Requires the sibling annex, built once per index:
| Option | Description |
|---|---|
| --- | --- |
| `--sibling-annex` | Build (or rebuild) the sibling-count/minorant annex — prerequisite for every option in this section, and for a `snp-*` `--distance` value |
| `--sibling-stats` | Write `<prefix>_siblings.csv`: the family-size distribution, per genome and globally |
| `--sibling-hist` | Print the global family-size histogram (1-4 members) only |
@@ -138,7 +138,7 @@ A family is eligible for a genome pair $(i,j)$ only if both genomes carry exactl
`<prefix>_siblings.csv` — family size = number of distinct central bases observed at a family (1-4).
| Column | Meaning |
|---|---|
| --- | --- |
| `genome` | genome label, or the literal `global` for the last row |
| `1`, `2`, `3`, `4` | for a genome row: number of families of that size where the genome carries ≥ 1 member. For the `global` row: the actual deduplicated family-size histogram — not the sum of the rows above |
@@ -153,7 +153,7 @@ A family is eligible for a genome pair $(i,j)$ only if both genomes carry exactl
`<prefix>_alignment.fasta` — one record per non-excluded genome, one column per variable family (family size ≥ 2). Each site is IUPAC-coded from the genome's presence mask at that family: a single observed form → the plain base; several forms → the matching IUPAC ambiguity code; no form → `-`.
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--subsample N` | none (mandatory here) | Target number of families to sample |
| `--free-loss` | off | Treat a genome carrying none of a family's observed members as missing data (`?`) instead of `-` |
| `--no-ambiguity` | off | Treat a genome carrying more than one member of a family as missing data (`?`) instead of an IUPAC ambiguity code |
@@ -171,7 +171,7 @@ Every invocation, with or without `--session`, draws its own fresh random sample
### `--session`: reusing a sample across separate commands
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--session DIR` | none | Persist the sample (and, for `--sankoff`, its calibration/alignment) in `DIR` so a later, separate `obikmer phylo` invocation with the exact same selection parameters restores it instead of resampling |
| `--session-force` | off | With `--session DIR`: overwrite its saved parameters and cached sample instead of erroring out when this run's parameters don't match. No effect without `--session` |
@@ -188,7 +188,7 @@ Without `--subsample`, every variable family (family size ≥ 2) is used. With `
`<prefix>_entropy.csv` has one row per family visited:
| Column | Meaning |
|---|---|
| --- | --- |
| `layer` | an internal index-layer identifier — stable within one run, not meaningful across indexes |
| `family_idx` | the family's position within that layer |
| `entropy15` | Shannon entropy (bits) over the 16 possible states (the 15 non-empty subsets of `{A,C,G,T}`), genomes absent from the family excluded from the count |
@@ -207,7 +207,7 @@ The first `phylo` run on a given index that uses `--entropy`/`--entropy-sd` pays
## Sankoff calibration and phylogenetic exports
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--sankoff` | off | Calibrate a 16-state parsimony cost matrix and matching pseudo-alignment. Requires `--subsample N` |
| `--sankoff-ratio-ceiling` | `0.5` | Exclude genome pairs whose raw SNP ratio exceeds this value from the base-composition part of the calibration |
| `--free-loss` | off | Recode a family's non-detection as the `?` missing-data symbol instead of an ordinary, costed state, throughout `--sankoff` and every export built from it |
@@ -242,7 +242,7 @@ With `-o/--output PREFIX`, the relevant subset of the files below is written. Wi
### Distance matrix
| File | Written by | Format | Content |
|---|---|---|---|
| --- | --- | --- | --- |
| `<prefix>_dist.phy` | always, unless `--csv` | relaxed PHYLIP | the `--distance` matrix |
| `<prefix>_dist.csv` | `--csv` | CSV matrix | the `--distance` matrix, 6 decimals |
| `<prefix>_shared.csv` | `--shared-kmers` | CSV matrix | shared-kmer count per genome pair (integers) |
@@ -254,7 +254,7 @@ CSV matrix layout (`_dist.csv`, `_shared.csv`, `_family_overlap.csv`): header `g
### Central-position SNP model
| File | Written by | Format | Content |
|---|---|---|---|
| --- | --- | --- | --- |
| `<prefix>_siblings.csv` | `--sibling-stats` | CSV table | family-size distribution, per genome and global |
| `<prefix>_family_overlap.csv` | `--family-overlap` | CSV matrix | variable families both genomes of a pair carry a call for |
| `<prefix>_entropy.csv` | `--shannon` | CSV table | per-family Shannon entropy, see "Sampling at scale" above |
@@ -263,7 +263,7 @@ CSV matrix layout (`_dist.csv`, `_shared.csv`, `_family_overlap.csv`): header `g
### Sankoff calibration and exports
| File | Written by | Format | Content |
|---|---|---|---|
| --- | --- | --- | --- |
| `<prefix>_sankoff_matrix.csv` | `--sankoff`/`--tnt`/`--phyg`/`--iqtree` | CSV matrix | calibrated 16×16 cost matrix |
| `<prefix>_sankoff_params.yaml` | same flags | YAML | calibration report (raw tallies + derived probabilities) |
| `<prefix>_sankoff.fasta` | same flags | FASTA | Sankoff-recoded pseudo-alignment, header carries an `n_sites` annotation |
@@ -279,7 +279,7 @@ CSV matrix layout (`_dist.csv`, `_shared.csv`, `_family_overlap.csv`): header `g
**`_sankoff_params.yaml`**:
| Key | Meaning |
|---|---|
| --- | --- |
| `ratio_ceiling` | the `--sankoff-ratio-ceiling` value used |
| `cardinality_transitions` | 5×5 list of `{from, to, count, probability}`, family cardinality (0-4 observed forms) |
| `composition_transitions` | 4×4 list of `{from, to, count, probability}`, base letters `A/C/G/T`, single-copy substitutions |
@@ -295,9 +295,8 @@ CSV matrix layout (`_dist.csv`, `_shared.csv`, `_family_overlap.csv`): header `g
**`_iqtree.model`** (`--iqtree`) — lower-triangular exchangeability matrix $R(a,b) = e^{-\text{cost}(a,b)}$ (one row of increasing length per state, whitespace-separated, PAML order), followed by one line of empirical state frequencies. Only states actually occurring in the alignment are kept, compactly renumbered `0..k-1`.
**`_iqtree.fasta`** (`--iqtree`) — alignment recoded to that same compact `0..k-1` alphabet (symbols `0-9A-F`). Under `--free-loss`, non-detection becomes `?` and columns left non-informative once missing calls are ignored are dropped first (required for `+ASC`); with `--iqtree-min-freq` also set (the default), any state rarer than that threshold is folded into the same `?` treatment, and non-informative columns are re-checked and dropped again. Run with:
```
iqtree3 -s <prefix>_iqtree.fasta --seqtype MORPH -m <prefix>_iqtree.model+ASC --prefix <prefix>_iqtree -T AUTO
```
iqtree3 -s <prefix>_iqtree.fasta --seqtype MORPH -m <prefix>_iqtree.model+ASC --prefix <prefix>_iqtree -T AUTO
**`_iqtree_states.csv`** (`--iqtree`) — one row per state actually kept in `_iqtree.model`/`_iqtree.fasta` (header `iqtree_symbol,canonical_symbol,frequency`): `iqtree_symbol` is the compact `0-9A-F` symbol as written in those two files, `canonical_symbol` is the matching `_sankoff_matrix.csv` state, `frequency` is that state's empirical frequency at full precision. Under `--free-loss`, absent (`0`/`?`) is never a kept state, so it never appears here — nor does any state `--iqtree-min-freq` folded away for being too rare.
+3 -5
@@ -5,7 +5,7 @@ Several commands ([`filter`](usage-filter), [`select`](usage-select), [`dump`](u
## Predicate syntax
| Form | Meaning |
|---|---|
| --- | --- |
| `*` or `all` | Matches every genome (case-insensitive) |
| `key=v1\|v2` | Genome's `key` metadata equals one of the listed values |
| `key!=v` | Genome's `key` metadata does not equal `v` |
@@ -20,9 +20,7 @@ Multiple `--ingroup` predicates are combined with AND; multiple `--outgroup` pre
A metadata value is treated as a taxonomy path when it starts with the literal prefix `taxonomy:/`; any other value is treated as a plain string and only supports `=`/`!=`.
```
taxonomy:/segment1@rank1/segment2@rank2/...
```
taxonomy:/segment1@rank1/segment2@rank2/...
Each segment is a name, optionally annotated with a rank (e.g. `@family`, `@genus`, `@species`); ranks are optional and can be mixed within a path. The `@` character is reserved inside taxonomy paths and cannot appear in segment names or rank labels.
@@ -31,7 +29,7 @@ Each segment is a name, optionally annotated with a rank (e.g. `@family`, `@genu
Matching compares segment names only (ranks are informational, not part of the match), with anchoring controlled by leading/trailing `/`:
| Pattern | Matches |
|---|---|
| --- | --- |
| `A/B` | anywhere in the path |
| `/A/B` | at the start of the path (prefix) |
| `A/B$` | at the end of the path (suffix) |
+2 -2
@@ -9,14 +9,14 @@ obikmer query INDEX INPUTS... [OPTIONS]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INDEX` | Index directory to query against |
| `INPUTS...` | Input sequence files (FASTA/FASTQ, gzip optional); at least one required |
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `--detail` | off | Report per-position, per-genome coverage vectors in the output |
| `--count-missing` | off | Also count query kmers absent from the index |
| `--force-presence` | off | Report presence (0/1) per genome instead of raw counts |
+2 -2
@@ -9,13 +9,13 @@ obikmer select SOURCE --output OUTPUT [OPTIONS]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `SOURCE` | Source index directory |
## Options
| Option | Default | Description |
|---|---|---|
| --- | --- | --- |
| `-o, --output` | — | Output index directory (required) |
| `-f, --force` | off | Overwrite an existing output directory |
| `--group NAME:PRED` | none | Define a named group of genomes by predicate (repeatable; mutually exclusive with `--aggregate-by`) |
+3 -3
@@ -9,14 +9,14 @@ obikmer superkmer [OPTIONS] [INPUTS...]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INPUTS...` | Input sequence files or directories (FASTA/FASTQ/GenBank, gzip optional). If omitted, reads from stdin. |
## Options
| Option | Default | Description |
|---|---|---|
| `-k, --kmer-size` | `31` | Kmer size (must be odd, in [11, 31]) |
| --- | --- | --- |
| `-k, --kmer-size` | `31` | Kmer size (must be odd, in \[11, 31\]) |
| `-m, --minimizer-size` | `11` | Minimizer size (must be odd, in $[3, k-1]$) |
| `--theta` | `0.7` | Entropy threshold; kmers with a normalized entropy at or below this value are excluded |
| `--level-max` | `6` | Maximum sub-word size used for the entropy score |
+1 -1
@@ -9,7 +9,7 @@ obikmer unitig INDEX [OPTIONS]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INDEX` | Index directory |
## Options
+2 -2
@@ -9,13 +9,13 @@ obikmer utils INDEXES... [OPTIONS]
## Arguments
| Argument | Description |
|---|---|
| --- | --- |
| `INDEXES...` | One or more index directories |
## Options
| Option | Scope | Description |
|---|---|---|
| --- | --- | --- |
| `--new-label NEW=OLD` | single index only | Rename a genome label |
| `--upgrade-index` | single index only | Add any missing layer metadata files to an older index |
| `--bits-per-kmer` | single index only | Print bits-per-kmer statistics |