Update indexing constraints and CLI options

This commit introduces new constraints for kmer size and minimizer selection, defines super-kmers, and adds extensive new command-line options for filtering, conversion, merging, and packing indices. Documentation across the codebase has also been updated.
Eric Coissac committed 2026-09-12 14:00:23 +02:00
1 parent c1e139c597
commit 5313788f7c
22 files changed
+100 -107

No files matched your search

+3 -3
@@ -15,7 +15,7 @@ All functionality is exposed through a single binary, `obikmer`, organized as su
## Commands ## Commands
| Command | Purpose | | Command | Purpose |
|---|---| | --- | --- |
| [`superkmer`](usage-superkmer) | Extract super-kmers from a sequence file and write them to stdout | | [`superkmer`](usage-superkmer) | Extract super-kmers from a sequence file and write them to stdout |
| [`index`](usage-index_command) | Build a genome index | | [`index`](usage-index_command) | Build a genome index |
| [`merge`](usage-merge) | Merge multiple indexes into one | | [`merge`](usage-merge) | Merge multiple indexes into one |
@@ -48,7 +48,7 @@ See [Genome predicates and taxonomy paths](usage-predicates) for the selection l
These constraints are checked at startup; an invalid value exits immediately with an error. These constraints are checked at startup; an invalid value exits immediately with an error.
| Parameter | Constraint | Reason | | Parameter | Constraint | Reason |
|---|---|---| | --- | --- | --- |
| $k$ (`--kmer-size`) | odd, $k \in [11, 31]$ | odd length guarantees the canonical form is always well defined; the range keeps a kmer within a 64-bit word while retaining specificity | | $k$ (`--kmer-size`) | odd, $k \in [11, 31]$ | odd length guarantees the canonical form is always well defined; the range keeps a kmer within a 64-bit word while retaining specificity |
| $m$ (`--minimizer-size`) | odd, $3 \le m \le k-1$ | same palindrome argument as $k$; must be strictly shorter than the kmer | | $m$ (`--minimizer-size`) | odd, $3 \le m \le k-1$ | same palindrome argument as $k$; must be strictly shorter than the kmer |
| $z$ (`-z`, approximate evidence only) | $z \le k-1$ | the effective indexed kmer size is $k-z+1$ | | $z$ (`-z`, approximate evidence only) | $z \le k-1$ | the effective indexed kmer size is $k-z+1$ |
@@ -58,7 +58,7 @@ These constraints are checked at startup; an invalid value exits immediately wit
Genome labels are arbitrary Unicode strings, with the following restrictions: Genome labels are arbitrary Unicode strings, with the following restrictions:
| Character | Forbidden | Reason | | Character | Forbidden | Reason |
|---|---|---| | --- | --- | --- |
| `/` | yes | filesystem path separator | | `/` | yes | filesystem path separator |
| `=` | yes | separator used by `--new-label` | | `=` | yes | separator used by `--new-label` |
| `\0` | yes | null byte | | `\0` | yes | null byte |
+18 -20
@@ -26,26 +26,24 @@ Two verification modes are available, selected at build time (`index --approx`)
## On-disk layout ## On-disk layout
``` <index_root>/
<index_root>/ index.meta global configuration (k, minimizer size, partition count,
index.meta global configuration (k, minimizer size, partition count, evidence mode, whether counts are stored) and genome list/metadata
evidence mode, whether counts are stored) and genome list/metadata scatter.done / count.done / index.done build-progress sentinels
scatter.done / count.done / index.done build-progress sentinels spectrums/<label>.json per-genome kmer frequency histogram
spectrums/<label>.json per-genome kmer frequency histogram partitions/
partitions/ part_00000/ ... part_NNNNN/
part_00000/ ... part_NNNNN/ index/
index/ meta.json number of layers in this partition
meta.json number of layers in this partition layer_0/
layer_0/ unitigs.bin reconstructible kmer sequence data — always kept
unitigs.bin reconstructible kmer sequence data — always kept unitigs.bin.idx random-access index into unitigs.bin (exact evidence only)
unitigs.bin.idx random-access index into unitigs.bin (exact evidence only) mphf.bin the minimal perfect hash function
mphf.bin the minimal perfect hash function evidence.bin exact evidence (exact mode only)
evidence.bin exact evidence (exact mode only) fingerprint.bin approximate evidence (approximate mode only)
fingerprint.bin approximate evidence (approximate mode only) counts/ per-genome kmer counts (if counts were requested)
counts/ per-genome kmer counts (if counts were requested) presence/ per-genome presence/absence bits
presence/ per-genome presence/absence bits layer_1/, layer_2/, ... added by later merges, same internal structure
layer_1/, layer_2/, ... added by later merges, same internal structure
```
`unitigs.bin` is the only file from which the indexed kmer content can be fully recovered; it is always retained. Every other file (MPHF, evidence, counts) is derived from it. `unitigs.bin` is the only file from which the indexed kmer content can be fully recovered; it is always retained. Every other file (MPHF, evidence, counts) is derived from it.
+5 -5
@@ -5,11 +5,11 @@
Every nucleotide is encoded on 2 bits, most-significant-bit first within each word: Every nucleotide is encoded on 2 bits, most-significant-bit first within each word:
| Base | Encoding | | Base | Encoding |
|------|----------| | --- | --- |
| A | `00` | | A | `00` |
| C | `01` | | C | `01` |
| G | `10` | | G | `10` |
| T | `11` | | T | `11` |
The Watson-Crick complement of a base is its bitwise NOT on 2 bits: $\text{complement}(base) = \lnot base \mathbin{\&} \texttt{0b11}$. The Watson-Crick complement of a base is its bitwise NOT on 2 bits: $\text{complement}(base) = \lnot base \mathbin{\&} \texttt{0b11}$.
+1 -1
@@ -22,7 +22,7 @@ A value near 0 indicates low complexity (e.g. a homopolymer run); near 1 indicat
## Final score ## Final score
The filter evaluates $\hat{H}(ws)$ for every word size from 1 to ws_max and keeps the minimum: The filter evaluates $\hat{H}(ws)$ for every word size from 1 to ws\_max and keeps the minimum:
$$\text{entropy}(kmer) = \min_{ws=1}^{ws_{\max}} \hat{H}(ws)$$ $$\text{entropy}(kmer) = \min_{ws=1}^{ws_{\max}} \hat{H}(ws)$$
+6 -8
@@ -6,9 +6,7 @@ An index is split into a fixed number of **partitions**, each handling an indepe
The canonical minimizer of a super-kmer (see [Minimizer selection](theory-minimizer_selection)) is hashed to produce a $p$-bit routing value that selects the destination partition: The canonical minimizer of a super-kmer (see [Minimizer selection](theory-minimizer_selection)) is hashed to produce a $p$-bit routing value that selects the destination partition:
``` canonical minimizer → hash(minimizer) → p-bit value → partition index
canonical minimizer → hash(minimizer) → p-bit value → partition index
```
Within a partition, kmers are indexed as plain values via a minimal perfect hash function (see [On-disk storage](formats-index_layout)); the minimizer plays no further role once a super-kmer has reached its partition. Within a partition, kmers are indexed as plain values via a minimal perfect hash function (see [On-disk storage](formats-index_layout)); the minimizer plays no further role once a super-kmer has reached its partition.
@@ -17,10 +15,10 @@ Within a partition, kmers are indexed as plain values via a minimal perfect hash
Even though $H$ already makes minimizer values well-distributed (see [Minimizer selection](theory-minimizer_selection)), choosing $p$ well below the number of bits available in the minimizer ($2m$) leaves a comfortable entropy margin, provided the number of distinct minimizers actually observed is much larger than the number of partitions. Even though $H$ already makes minimizer values well-distributed (see [Minimizer selection](theory-minimizer_selection)), choosing $p$ well below the number of bits available in the minimizer ($2m$) leaves a comfortable entropy margin, provided the number of distinct minimizers actually observed is much larger than the number of partitions.
| Minimizer size $m$ | Minimizer bits ($2m$) | Typical partition-index bits $p$ | Partitions | | Minimizer size $m$ | Minimizer bits ($2m$) | Typical partition-index bits $p$ | Partitions |
|----|-----------|-----------|------------| | --- | --- | --- | --- |
| 9 | 18 | 6–8 | 64–256 | | 9 | 18 | 6–8 | 64–256 |
| 11 | 22 | 8–10 | 256–1 024 | | 11 | 22 | 8–10 | 256–1 024 |
| 13 | 26 | 10–12 | 1 024–4 096| | 13 | 26 | 10–12 | 1 024–4 096 |
| 15 | 30 | 10–14 | 1 024–16 384| | 15 | 30 | 10–14 | 1 024–16 384 |
The number of partitions must satisfy $p \le 2m$, and in practice $p$ is chosen well below that bound to leave a comfortable entropy margin. For $k=31$, $m=13$, $p=10$ (1024 partitions), partition load is well balanced on real genomic data. The number of partitions must satisfy $p \le 2m$, and in practice $p$ is chosen well below that bound to leave a comfortable entropy margin. For $k=31$, $m=13$, $p=10$ (1024 partitions), partition load is well balanced on real genomic data.
+4 -4
@@ -11,7 +11,7 @@ A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary. A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately ([Golan & Shur 2025](#ref-Golan2025-xf); [Zheng *et al.* 2020](#ref-Zheng2020-ji)): For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately ([Golan & Shur 2025](#ref-Golan2025-xf); [Zheng et al. 2020](#ref-Zheng2020-ji)):
$$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$ $$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$
@@ -25,17 +25,17 @@ Super-kmers are the unit of work used throughout construction and querying: sequ
## Bibliography ## Bibliography
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0"> <div id="refs" class="references csl-bib-body hanging-indent">
<div id="ref-Golan2025-xf" class="csl-entry"> <div id="ref-Golan2025-xf" class="csl-entry">
Golan, S. & Shur, A.M. (2025). [Expected density of random minimizers](https://doi.org/10.1007/978-3-031-82670-2\_25). In: *Lecture notes in computer science*, Lecture notes in computer science. Springer Nature Switzerland, Cham, pp. 347–360. Golan, S. & Shur, A.M. (2025). <a href="https://doi.org/10.1007/978-3-031-82670-2\_25">Expected density of random minimizers</a>. In: <span style="font-style: italic;">Lecture Notes in Computer Science</span>, Lecture Notes in Computer Science. Springer Nature Switzerland, Cham, pp. 347–360.
</div> </div>
<div id="ref-Zheng2020-ji" class="csl-entry"> <div id="ref-Zheng2020-ji" class="csl-entry">
Zheng, H., Kingsford, C. & Marçais, G. (2020). [Improved design and analysis of practical minimizers](https://doi.org/10.1093/bioinformatics/btaa472). *Bioinformatics (Oxford, England)*, 36, i119–i127. Zheng, H., Kingsford, C. & Marçais, G. (2020). <a href="https://doi.org/10.1093/bioinformatics/btaa472">Improved design and analysis of practical minimizers</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 36, i119–i127.
</div> </div>
+17 -17
@@ -8,7 +8,7 @@ The minimizer partitions a sequence into super-kmers: maximal runs of overlappin
## Hash-based ("random") minimizer ## Hash-based ("random") minimizer
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions ([Golan & Shur 2025](#ref-Golan2025-xf); [Kille *et al.* 2023](#ref-Kille2023-px); [Pan & Reinert 2024](#ref-Pan2024-hb); [Zheng *et al.* 2020](#ref-Zheng2020-ji), [2021](#ref-Zheng2021-cc)). `obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions ([Golan & Shur 2025](#ref-Golan2025-xf); [Kille et al. 2023](#ref-Kille2023-px); [Pan & Reinert 2024](#ref-Pan2024-hb); [Zheng et al. 2020](#ref-Zheng2020-ji); [2021](#ref-Zheng2021-cc)).
Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition. Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition.
@@ -24,7 +24,7 @@ The hash function is a 64-bit mixing function (splitmix64-style finalizer) appli
$$H(x) = \text{mix64}(x \oplus s)$$ $$H(x) = \text{mix64}(x \oplus s)$$
``` text ```text
H(x): H(x):
x ← x ⊕ s x ← x ⊕ s
x ← x ⊕ (x >> 30) x ← x ⊕ (x >> 30)
@@ -36,15 +36,15 @@ H(x):
The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes; if one of them happened to be the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$, since $\text{mix64}(0) = 0$ is a fixed point — it would win far more windows than the composition-uniform behavior established above predicts, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. Exhaustive checks confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$: The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes; if one of them happened to be the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$, since $\text{mix64}(0) = 0$ is a fixed point — it would win far more windows than the composition-uniform behavior established above predicts, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. Exhaustive checks confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$:
| $m$ | argmin (canonical) | decoded sequence | minimal period | | $m$ | argmin (canonical) | decoded sequence | minimal period |
|-----|--------------------|-------------------|----------------| | --- | --- | --- | --- |
| 3 | 16 | `CAA` | 3 | | 3 | 16 | `CAA` | 3 |
| 5 | 78 | `ACATG` | 5 | | 5 | 78 | `ACATG` | 5 |
| 7 | 5512 | `CCCGAGA` | 7 | | 7 | 5512 | `CCCGAGA` | 7 |
| 9 | 108760 | `CGGGATCGA` | 9 | | 9 | 108760 | `CGGGATCGA` | 9 |
| 11 | 179014 | `AAGGTGTCACG` | 11 | | 11 | 179014 | `AAGGTGTCACG` | 11 |
| 13 | 33759044 | `GAAATACTTCACA` | 13 | | 13 | 33759044 | `GAAATACTTCACA` | 13 |
| 15 | 29869313 | `AACTACTTACCAAAC` | 15 | | 15 | 29869313 | `AACTACTTACCAAAC` | 15 |
If the minimum period length is found to be $m$ as observed in the above table, then the sequence is aperiodic, as no shorter period can be identified. If the minimum period length is found to be $m$ as observed in the above table, then the sequence is aperiodic, as no shorter period can be identified.
@@ -58,35 +58,35 @@ See [Partitioning and indexing architecture](theory-indexing_architecture) for m
## Bibliography ## Bibliography
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0"> <div id="refs" class="references csl-bib-body hanging-indent">
<div id="ref-Golan2025-xf" class="csl-entry"> <div id="ref-Golan2025-xf" class="csl-entry">
Golan, S. & Shur, A.M. (2025). [Expected density of random minimizers](https://doi.org/10.1007/978-3-031-82670-2\_25). In: *Lecture notes in computer science*, Lecture notes in computer science. Springer Nature Switzerland, Cham, pp. 347–360. Golan, S. & Shur, A.M. (2025). <a href="https://doi.org/10.1007/978-3-031-82670-2\_25">Expected density of random minimizers</a>. In: <span style="font-style: italic;">Lecture Notes in Computer Science</span>, Lecture Notes in Computer Science. Springer Nature Switzerland, Cham, pp. 347–360.
</div> </div>
<div id="ref-Kille2023-px" class="csl-entry"> <div id="ref-Kille2023-px" class="csl-entry">
Kille, B., Garrison, E., Treangen, T.J. & Phillippy, A.M. (2023). [Minmers are a generalization of minimizers that enable unbiased local jaccard estimation](https://doi.org/10.1093/bioinformatics/btad512). *Bioinformatics (Oxford, England)*, 39. Kille, B., Garrison, E., Treangen, T.J. & Phillippy, A.M. (2023). <a href="https://doi.org/10.1093/bioinformatics/btad512">Minmers are a generalization of minimizers that enable unbiased local Jaccard estimation</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 39.
</div> </div>
<div id="ref-Pan2024-hb" class="csl-entry"> <div id="ref-Pan2024-hb" class="csl-entry">
Pan, C. & Reinert, K. (2024). [A simple refined DNA minimizer operator enables 2-fold faster computation](https://doi.org/10.1093/bioinformatics/btae045). *Bioinformatics (Oxford, England)*, 40. Pan, C. & Reinert, K. (2024). <a href="https://doi.org/10.1093/bioinformatics/btae045">A simple refined DNA minimizer operator enables 2-fold faster computation</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 40.
</div> </div>
<div id="ref-Zheng2020-ji" class="csl-entry"> <div id="ref-Zheng2020-ji" class="csl-entry">
Zheng, H., Kingsford, C. & Marçais, G. (2020). [Improved design and analysis of practical minimizers](https://doi.org/10.1093/bioinformatics/btaa472). *Bioinformatics (Oxford, England)*, 36, i119–i127. Zheng, H., Kingsford, C. & Marçais, G. (2020). <a href="https://doi.org/10.1093/bioinformatics/btaa472">Improved design and analysis of practical minimizers</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 36, i119–i127.
</div> </div>
<div id="ref-Zheng2021-cc" class="csl-entry"> <div id="ref-Zheng2021-cc" class="csl-entry">
Zheng, H., Kingsford, C. & Marçais, G. (2021). [Sequence-specific minimizers via polar sets](https://doi.org/10.1093/bioinformatics/btab313). *Bioinformatics (Oxford, England)*, 37, i187–i195. Zheng, H., Kingsford, C. & Marçais, G. (2021). <a href="https://doi.org/10.1093/bioinformatics/btab313">Sequence-specific minimizers via polar sets</a>. <span style="font-style: italic;">Bioinformatics (Oxford, England)</span>, 37, i187–i195.
</div> </div>
+2 -2
@@ -10,13 +10,13 @@ obikmer annotate INDEX --dump
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INDEX` | Index directory to annotate (modified in place) | | `INDEX` | Index directory to annotate (modified in place) |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--csv` | — | CSV file of metadata to apply (must contain an id column); required unless `--dump` is used | | `--csv` | — | CSV file of metadata to apply (must contain an id column); required unless `--dump` is used |
| `--sep` | `,` | CSV field separator | | `--sep` | `,` | CSV field separator |
| `--id-col` | `id` | Name of the column containing genome labels | | `--id-col` | `id` | Name of the column containing genome labels |
+2 -2
@@ -9,7 +9,7 @@ obikmer convert INDEX (--exact-evidence | --approx-evidence BITS | --hybrid-evid
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INDEX` | Index directory to convert (modified in place) | | `INDEX` | Index directory to convert (modified in place) |
## Options ## Options
@@ -17,7 +17,7 @@ obikmer convert INDEX (--exact-evidence | --approx-evidence BITS | --hybrid-evid
Exactly one of the first three is required: Exactly one of the first three is required:
| Option | Description | | Option | Description |
|---|---| | --- | --- |
| `--exact-evidence` | Convert to exact evidence (zero false positives) | | `--exact-evidence` | Convert to exact evidence (zero false positives) |
| `--approx-evidence BITS` | Convert to approximate (fingerprint-only) evidence; `BITS` = fingerprint bits per slot (b) | | `--approx-evidence BITS` | Convert to approximate (fingerprint-only) evidence; `BITS` = fingerprint bits per slot (b) |
| `--hybrid-evidence` | Convert to hybrid evidence (both exact and approximate bundles kept) | | `--hybrid-evidence` | Convert to hybrid evidence (both exact and approximate bundles kept) |
+2 -2
@@ -9,13 +9,13 @@ obikmer dump INDEX [OPTIONS]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INDEX` | Index directory to dump | | `INDEX` | Index directory to dump |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--force-presence` | off | Output presence/absence (0/1) even if the index stores counts | | `--force-presence` | off | Output presence/absence (0/1) even if the index stores counts |
| `--debug` | off | Prefix each row with the partition and layer columns | | `--debug` | off | Prefix each row with the partition and layer columns |
| `--head N` | none | Limit output to the first N kmers | | `--head N` | none | Limit output to the first N kmers |
+1 -1
@@ -9,7 +9,7 @@ obikmer estimate [OPTIONS]
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `-k, --kmer-size` | `31` | Kmer size used at query time (matches `index`'s `--kmer-size`) | | `-k, --kmer-size` | `31` | Kmer size used at query time (matches `index`'s `--kmer-size`) |
| `-z, --findere-z` | none | Findere z parameter | | `-z, --findere-z` | none | Findere z parameter |
| `--evidence-bits` | none | Fingerprint bits per slot (b) | | `--evidence-bits` | none | Fingerprint bits per slot (b) |
+3 -3
@@ -9,13 +9,13 @@ obikmer filter SOURCE -o OUTPUT [OPTIONS]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `SOURCE` | Source index directory | | `SOURCE` | Source index directory |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `-o, --output` | — (required) | Output index directory | | `-o, --output` | — (required) | Output index directory |
| `-f, --force` | off | Overwrite an existing output directory | | `-f, --force` | off | Overwrite an existing output directory |
| `--presence` | off | Output presence/absence instead of counts | | `--presence` | off | Output presence/absence instead of counts |
@@ -27,7 +27,7 @@ obikmer filter SOURCE -o OUTPUT [OPTIONS]
## Predicate options ## Predicate options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--ingroup` | none | Ingroup predicate (repeatable; each occurrence is ANDed) | | `--ingroup` | none | Ingroup predicate (repeatable; each occurrence is ANDed) |
| `--outgroup` | none | Outgroup predicate (repeatable; each occurrence is ORed) | | `--outgroup` | none | Outgroup predicate (repeatable; each occurrence is ORed) |
| `--min-count` | 0, or group size + N if negative | Minimum number of ingroup genomes carrying the kmer | | `--min-count` | 0, or group size + N if negative | Minimum number of ingroup genomes carrying the kmer |
+3 -3
@@ -9,18 +9,18 @@ obikmer index -o OUTPUT [OPTIONS] [INPUTS...]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INPUTS...` | Input sequence files or directories (FASTA/FASTQ/GenBank, gzip optional). If omitted, reads from stdin. | | `INPUTS...` | Input sequence files or directories (FASTA/FASTQ/GenBank, gzip optional). If omitted, reads from stdin. |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `-o, --output` | — (required) | Output index directory | | `-o, --output` | — (required) | Output index directory |
| `--force` | off | Overwrite an existing output directory | | `--force` | off | Overwrite an existing output directory |
| `--label` | input file name without extension | Genome label stored in the index | | `--label` | input file name without extension | Genome label stored in the index |
| `--meta KEY=VALUE` | none | Attach a categorical metadata field to the genome (repeatable) | | `--meta KEY=VALUE` | none | Attach a categorical metadata field to the genome (repeatable) |
| `-k, --kmer-size` | `31` | Kmer size (odd, in [11, 31]) | | `-k, --kmer-size` | `31` | Kmer size (odd, in \[11, 31\]) |
| `-m, --minimizer-size` | `11` | Minimizer size (odd, in $[3, k-1]$) | | `-m, --minimizer-size` | `11` | Minimizer size (odd, in $[3, k-1]$) |
| `--theta` | `0.7` | Entropy threshold for the low-complexity filter | | `--theta` | `0.7` | Entropy threshold for the low-complexity filter |
| `--level-max` | `6` | Maximum sub-word size for the entropy score | | `--level-max` | `6` | Maximum sub-word size for the entropy score |
+2 -2
@@ -9,13 +9,13 @@ obikmer merge -o OUTPUT SOURCE... [OPTIONS]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `SOURCE...` | Index directories to merge (at least one required) | | `SOURCE...` | Index directories to merge (at least one required) |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `-o, --output` | — (required) | Output index directory | | `-o, --output` | — (required) | Output index directory |
| `--force` | off | Overwrite an existing output directory | | `--force` | off | Overwrite an existing output directory |
| `--force-presence` | off | Store the merged index as presence/absence even if all sources have counts | | `--force-presence` | off | Store the merged index as presence/absence even if all sources have counts |
+2 -2
@@ -9,13 +9,13 @@ obikmer pack INDEX [--sparse]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INDEX` | Index directory to pack (modified in place) | | `INDEX` | Index directory to pack (modified in place) |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--sparse` | off | Pack presence/absence and count matrices into a sparse, deduplicated format instead of the dense one | | `--sparse` | off | Pack presence/absence and count matrices into a sparse, deduplicated format instead of the dense one |
The index directory is locked for exclusive access while packing. The index directory is locked for exclusive access while packing.
+16 -17
@@ -9,13 +9,13 @@ obikmer phylo INDEX [OPTIONS]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INDEX` | Index directory | | `INDEX` | Index directory |
## Distance matrix (`--distance`) ## Distance matrix (`--distance`)
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--distance` | `jaccard` | See the two tables below for the full list of accepted values | | `--distance` | `jaccard` | See the two tables below for the full list of accepted values |
| `--gamma-shape ALPHA\|auto` | none | Rate-heterogeneity correction, for `snp-*` values that support it (see below). Either a fixed $\alpha$ or `auto`/`estimate` to fit it from the data (see "Automatic $\alpha$ estimation" below). No effect on the other values; rejected if given together with a value that doesn't support it | | `--gamma-shape ALPHA\|auto` | none | Rate-heterogeneity correction, for `snp-*` values that support it (see below). Either a fixed $\alpha$ or `auto`/`estimate` to fit it from the data (see "Automatic $\alpha$ estimation" below). No effect on the other values; rejected if given together with a value that doesn't support it |
| `--presence-threshold` | `1` | Minimum count for a kmer to be considered present, for `jaccard`/`mash` on a count index | | `--presence-threshold` | `1` | Minimum count for a kmer to be considered present, for `jaccard`/`mash` on a count index |
@@ -30,7 +30,7 @@ Every value routes to one of two independent computations:
### Whole-index metrics ### Whole-index metrics
| Value | Definition | | Value | Definition |
|---|---| | --- | --- |
| `jaccard` | $D = 1 - \dfrac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}$ over the sets of kmers present in each genome | | `jaccard` | $D = 1 - \dfrac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}$ over the sets of kmers present in each genome |
| `mash` | derived from the Jaccard distance via $D = -\dfrac{1}{k} \ln\!\left(\dfrac{2J}{1+J}\right)$ where $J = 1 - D_{\text{jaccard}}$ and $k$ is the index's kmer size; clamped to 1.0 when $J \le 0$ | | `mash` | derived from the Jaccard distance via $D = -\dfrac{1}{k} \ln\!\left(\dfrac{2J}{1+J}\right)$ where $J = 1 - D_{\text{jaccard}}$ and $k$ is the index's kmer size; clamped to 1.0 when $J \le 0$ |
| `hamming` | number of kmer positions where presence differs between the two genomes (presence index only, not normalized): $D = \sum_i \mathbb{1}[a_i \ne b_i]$ | | `hamming` | number of kmer positions where presence differs between the two genomes (presence index only, not normalized): $D = \sum_i \mathbb{1}[a_i \ne b_i]$ |
@@ -112,7 +112,7 @@ Without `-o`, the matrix goes to stdout in relaxed-PHYLIP format (`n` on the fir
## `--exclude-genome`, `--min-shared-family` ## `--exclude-genome`, `--min-shared-family`
| Option | Description | | Option | Description |
|---|---| | --- | --- |
| `--exclude-genome LABEL` | Exclude a genome (repeatable). Drops its row/column from the distance/shared-kmer matrix output, and removes it from the sampling used by `--pseudo-alignment`/`--sankoff`/a `snp-*` `--distance` value. Does not change the value computed for any remaining pair | | `--exclude-genome LABEL` | Exclude a genome (repeatable). Drops its row/column from the distance/shared-kmer matrix output, and removes it from the sampling used by `--pseudo-alignment`/`--sankoff`/a `snp-*` `--distance` value. Does not change the value computed for any remaining pair |
| `--min-shared-family N` | Auto-exclude, on top of `--exclude-genome`, any genome whose mean shared-family count against every other genome (see "Family Overlap" below) falls below `N`. Applies only to `--pseudo-alignment`/`--sankoff`/`snp-*` `--distance` — never to the whole-index metrics or their matrix/NJ/UPGMA output | | `--min-shared-family N` | Auto-exclude, on top of `--exclude-genome`, any genome whose mean shared-family count against every other genome (see "Family Overlap" below) falls below `N`. Applies only to `--pseudo-alignment`/`--sankoff`/`snp-*` `--distance` — never to the whole-index metrics or their matrix/NJ/UPGMA output |
@@ -123,7 +123,7 @@ Neighbor-Joining and UPGMA trees (`--nj`/`--upgma`) are always built from every
Requires the sibling annex, built once per index: Requires the sibling annex, built once per index:
| Option | Description | | Option | Description |
|---|---| | --- | --- |
| `--sibling-annex` | Build (or rebuild) the sibling-count/minorant annex — prerequisite for every option in this section, and for a `snp-*` `--distance` value | | `--sibling-annex` | Build (or rebuild) the sibling-count/minorant annex — prerequisite for every option in this section, and for a `snp-*` `--distance` value |
| `--sibling-stats` | Write `<prefix>_siblings.csv`: the family-size distribution, per genome and globally | | `--sibling-stats` | Write `<prefix>_siblings.csv`: the family-size distribution, per genome and globally |
| `--sibling-hist` | Print the global family-size histogram (1-4 members) only | | `--sibling-hist` | Print the global family-size histogram (1-4 members) only |
@@ -138,7 +138,7 @@ A family is eligible for a genome pair $(i,j)$ only if both genomes carry exactl
`<prefix>_siblings.csv` — family size = number of distinct central bases observed at a family (1-4). `<prefix>_siblings.csv` — family size = number of distinct central bases observed at a family (1-4).
| Column | Meaning | | Column | Meaning |
|---|---| | --- | --- |
| `genome` | genome label, or the literal `global` for the last row | | `genome` | genome label, or the literal `global` for the last row |
| `1`, `2`, `3`, `4` | for a genome row: number of families of that size where the genome carries ≥ 1 member. For the `global` row: the actual deduplicated family-size histogram — not the sum of the rows above | | `1`, `2`, `3`, `4` | for a genome row: number of families of that size where the genome carries ≥ 1 member. For the `global` row: the actual deduplicated family-size histogram — not the sum of the rows above |
@@ -153,7 +153,7 @@ A family is eligible for a genome pair $(i,j)$ only if both genomes carry exactl
`<prefix>_alignment.fasta` — one record per non-excluded genome, one column per variable family (family size ≥ 2). Each site is IUPAC-coded from the genome's presence mask at that family: a single observed form → the plain base; several forms → the matching IUPAC ambiguity code; no form → `-`. `<prefix>_alignment.fasta` — one record per non-excluded genome, one column per variable family (family size ≥ 2). Each site is IUPAC-coded from the genome's presence mask at that family: a single observed form → the plain base; several forms → the matching IUPAC ambiguity code; no form → `-`.
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--subsample N` | none (mandatory here) | Target number of families to sample | | `--subsample N` | none (mandatory here) | Target number of families to sample |
| `--free-loss` | off | Treat a genome carrying none of a family's observed members as missing data (`?`) instead of `-` | | `--free-loss` | off | Treat a genome carrying none of a family's observed members as missing data (`?`) instead of `-` |
| `--no-ambiguity` | off | Treat a genome carrying more than one member of a family as missing data (`?`) instead of an IUPAC ambiguity code | | `--no-ambiguity` | off | Treat a genome carrying more than one member of a family as missing data (`?`) instead of an IUPAC ambiguity code |
@@ -171,7 +171,7 @@ Every invocation, with or without `--session`, draws its own fresh random sample
### `--session`: reusing a sample across separate commands ### `--session`: reusing a sample across separate commands
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--session DIR` | none | Persist the sample (and, for `--sankoff`, its calibration/alignment) in `DIR` so a later, separate `obikmer phylo` invocation with the exact same selection parameters restores it instead of resampling | | `--session DIR` | none | Persist the sample (and, for `--sankoff`, its calibration/alignment) in `DIR` so a later, separate `obikmer phylo` invocation with the exact same selection parameters restores it instead of resampling |
| `--session-force` | off | With `--session DIR`: overwrite its saved parameters and cached sample instead of erroring out when this run's parameters don't match. No effect without `--session` | | `--session-force` | off | With `--session DIR`: overwrite its saved parameters and cached sample instead of erroring out when this run's parameters don't match. No effect without `--session` |
@@ -188,7 +188,7 @@ Without `--subsample`, every variable family (family size ≥ 2) is used. With `
`<prefix>_entropy.csv` has one row per family visited: `<prefix>_entropy.csv` has one row per family visited:
| Column | Meaning | | Column | Meaning |
|---|---| | --- | --- |
| `layer` | an internal index-layer identifier — stable within one run, not meaningful across indexes | | `layer` | an internal index-layer identifier — stable within one run, not meaningful across indexes |
| `family_idx` | the family's position within that layer | | `family_idx` | the family's position within that layer |
| `entropy15` | Shannon entropy (bits) over the 16 possible states (the 15 non-empty subsets of `{A,C,G,T}`), genomes absent from the family excluded from the count | | `entropy15` | Shannon entropy (bits) over the 16 possible states (the 15 non-empty subsets of `{A,C,G,T}`), genomes absent from the family excluded from the count |
@@ -207,7 +207,7 @@ The first `phylo` run on a given index that uses `--entropy`/`--entropy-sd` pays
## Sankoff calibration and phylogenetic exports ## Sankoff calibration and phylogenetic exports
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--sankoff` | off | Calibrate a 16-state parsimony cost matrix and matching pseudo-alignment. Requires `--subsample N` | | `--sankoff` | off | Calibrate a 16-state parsimony cost matrix and matching pseudo-alignment. Requires `--subsample N` |
| `--sankoff-ratio-ceiling` | `0.5` | Exclude genome pairs whose raw SNP ratio exceeds this value from the base-composition part of the calibration | | `--sankoff-ratio-ceiling` | `0.5` | Exclude genome pairs whose raw SNP ratio exceeds this value from the base-composition part of the calibration |
| `--free-loss` | off | Recode a family's non-detection as the `?` missing-data symbol instead of an ordinary, costed state, throughout `--sankoff` and every export built from it | | `--free-loss` | off | Recode a family's non-detection as the `?` missing-data symbol instead of an ordinary, costed state, throughout `--sankoff` and every export built from it |
@@ -242,7 +242,7 @@ With `-o/--output PREFIX`, the relevant subset of the files below is written. Wi
### Distance matrix ### Distance matrix
| File | Written by | Format | Content | | File | Written by | Format | Content |
|---|---|---|---| | --- | --- | --- | --- |
| `<prefix>_dist.phy` | always, unless `--csv` | relaxed PHYLIP | the `--distance` matrix | | `<prefix>_dist.phy` | always, unless `--csv` | relaxed PHYLIP | the `--distance` matrix |
| `<prefix>_dist.csv` | `--csv` | CSV matrix | the `--distance` matrix, 6 decimals | | `<prefix>_dist.csv` | `--csv` | CSV matrix | the `--distance` matrix, 6 decimals |
| `<prefix>_shared.csv` | `--shared-kmers` | CSV matrix | shared-kmer count per genome pair (integers) | | `<prefix>_shared.csv` | `--shared-kmers` | CSV matrix | shared-kmer count per genome pair (integers) |
@@ -254,7 +254,7 @@ CSV matrix layout (`_dist.csv`, `_shared.csv`, `_family_overlap.csv`): header `g
### Central-position SNP model ### Central-position SNP model
| File | Written by | Format | Content | | File | Written by | Format | Content |
|---|---|---|---| | --- | --- | --- | --- |
| `<prefix>_siblings.csv` | `--sibling-stats` | CSV table | family-size distribution, per genome and global | | `<prefix>_siblings.csv` | `--sibling-stats` | CSV table | family-size distribution, per genome and global |
| `<prefix>_family_overlap.csv` | `--family-overlap` | CSV matrix | variable families both genomes of a pair carry a call for | | `<prefix>_family_overlap.csv` | `--family-overlap` | CSV matrix | variable families both genomes of a pair carry a call for |
| `<prefix>_entropy.csv` | `--shannon` | CSV table | per-family Shannon entropy, see "Sampling at scale" above | | `<prefix>_entropy.csv` | `--shannon` | CSV table | per-family Shannon entropy, see "Sampling at scale" above |
@@ -263,7 +263,7 @@ CSV matrix layout (`_dist.csv`, `_shared.csv`, `_family_overlap.csv`): header `g
### Sankoff calibration and exports ### Sankoff calibration and exports
| File | Written by | Format | Content | | File | Written by | Format | Content |
|---|---|---|---| | --- | --- | --- | --- |
| `<prefix>_sankoff_matrix.csv` | `--sankoff`/`--tnt`/`--phyg`/`--iqtree` | CSV matrix | calibrated 16×16 cost matrix | | `<prefix>_sankoff_matrix.csv` | `--sankoff`/`--tnt`/`--phyg`/`--iqtree` | CSV matrix | calibrated 16×16 cost matrix |
| `<prefix>_sankoff_params.yaml` | same flags | YAML | calibration report (raw tallies + derived probabilities) | | `<prefix>_sankoff_params.yaml` | same flags | YAML | calibration report (raw tallies + derived probabilities) |
| `<prefix>_sankoff.fasta` | same flags | FASTA | Sankoff-recoded pseudo-alignment, header carries an `n_sites` annotation | | `<prefix>_sankoff.fasta` | same flags | FASTA | Sankoff-recoded pseudo-alignment, header carries an `n_sites` annotation |
@@ -279,7 +279,7 @@ CSV matrix layout (`_dist.csv`, `_shared.csv`, `_family_overlap.csv`): header `g
**`_sankoff_params.yaml`**: **`_sankoff_params.yaml`**:
| Key | Meaning | | Key | Meaning |
|---|---| | --- | --- |
| `ratio_ceiling` | the `--sankoff-ratio-ceiling` value used | | `ratio_ceiling` | the `--sankoff-ratio-ceiling` value used |
| `cardinality_transitions` | 5×5 list of `{from, to, count, probability}`, family cardinality (0-4 observed forms) | | `cardinality_transitions` | 5×5 list of `{from, to, count, probability}`, family cardinality (0-4 observed forms) |
| `composition_transitions` | 4×4 list of `{from, to, count, probability}`, base letters `A/C/G/T`, single-copy substitutions | | `composition_transitions` | 4×4 list of `{from, to, count, probability}`, base letters `A/C/G/T`, single-copy substitutions |
@@ -295,9 +295,8 @@ CSV matrix layout (`_dist.csv`, `_shared.csv`, `_family_overlap.csv`): header `g
**`_iqtree.model`** (`--iqtree`) — lower-triangular exchangeability matrix $R(a,b) = e^{-\text{cost}(a,b)}$ (one row of increasing length per state, whitespace-separated, PAML order), followed by one line of empirical state frequencies. Only states actually occurring in the alignment are kept, compactly renumbered `0..k-1`. **`_iqtree.model`** (`--iqtree`) — lower-triangular exchangeability matrix $R(a,b) = e^{-\text{cost}(a,b)}$ (one row of increasing length per state, whitespace-separated, PAML order), followed by one line of empirical state frequencies. Only states actually occurring in the alignment are kept, compactly renumbered `0..k-1`.
**`_iqtree.fasta`** (`--iqtree`) — alignment recoded to that same compact `0..k-1` alphabet (symbols `0-9A-F`). Under `--free-loss`, non-detection becomes `?` and columns left non-informative once missing calls are ignored are dropped first (required for `+ASC`); with `--iqtree-min-freq` also set (the default), any state rarer than that threshold is folded into the same `?` treatment, and non-informative columns are re-checked and dropped again. Run with: **`_iqtree.fasta`** (`--iqtree`) — alignment recoded to that same compact `0..k-1` alphabet (symbols `0-9A-F`). Under `--free-loss`, non-detection becomes `?` and columns left non-informative once missing calls are ignored are dropped first (required for `+ASC`); with `--iqtree-min-freq` also set (the default), any state rarer than that threshold is folded into the same `?` treatment, and non-informative columns are re-checked and dropped again. Run with:
```
iqtree3 -s <prefix>_iqtree.fasta --seqtype MORPH -m <prefix>_iqtree.model+ASC --prefix <prefix>_iqtree -T AUTO iqtree3 -s <prefix>_iqtree.fasta --seqtype MORPH -m <prefix>_iqtree.model+ASC --prefix <prefix>_iqtree -T AUTO
```
**`_iqtree_states.csv`** (`--iqtree`) — one row per state actually kept in `_iqtree.model`/`_iqtree.fasta` (header `iqtree_symbol,canonical_symbol,frequency`): `iqtree_symbol` is the compact `0-9A-F` symbol as written in those two files, `canonical_symbol` is the matching `_sankoff_matrix.csv` state, `frequency` is that state's empirical frequency at full precision. Under `--free-loss`, absent (`0`/`?`) is never a kept state, so it never appears here — nor does any state `--iqtree-min-freq` folded away for being too rare. **`_iqtree_states.csv`** (`--iqtree`) — one row per state actually kept in `_iqtree.model`/`_iqtree.fasta` (header `iqtree_symbol,canonical_symbol,frequency`): `iqtree_symbol` is the compact `0-9A-F` symbol as written in those two files, `canonical_symbol` is the matching `_sankoff_matrix.csv` state, `frequency` is that state's empirical frequency at full precision. Under `--free-loss`, absent (`0`/`?`) is never a kept state, so it never appears here — nor does any state `--iqtree-min-freq` folded away for being too rare.
+3 -5
@@ -5,7 +5,7 @@ Several commands ([`filter`](usage-filter), [`select`](usage-select), [`dump`](u
## Predicate syntax ## Predicate syntax
| Form | Meaning | | Form | Meaning |
|---|---| | --- | --- |
| `*` or `all` | Matches every genome (case-insensitive) | | `*` or `all` | Matches every genome (case-insensitive) |
| `key=v1\|v2` | Genome's `key` metadata equals one of the listed values | | `key=v1\|v2` | Genome's `key` metadata equals one of the listed values |
| `key!=v` | Genome's `key` metadata does not equal `v` | | `key!=v` | Genome's `key` metadata does not equal `v` |
@@ -20,9 +20,7 @@ Multiple `--ingroup` predicates are combined with AND; multiple `--outgroup` pre
A metadata value is treated as a taxonomy path when it starts with the literal prefix `taxonomy:/`; any other value is treated as a plain string and only supports `=`/`!=`. A metadata value is treated as a taxonomy path when it starts with the literal prefix `taxonomy:/`; any other value is treated as a plain string and only supports `=`/`!=`.
``` taxonomy:/segment1@rank1/segment2@rank2/...
taxonomy:/segment1@rank1/segment2@rank2/...
```
Each segment is a name, optionally annotated with a rank (e.g. `@family`, `@genus`, `@species`); ranks are optional and can be mixed within a path. The `@` character is reserved inside taxonomy paths and cannot appear in segment names or rank labels. Each segment is a name, optionally annotated with a rank (e.g. `@family`, `@genus`, `@species`); ranks are optional and can be mixed within a path. The `@` character is reserved inside taxonomy paths and cannot appear in segment names or rank labels.
@@ -31,7 +29,7 @@ Each segment is a name, optionally annotated with a rank (e.g. `@family`, `@genu
Matching compares segment names only (ranks are informational, not part of the match), with anchoring controlled by leading/trailing `/`: Matching compares segment names only (ranks are informational, not part of the match), with anchoring controlled by leading/trailing `/`:
| Pattern | Matches | | Pattern | Matches |
|---|---| | --- | --- |
| `A/B` | anywhere in the path | | `A/B` | anywhere in the path |
| `/A/B` | at the start of the path (prefix) | | `/A/B` | at the start of the path (prefix) |
| `A/B$` | at the end of the path (suffix) | | `A/B$` | at the end of the path (suffix) |
+2 -2
@@ -9,14 +9,14 @@ obikmer query INDEX INPUTS... [OPTIONS]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INDEX` | Index directory to query against | | `INDEX` | Index directory to query against |
| `INPUTS...` | Input sequence files (FASTA/FASTQ, gzip optional); at least one required | | `INPUTS...` | Input sequence files (FASTA/FASTQ, gzip optional); at least one required |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `--detail` | off | Report per-position, per-genome coverage vectors in the output | | `--detail` | off | Report per-position, per-genome coverage vectors in the output |
| `--count-missing` | off | Also count query kmers absent from the index | | `--count-missing` | off | Also count query kmers absent from the index |
| `--force-presence` | off | Report presence (0/1) per genome instead of raw counts | | `--force-presence` | off | Report presence (0/1) per genome instead of raw counts |
+2 -2
@@ -9,13 +9,13 @@ obikmer select SOURCE --output OUTPUT [OPTIONS]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `SOURCE` | Source index directory | | `SOURCE` | Source index directory |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `-o, --output` | — | Output index directory (required) | | `-o, --output` | — | Output index directory (required) |
| `-f, --force` | off | Overwrite an existing output directory | | `-f, --force` | off | Overwrite an existing output directory |
| `--group NAME:PRED` | none | Define a named group of genomes by predicate (repeatable; mutually exclusive with `--aggregate-by`) | | `--group NAME:PRED` | none | Define a named group of genomes by predicate (repeatable; mutually exclusive with `--aggregate-by`) |
+3 -3
@@ -9,14 +9,14 @@ obikmer superkmer [OPTIONS] [INPUTS...]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INPUTS...` | Input sequence files or directories (FASTA/FASTQ/GenBank, gzip optional). If omitted, reads from stdin. | | `INPUTS...` | Input sequence files or directories (FASTA/FASTQ/GenBank, gzip optional). If omitted, reads from stdin. |
## Options ## Options
| Option | Default | Description | | Option | Default | Description |
|---|---|---| | --- | --- | --- |
| `-k, --kmer-size` | `31` | Kmer size (must be odd, in [11, 31]) | | `-k, --kmer-size` | `31` | Kmer size (must be odd, in \[11, 31\]) |
| `-m, --minimizer-size` | `11` | Minimizer size (must be odd, in $[3, k-1]$) | | `-m, --minimizer-size` | `11` | Minimizer size (must be odd, in $[3, k-1]$) |
| `--theta` | `0.7` | Entropy threshold; kmers with a normalized entropy at or below this value are excluded | | `--theta` | `0.7` | Entropy threshold; kmers with a normalized entropy at or below this value are excluded |
| `--level-max` | `6` | Maximum sub-word size used for the entropy score | | `--level-max` | `6` | Maximum sub-word size used for the entropy score |
+1 -1
@@ -9,7 +9,7 @@ obikmer unitig INDEX [OPTIONS]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INDEX` | Index directory | | `INDEX` | Index directory |
## Options ## Options
+2 -2
@@ -9,13 +9,13 @@ obikmer utils INDEXES... [OPTIONS]
## Arguments ## Arguments
| Argument | Description | | Argument | Description |
|---|---| | --- | --- |
| `INDEXES...` | One or more index directories | | `INDEXES...` | One or more index directories |
## Options ## Options
| Option | Scope | Description | | Option | Scope | Description |
|---|---|---| | --- | --- | --- |
| `--new-label NEW=OLD` | single index only | Rename a genome label | | `--new-label NEW=OLD` | single index only | Rename a genome label |
| `--upgrade-index` | single index only | Add any missing layer metadata files to an older index | | `--upgrade-index` | single index only | Add any missing layer metadata files to an older index |
| `--bits-per-kmer` | single index only | Print bits-per-kmer statistics | | `--bits-per-kmer` | single index only | Print bits-per-kmer statistics |