Switch internal links to native Markdown syntax

Gitea/Forgejo's [[page|label]] shortlink is post-processed on rendered HTML
text nodes: it cannot match when the label contains inline formatting (a
code span splits the surrounding text into separate DOM nodes), and even
plain labels had target/text swapped from the intended order. Native
Markdown links are real AST link nodes and have neither problem.
Eric Coissac committed 2026-09-12 10:59:39 +02:00
1 parent 385f666fb6
commit 2af47f4c1f
16 files changed
+69 -143

No files matched your search

+21 -21
@@ -6,37 +6,37 @@ All functionality is exposed through a single binary, `obikmer`, organized as su
## Core principles
- Kmers are of fixed, odd length $k$, chosen at index-construction time in the range $[11, 31]$ (see [[theory-kmers_and_superkmers\|Kmers and super-kmers]]).
- Each kmer fits in a 64-bit word using a 2-bit-per-base encoding (see [[theory-encoding\|DNA encoding]]).
- Kmers are of fixed, odd length $k$, chosen at index-construction time in the range $[11, 31]$ (see [Kmers and super-kmers](theory-kmers_and_superkmers)).
- Each kmer fits in a 64-bit word using a 2-bit-per-base encoding (see [DNA encoding](theory-encoding)).
- Kmers are handled in **canonical form** ($\text{canonical}(kmer) = \min(kmer, \text{revcomp}(kmer))$), making counting strand-independent.
- Sequences are decomposed into **super-kmers** before storage, anchored on a hash-selected **minimizer** (see [[theory-minimizer_selection\|Minimizer selection]]), then routed to one of several **partitions** for parallel, memory-bounded processing (see [[theory-indexing_architecture\|Partitioning and indexing architecture]]).
- Low-complexity kmers can be filtered out at index-construction time using an entropy-based score (see [[theory-entropy_filter\|Low-complexity kmer filter]]).
- Sequences are decomposed into **super-kmers** before storage, anchored on a hash-selected **minimizer** (see [Minimizer selection](theory-minimizer_selection)), then routed to one of several **partitions** for parallel, memory-bounded processing (see [Partitioning and indexing architecture](theory-indexing_architecture)).
- Low-complexity kmers can be filtered out at index-construction time using an entropy-based score (see [Low-complexity kmer filter](theory-entropy_filter)).
## Commands
| Command | Purpose |
|---|---|
| [[usage-superkmer\|`superkmer`]] | Extract super-kmers from a sequence file and write them to stdout |
| [[usage-index_command\|`index`]] | Build a genome index |
| [[usage-merge\|`merge`]] | Merge multiple indexes into one |
| [[usage-filter\|`filter`]] | Retain only kmers matching ingroup/outgroup predicates |
| [[usage-select\|`select`]] | Project and/or aggregate genome columns of an index |
| [[usage-query\|`query`]] | Query an index with sequences and annotate matches |
| [[usage-dump\|`dump`]] | Dump indexed kmers as CSV |
| [[usage-annotate\|`annotate`]] | Add, update, or dump genome metadata |
| [[usage-phylo\|`phylo`]] | Compute pairwise genome distances, trees, and phylogenetic exports |
| [[usage-unitig\|`unitig`]] | Dump the unitigs of an index as FASTA |
| [[usage-estimate\|`estimate`]] | Estimate approximate-index parameters before indexing |
| [[usage-convert\|`convert`]] | Convert an index's evidence representation (exact/approximate/hybrid), in place |
| [[usage-utils\|`utils`]] | Miscellaneous index maintenance and inspection utilities |
| [[usage-pack\|`pack`]] | Pack per-column matrix files into a single-file format |
| [`superkmer`](usage-superkmer) | Extract super-kmers from a sequence file and write them to stdout |
| [`index`](usage-index_command) | Build a genome index |
| [`merge`](usage-merge) | Merge multiple indexes into one |
| [`filter`](usage-filter) | Retain only kmers matching ingroup/outgroup predicates |
| [`select`](usage-select) | Project and/or aggregate genome columns of an index |
| [`query`](usage-query) | Query an index with sequences and annotate matches |
| [`dump`](usage-dump) | Dump indexed kmers as CSV |
| [`annotate`](usage-annotate) | Add, update, or dump genome metadata |
| [`phylo`](usage-phylo) | Compute pairwise genome distances, trees, and phylogenetic exports |
| [`unitig`](usage-unitig) | Dump the unitigs of an index as FASTA |
| [`estimate`](usage-estimate) | Estimate approximate-index parameters before indexing |
| [`convert`](usage-convert) | Convert an index's evidence representation (exact/approximate/hybrid), in place |
| [`utils`](usage-utils) | Miscellaneous index maintenance and inspection utilities |
| [`pack`](usage-pack) | Pack per-column matrix files into a single-file format |
See [[usage-predicates\|Genome predicates and taxonomy paths]] for the selection language shared by `filter`, `select`, `dump`, and `unitig`.
See [Genome predicates and taxonomy paths](usage-predicates) for the selection language shared by `filter`, `select`, `dump`, and `unitig`.
## Further reading
- [[formats-index_layout\|Index construction and on-disk layout]]
- [[architecture\|Architecture notes for advanced use]] — parallel execution, NUMA awareness, index dimensioning
- [Index construction and on-disk layout](formats-index_layout)
- [Architecture notes for advanced use](architecture) — parallel execution, NUMA awareness, index dimensioning
## Input formats
-6
@@ -1,6 +0,0 @@
# Test
- [[usage-superkmer|wikilink plain]]
- [[usage-superkmer|`wikilink code`]]
- [markdown link plain](usage-superkmer)
- [`markdown link code`](usage-superkmer)
+24 -24
@@ -1,29 +1,29 @@
# Wiki sidebar
- [[Home\|Home]]
- [[installation\|Installation]]
- [Home](Home)
- [Installation](installation)
## Theory
- [[theory-kmers_and_superkmers\|Kmers and super-kmers]]
- [[theory-encoding\|DNA encoding]]
- [[theory-entropy_filter\|Low-complexity kmer filter]]
- [[theory-minimizer_selection\|Minimizer selection]]
- [[theory-indexing_architecture\|Partitioning and indexing architecture]]
- [Kmers and super-kmers](theory-kmers_and_superkmers)
- [DNA encoding](theory-encoding)
- [Low-complexity kmer filter](theory-entropy_filter)
- [Minimizer selection](theory-minimizer_selection)
- [Partitioning and indexing architecture](theory-indexing_architecture)
## Usage
- [[usage-superkmer\|superkmer]]
- [[usage-index_command\|index]]
- [[usage-merge\|merge]]
- [[usage-filter\|filter]]
- [[usage-select\|select]]
- [[usage-query\|query]]
- [[usage-dump\|dump]]
- [[usage-annotate\|annotate]]
- [[usage-phylo\|phylo]]
- [[usage-unitig\|unitig]]
- [[usage-estimate\|estimate]]
- [[usage-convert\|convert]]
- [[usage-utils\|utils]]
- [[usage-pack\|pack]]
- [[usage-predicates\|Predicates and taxonomy paths]]
- [superkmer](usage-superkmer)
- [index](usage-index_command)
- [merge](usage-merge)
- [filter](usage-filter)
- [select](usage-select)
- [query](usage-query)
- [dump](usage-dump)
- [annotate](usage-annotate)
- [phylo](usage-phylo)
- [unitig](usage-unitig)
- [estimate](usage-estimate)
- [convert](usage-convert)
- [utils](usage-utils)
- [pack](usage-pack)
- [Predicates and taxonomy paths](usage-predicates)
## Formats
- [[formats-index_layout\|Index construction and on-disk layout]]
- [[architecture\|Architecture notes]]
- [Index construction and on-disk layout](formats-index_layout)
- [Architecture notes](architecture)
+3 -3
@@ -1,6 +1,6 @@
# Architecture notes for advanced use
This page describes execution-level behavior relevant to sizing and running `obikmer` on large datasets or multi-socket machines. It complements the [[formats-index_layout\|index format]] and [[theory-indexing_architecture\|theory]] pages.
This page describes execution-level behavior relevant to sizing and running `obikmer` on large datasets or multi-socket machines. It complements the [index format](formats-index_layout) and [theory](theory-indexing_architecture) pages.
## Sequence invariant
@@ -8,7 +8,7 @@ Every input sequence is treated purely as a compact representation of a set of o
- Only the `A`/`C`/`G`/`T` alphabet (case-insensitive) is recognized; a sequence is cut at any other character (including IUPAC ambiguity codes), so runs containing them are not represented in the index.
- Sequences are internally processed in chunks of at most 256 nucleotides; a chunk shorter than k is dropped. This is invisible to the user beyond the ACGT-only, minimum-length-k constraints above.
- Kmers are always handled in canonical form (see [[theory-encoding\|DNA encoding]]), so the tool is strand-agnostic throughout: a kmer and its reverse complement are always the same entry.
- Kmers are always handled in canonical form (see [DNA encoding](theory-encoding)), so the tool is strand-agnostic throughout: a kmer and its reverse complement are always the same entry.
## Index dimensioning
@@ -30,4 +30,4 @@ No CLI flag controls this directly; it is fully automatic at runtime. NUMA-aware
## Kmer filtering (`filter`)
[[usage-filter\|`filter`]] evaluates predicates against the genome metadata matrix directly whenever every active filter can be expressed as a column-level test (e.g. "any outgroup column non-zero"), producing a per-slot keep/drop decision without touching kmer sequence data at all. If any active filter cannot be expressed this way, evaluation falls back to a per-kmer, row-level check. Either way, the result is always written as a single, freshly compacted layer (`unitigs.bin` and the MPHF are rebuilt from the surviving kmers), never as an additional layer on top of the source index.
[`filter`](usage-filter) evaluates predicates against the genome metadata matrix directly whenever every active filter can be expressed as a column-level test (e.g. "any outgroup column non-zero"), producing a per-slot keep/drop decision without touching kmer sequence data at all. If any active filter cannot be expressed this way, evaluation falls back to a per-kmer, row-level check. Either way, the result is always written as a single, freshly compacted layer (`unitigs.bin` and the MPHF are rebuilt from the surviving kmers), never as an additional layer on top of the source index.
+6 -6
@@ -2,9 +2,9 @@
## Construction pipeline
Building an index ([[usage-index_command\|`index`]]) proceeds through a fixed sequence of phases, each operating independently per partition (see [[theory-indexing_architecture\|Partitioning and indexing architecture]]):
Building an index ([`index`](usage-index_command)) proceeds through a fixed sequence of phases, each operating independently per partition (see [Partitioning and indexing architecture](theory-indexing_architecture)):
1. **Scatter.** A single streaming pass over the input. Each sequence fragment is cut at non-ACGT bases, passed through the low-complexity entropy filter (see [[theory-entropy_filter\|Low-complexity kmer filter]]), and any resulting segment shorter than k is dropped. Surviving segments are decomposed into super-kmers, canonicalized, and routed by `hash(minimizer) mod n_partitions` into one file per partition.
1. **Scatter.** A single streaming pass over the input. Each sequence fragment is cut at non-ACGT bases, passed through the low-complexity entropy filter (see [Low-complexity kmer filter](theory-entropy_filter)), and any resulting segment shorter than k is dropped. Surviving segments are decomposed into super-kmers, canonicalized, and routed by `hash(minimizer) mod n_partitions` into one file per partition.
2. **Dereplication.** Within each partition, identical super-kmer sequences are merged and their occurrence counts summed. This count is per super-kmer, not per kmer — a kmer's true abundance is the sum of the counts of every super-kmer containing it.
3. **Exact counting.** Every kmer in every dereplicated super-kmer is enumerated and its exact total count computed. A per-genome kmer frequency spectrum is produced at this stage.
4. **Quorum filtering.** Kmers outside the `--min-abundance`/`--max-abundance` range are dropped, and super-kmers are recompacted around the surviving kmer set.
@@ -19,10 +19,10 @@ Each partition's surviving kmers are mapped to a dense range of integer slots by
## Evidence: exact vs. approximate
Two verification modes are available, selected at build time (`index --approx`) and convertible afterwards ([[usage-convert\|`convert`]]):
Two verification modes are available, selected at build time (`index --approx`) and convertible afterwards ([`convert`](usage-convert)):
- **Exact** (default): the hashed slot stores a pointer back into the partition's unitig data. At query time the kmer is reconstructed from that location and compared directly to the query. Zero false positives, at the cost of one extra random read per lookup.
- **Approximate** (`--approx`): the slot stores a short fingerprint (`--evidence-bits` bits) instead of a pointer; verification is a single fingerprint comparison. This trades a small, bounded false-positive rate ($1/2^b$ per kmer, reduced further to about $1/2^{b \cdot z}$ for a read requiring $z$ consecutive matching kmers via the `-z`/`--findere-z` parameter) for lower memory and disk usage, since no reconstruction index is needed. See [[usage-estimate\|`estimate`]] to explore this trade-off before building.
- **Approximate** (`--approx`): the slot stores a short fingerprint (`--evidence-bits` bits) instead of a pointer; verification is a single fingerprint comparison. This trades a small, bounded false-positive rate ($1/2^b$ per kmer, reduced further to about $1/2^{b \cdot z}$ for a read requiring $z$ consecutive matching kmers via the `-z`/`--findere-z` parameter) for lower memory and disk usage, since no reconstruction index is needed. See [`estimate`](usage-estimate) to explore this trade-off before building.
## On-disk layout
@@ -49,8 +49,8 @@ Two verification modes are available, selected at build time (`index --approx`)
`unitigs.bin` is the only file from which the indexed kmer content can be fully recovered; it is always retained. Every other file (MPHF, evidence, counts) is derived from it.
A **layer** corresponds to one increment of kmer content added to a partition — most commonly, one [[usage-merge\|`merge`]] operation that introduces kmers not already present in the index. Genomes already present in the index simply gain new columns in the existing layers' count/presence data; only genuinely new kmer content is assembled into a new layer. Because of this, merging cost scales with the novel kmer content being added, not with the accumulated size of the index. A query against an index with several layers checks each layer's MPHF in turn.
A **layer** corresponds to one increment of kmer content added to a partition — most commonly, one [`merge`](usage-merge) operation that introduces kmers not already present in the index. Genomes already present in the index simply gain new columns in the existing layers' count/presence data; only genuinely new kmer content is assembled into a new layer. Because of this, merging cost scales with the novel kmer content being added, not with the accumulated size of the index. A query against an index with several layers checks each layer's MPHF in turn.
Sources merged together must share the same kmer size, minimizer size, partition count, and evidence mode (including matching approximate-mode parameters); mismatches are rejected rather than silently reconciled — [[usage-convert\|`convert`]] one of the sources first if needed.
Sources merged together must share the same kmer size, minimizer size, partition count, and evidence mode (including matching approximate-mode parameters); mismatches are rejected rather than silently reconciled — [`convert`](usage-convert) one of the sources first if needed.
`obikmer pack` consolidates a partition's per-column files (counts/presence) into a single file, reducing the number of file opens needed at query time.
-68
@@ -1,68 +0,0 @@
# obikmer
`obikmer` is a command-line tool for counting, indexing, querying and comparing DNA sequences represented as kmer sets. It targets individual genome datasets of tens of gigabases, with an emphasis on computational, memory, and disk efficiency.
All functionality is exposed through a single binary, `obikmer`, organized as subcommands.
## Core principles
- Kmers are of fixed, odd length $k$, chosen at index-construction time in the range $[11, 31]$ (see [[theory-kmers_and_superkmers|Kmers and super-kmers]]).
- Each kmer fits in a 64-bit word using a 2-bit-per-base encoding (see [[theory-encoding|DNA encoding]]).
- Kmers are handled in **canonical form** ($\text{canonical}(kmer) = \min(kmer, \text{revcomp}(kmer))$), making counting strand-independent.
- Sequences are decomposed into **super-kmers** before storage, anchored on a hash-selected **minimizer** (see [[theory-minimizer_selection|Minimizer selection]]), then routed to one of several **partitions** for parallel, memory-bounded processing (see [[theory-indexing_architecture|Partitioning and indexing architecture]]).
- Low-complexity kmers can be filtered out at index-construction time using an entropy-based score (see [[theory-entropy_filter|Low-complexity kmer filter]]).
## Commands
| Command | Purpose |
|---|---|
| [[usage-superkmer|`superkmer`]] | Extract super-kmers from a sequence file and write them to stdout |
| [[usage-index_command|`index`]] | Build a genome index |
| [[usage-merge|`merge`]] | Merge multiple indexes into one |
| [[usage-filter|`filter`]] | Retain only kmers matching ingroup/outgroup predicates |
| [[usage-select|`select`]] | Project and/or aggregate genome columns of an index |
| [[usage-query|`query`]] | Query an index with sequences and annotate matches |
| [[usage-dump|`dump`]] | Dump indexed kmers as CSV |
| [[usage-annotate|`annotate`]] | Add, update, or dump genome metadata |
| [[usage-phylo|`phylo`]] | Compute pairwise genome distances, trees, and phylogenetic exports |
| [[usage-unitig|`unitig`]] | Dump the unitigs of an index as FASTA |
| [[usage-estimate|`estimate`]] | Estimate approximate-index parameters before indexing |
| [[usage-convert|`convert`]] | Convert an index's evidence representation (exact/approximate/hybrid), in place |
| [[usage-utils|`utils`]] | Miscellaneous index maintenance and inspection utilities |
| [[usage-pack|`pack`]] | Pack per-column matrix files into a single-file format |
See [[usage-predicates|Genome predicates and taxonomy paths]] for the selection language shared by `filter`, `select`, `dump`, and `unitig`.
## Further reading
- [[formats-index_layout|Index construction and on-disk layout]]
- [[architecture|Architecture notes for advanced use]] — parallel execution, NUMA awareness, index dimensioning
## Input formats
- `superkmer` and `index`: FASTA (`.fa`, `.fasta`), FASTQ (`.fq`, `.fastq`), GenBank flat file (`.gb`, `.gbk`, `.gbff`), all optionally gzip-compressed; directories are expanded recursively; streaming stdin via `-` or when no input path is given.
- `query`: FASTA or FASTQ, optionally gzip-compressed; streaming stdin the same way.
## Parameter constraints
These constraints are checked at startup; an invalid value exits immediately with an error.
| Parameter | Constraint | Reason |
|---|---|---|
| $k$ (`--kmer-size`) | odd, $k \in [11, 31]$ | odd length guarantees the canonical form is always well defined; the range keeps a kmer within a 64-bit word while retaining specificity |
| $m$ (`--minimizer-size`) | odd, $3 \le m \le k-1$ | same palindrome argument as $k$; must be strictly shorter than the kmer |
| $z$ (`-z`, approximate evidence only) | $z \le k-1$ | the effective indexed kmer size is $k-z+1$ |
## Genome label constraints
Genome labels are arbitrary Unicode strings, with the following restrictions:
| Character | Forbidden | Reason |
|---|---|---|
| `/` | yes | filesystem path separator |
| `=` | yes | separator used by `--new-label` |
| `\0` | yes | null byte |
| `\n`, `\r`, `\t` | yes | would break CSV output |
| spaces | allowed | quote in the shell, e.g. `--new-label 'new label=old label'` |
Empty labels are rejected. A label derived automatically from the input file name (when `--label` is omitted) is not validated, since it is already filesystem-safe.
+2 -2
@@ -4,7 +4,7 @@ An index is split into a fixed number of **partitions**, each handling an indepe
## Routing
The canonical minimizer of a super-kmer (see [[theory-minimizer_selection\|Minimizer selection]]) is hashed to produce a $p$-bit routing value that selects the destination partition:
The canonical minimizer of a super-kmer (see [Minimizer selection](theory-minimizer_selection)) is hashed to produce a $p$-bit routing value that selects the destination partition:
```
canonical minimizer → hash(minimizer) → p-bit value → partition index
@@ -12,7 +12,7 @@ canonical minimizer → hash(minimizer) → p-bit value → partition index
The routing value is recomputed whenever it is needed (during construction and again at query time) rather than stored — it is not part of the on-disk super-kmer representation.
Within a partition, kmers are indexed as plain values via a minimal perfect hash function (see [[formats-index_layout\|On-disk storage]]); the minimizer plays no further role once a super-kmer has reached its partition.
Within a partition, kmers are indexed as plain values via a minimal perfect hash function (see [On-disk storage](formats-index_layout)); the minimizer plays no further role once a super-kmer has reached its partition.
## Why hashing is necessary
+2 -2
@@ -5,11 +5,11 @@
A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k$, both enforced when a command starts (an invalid value exits immediately with an error):
- $k \in [11, 31]$: long enough to be specific, short enough to fit in a single 64-bit word at 2 bits/base ($k \le 32$ is the hard limit; $k < 11$ gives insufficient specificity).
- $k$ **is odd**: an odd-length sequence can never equal its own reverse complement, so the two orientations of any kmer are always distinct. This is required for the canonical form (see [[theory-encoding\|DNA encoding]]) to be well defined.
- $k$ **is odd**: an odd-length sequence can never equal its own reverse complement, so the two orientations of any kmer are always distinct. This is required for the canonical form (see [DNA encoding](theory-encoding)) to be well defined.
## Super-kmers
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [[theory-minimizer_selection\|Minimizer selection]]). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately (@Zheng2020-ji; @Golan2025-xf):
+2 -2
@@ -4,7 +4,7 @@
A **minimizer** of a kmer window is the m-mer ($m < k$) that is smallest, among all $k - m + 1$ overlapping m-mers in the window, under a chosen ordering. The minimizer is always taken in canonical form (lexicographic minimum of forward and reverse complement) so that selection is strand-independent.
The minimizer partitions a sequence into super-kmers: maximal runs of overlapping kmers that share the same minimizer (see [[theory-kmers_and_superkmers\|Kmers and super-kmers]]).
The minimizer partitions a sequence into super-kmers: maximal runs of overlapping kmers that share the same minimizer (see [Kmers and super-kmers](theory-kmers_and_superkmers)).
## Hash-based ("random") minimizer
@@ -39,4 +39,4 @@ The hash used to select a minimizer within a window (the minimum of several hash
- **Selection** uses $H$ applied to every candidate m-mer in the window, keeping the minimum.
- **Partition routing** recomputes $H$ on the single selected minimizer only, once its position is fixed. This is a hash of one specific value, not the minimum of several, so it is uniformly distributed and safe to use directly for routing.
See [[theory-indexing_architecture\|Partitioning and indexing architecture]] for how the routing value is turned into a partition index.
See [Partitioning and indexing architecture](theory-indexing_architecture) for how the routing value is turned into a partition index.
+1 -1
@@ -26,4 +26,4 @@ Exactly one of the first three is required:
| `--fp FP` | Target false-positive rate per z-window (e.g. `0.01`); derives `b` or `z` when one of them isn't given directly |
| `--block-size N` | Block size for exact evidence's on-disk index (unitigs per block). Ignored when converting to pure approximate evidence. Default `1` |
See [[usage-index_command#exact-vs-approximate-evidence\|`index`]] for the exact/approximate trade-off and the underlying false-positive model, and [[usage-estimate\|`estimate`]] to explore parameters beforehand. The index directory is locked for exclusive access during conversion.
See [`index`](usage-index_command#exact-vs-approximate-evidence) for the exact/approximate trade-off and the underlying false-positive model, and [`estimate`](usage-estimate) to explore parameters beforehand. The index directory is locked for exclusive access during conversion.
+1 -1
@@ -20,6 +20,6 @@ obikmer dump INDEX [OPTIONS]
| `--debug` | off | Prefix each row with the partition and layer columns |
| `--head N` | none | Limit output to the first N kmers |
`dump` also accepts the shared [[usage-filter#predicate-options\|predicate options]] (`--ingroup`, `--outgroup`, `--min-count`, etc.) to restrict which kmers are dumped.
`dump` also accepts the shared [predicate options](usage-filter#predicate-options) (`--ingroup`, `--outgroup`, `--min-count`, etc.) to restrict which kmers are dumped.
Output is CSV on stdout.
+1 -1
@@ -40,7 +40,7 @@ obikmer filter SOURCE -o OUTPUT [OPTIONS]
| `--max-outgroup-frac` | `1.0` | Maximum fraction of outgroup genomes |
| `--presence-threshold` | `0` | Minimum count for a genome to be considered a carrier of a kmer |
See [[usage-predicates\|Genome predicates and taxonomy paths]] for the predicate syntax used by `--ingroup`/`--outgroup`.
See [Genome predicates and taxonomy paths](usage-predicates) for the predicate syntax used by `--ingroup`/`--outgroup`.
A negative `--min-count`/`--max-count` is interpreted as an offset from the group size — e.g. `--min-count=-1` means "all but one".
+2 -2
@@ -1,6 +1,6 @@
# index
Build a genome index from one or more sequence files. Construction proceeds in phases (scatter → dereplicate → count → layered MPHF), described in [[formats-index_layout\|On-disk storage]].
Build a genome index from one or more sequence files. Construction proceeds in phases (scatter → dereplicate → count → layered MPHF), described in [On-disk storage](formats-index_layout).
```bash
obikmer index -o OUTPUT [OPTIONS] [INPUTS...]
@@ -45,6 +45,6 @@ With `--approx`, evidence is stored as a compact **fingerprint** instead, tradin
$$FP = \frac{1}{2^{b \cdot z}}$$
where $b$ is `--evidence-bits` and $z$ is `--findere-z`. Any two of `-z`, `--evidence-bits`, `--fp` can be given and the third is derived; if none are given, defaults are $b=8$, $z=1$ ($FP \approx 1/256$). See [[usage-estimate\|`estimate`]] to explore this trade-off before building an index, and [[usage-convert\|`convert`]] to change an existing index's representation afterwards.
where $b$ is `--evidence-bits` and $z$ is `--findere-z`. Any two of `-z`, `--evidence-bits`, `--fp` can be given and the third is derived; if none are given, defaults are $b=8$, $z=1$ ($FP \approx 1/256$). See [`estimate`](usage-estimate) to explore this trade-off before building an index, and [`convert`](usage-convert) to change an existing index's representation afterwards.
`z` must be strictly less than k: the effective indexed kmer length under approximate evidence is k−z+1.
+1 -1
@@ -1,6 +1,6 @@
# Genome predicates and taxonomy paths
Several commands ([[usage-filter\|`filter`]], [[usage-select\|`select`]], [[usage-dump\|`dump`]], [[usage-unitig\|`unitig`]]) select or group genomes using the same predicate language over genome metadata (see [[usage-annotate\|`annotate`]] for attaching metadata to a genome).
Several commands ([`filter`](usage-filter), [`select`](usage-select), [`dump`](usage-dump), [`unitig`](usage-unitig)) select or group genomes using the same predicate language over genome metadata (see [`annotate`](usage-annotate) for attaching metadata to a genome).
## Predicate syntax
+2 -2
@@ -1,6 +1,6 @@
# select
Project and/or aggregate the genome columns of an index into a new index. Where [[usage-filter\|`filter`]] selects rows (kmers), `select` operates on columns (genomes): grouping several genomes into one aggregated column, reordering columns, or dropping some.
Project and/or aggregate the genome columns of an index into a new index. Where [`filter`](usage-filter) selects rows (kmers), `select` operates on columns (genomes): grouping several genomes into one aggregated column, reordering columns, or dropping some.
```bash
obikmer select SOURCE --output OUTPUT [OPTIONS]
@@ -33,7 +33,7 @@ obikmer select SOURCE --output OUTPUT [OPTIONS]
A `select` never changes the underlying kmer set — only the per-genome data (counts or presence) is rewritten, so an unaggregated pass-through column (a plain genome label in `--select`) is a cheap copy.
At least one output column must be defined; every name listed in `--select` must resolve to either a defined group or an existing genome label. See [[usage-predicates\|Genome predicates and taxonomy paths]] for the predicate syntax used by `--group`.
At least one output column must be defined; every name listed in `--select` must resolve to either a defined group or an existing genome label. See [Genome predicates and taxonomy paths](usage-predicates) for the predicate syntax used by `--group`.
## Disk usage
+1 -1
@@ -14,6 +14,6 @@ obikmer unitig INDEX [OPTIONS]
## Options
`unitig` accepts the shared [[usage-filter#predicate-options\|predicate options]] (`--ingroup`, `--outgroup`, `--min-count`, etc.) to restrict which kmers are included before the unitigs are enumerated.
`unitig` accepts the shared [predicate options](usage-filter#predicate-options) (`--ingroup`, `--outgroup`, `--min-count`, etc.) to restrict which kmers are included before the unitigs are enumerated.
Output is FASTA on stdout.