From 62c2b9628f8927584c37128731ed395dd2e98207 Mon Sep 17 00:00:00 2001 From: Eric Coissac Date: Sat, 12 Sep 2026 10:35:54 +0200 Subject: [PATCH] Refine indexing architecture and command usage documentation Updates the core documentation for kmer structure, super-kmer routing, and indexing layout. Introduces new command-line options for filtering, dumping, and conversion, clarifying the Exact/Approx trade-offs and constraints for sequence processing. --- Home.md | 42 ++++++++++++++++----------------- architecture.md | 6 ++--- formats-index_layout.md | 12 +++++----- theory-indexing_architecture.md | 4 ++-- theory-kmers_and_superkmers.md | 4 ++-- theory-minimizer_selection.md | 4 ++-- usage-convert.md | 2 +- usage-dump.md | 2 +- usage-filter.md | 2 +- usage-index_command.md | 4 ++-- usage-predicates.md | 2 +- usage-select.md | 4 ++-- usage-unitig.md | 2 +- 13 files changed, 45 insertions(+), 45 deletions(-) diff --git a/Home.md b/Home.md index 4e3ebb3..066da3d 100644 --- a/Home.md +++ b/Home.md @@ -6,37 +6,37 @@ All functionality is exposed through a single binary, `obikmer`, organized as su ## Core principles -- Kmers are of fixed, odd length $k$, chosen at index-construction time in the range $[11, 31]$ (see [[theory-kmers_and_superkmers|Kmers and super-kmers]]). -- Each kmer fits in a 64-bit word using a 2-bit-per-base encoding (see [[theory-encoding|DNA encoding]]). +- Kmers are of fixed, odd length $k$, chosen at index-construction time in the range $[11, 31]$ (see [[theory-kmers_and_superkmers\|Kmers and super-kmers]]). +- Each kmer fits in a 64-bit word using a 2-bit-per-base encoding (see [[theory-encoding\|DNA encoding]]). - Kmers are handled in **canonical form** ($\text{canonical}(kmer) = \min(kmer, \text{revcomp}(kmer))$), making counting strand-independent. -- Sequences are decomposed into **super-kmers** before storage, anchored on a hash-selected **minimizer** (see [[theory-minimizer_selection|Minimizer selection]]), then routed to one of several **partitions** for parallel, memory-bounded processing (see [[theory-indexing_architecture|Partitioning and indexing architecture]]). -- Low-complexity kmers can be filtered out at index-construction time using an entropy-based score (see [[theory-entropy_filter|Low-complexity kmer filter]]). +- Sequences are decomposed into **super-kmers** before storage, anchored on a hash-selected **minimizer** (see [[theory-minimizer_selection\|Minimizer selection]]), then routed to one of several **partitions** for parallel, memory-bounded processing (see [[theory-indexing_architecture\|Partitioning and indexing architecture]]). +- Low-complexity kmers can be filtered out at index-construction time using an entropy-based score (see [[theory-entropy_filter\|Low-complexity kmer filter]]). ## Commands | Command | Purpose | |---|---| -| [[usage-superkmer|`superkmer`]] | Extract super-kmers from a sequence file and write them to stdout | -| [[usage-index_command|`index`]] | Build a genome index | -| [[usage-merge|`merge`]] | Merge multiple indexes into one | -| [[usage-filter|`filter`]] | Retain only kmers matching ingroup/outgroup predicates | -| [[usage-select|`select`]] | Project and/or aggregate genome columns of an index | -| [[usage-query|`query`]] | Query an index with sequences and annotate matches | -| [[usage-dump|`dump`]] | Dump indexed kmers as CSV | -| [[usage-annotate|`annotate`]] | Add, update, or dump genome metadata | -| [[usage-phylo|`phylo`]] | Compute pairwise genome distances, trees, and phylogenetic exports | -| [[usage-unitig|`unitig`]] | Dump the unitigs of an index as FASTA | -| [[usage-estimate|`estimate`]] | Estimate approximate-index parameters before indexing | -| [[usage-convert|`convert`]] | Convert an index's evidence representation (exact/approximate/hybrid), in place | -| [[usage-utils|`utils`]] | Miscellaneous index maintenance and inspection utilities | -| [[usage-pack|`pack`]] | Pack per-column matrix files into a single-file format | +| [[usage-superkmer\|`superkmer`]] | Extract super-kmers from a sequence file and write them to stdout | +| [[usage-index_command\|`index`]] | Build a genome index | +| [[usage-merge\|`merge`]] | Merge multiple indexes into one | +| [[usage-filter\|`filter`]] | Retain only kmers matching ingroup/outgroup predicates | +| [[usage-select\|`select`]] | Project and/or aggregate genome columns of an index | +| [[usage-query\|`query`]] | Query an index with sequences and annotate matches | +| [[usage-dump\|`dump`]] | Dump indexed kmers as CSV | +| [[usage-annotate\|`annotate`]] | Add, update, or dump genome metadata | +| [[usage-phylo\|`phylo`]] | Compute pairwise genome distances, trees, and phylogenetic exports | +| [[usage-unitig\|`unitig`]] | Dump the unitigs of an index as FASTA | +| [[usage-estimate\|`estimate`]] | Estimate approximate-index parameters before indexing | +| [[usage-convert\|`convert`]] | Convert an index's evidence representation (exact/approximate/hybrid), in place | +| [[usage-utils\|`utils`]] | Miscellaneous index maintenance and inspection utilities | +| [[usage-pack\|`pack`]] | Pack per-column matrix files into a single-file format | -See [[usage-predicates|Genome predicates and taxonomy paths]] for the selection language shared by `filter`, `select`, `dump`, and `unitig`. +See [[usage-predicates\|Genome predicates and taxonomy paths]] for the selection language shared by `filter`, `select`, `dump`, and `unitig`. ## Further reading -- [[formats-index_layout|Index construction and on-disk layout]] -- [[architecture|Architecture notes for advanced use]] — parallel execution, NUMA awareness, index dimensioning +- [[formats-index_layout\|Index construction and on-disk layout]] +- [[architecture\|Architecture notes for advanced use]] — parallel execution, NUMA awareness, index dimensioning ## Input formats diff --git a/architecture.md b/architecture.md index 861a762..4110d07 100644 --- a/architecture.md +++ b/architecture.md @@ -1,6 +1,6 @@ # Architecture notes for advanced use -This page describes execution-level behavior relevant to sizing and running `obikmer` on large datasets or multi-socket machines. It complements the [[formats-index_layout|index format]] and [[theory-indexing_architecture|theory]] pages. +This page describes execution-level behavior relevant to sizing and running `obikmer` on large datasets or multi-socket machines. It complements the [[formats-index_layout\|index format]] and [[theory-indexing_architecture\|theory]] pages. ## Sequence invariant @@ -8,7 +8,7 @@ Every input sequence is treated purely as a compact representation of a set of o - Only the `A`/`C`/`G`/`T` alphabet (case-insensitive) is recognized; a sequence is cut at any other character (including IUPAC ambiguity codes), so runs containing them are not represented in the index. - Sequences are internally processed in chunks of at most 256 nucleotides; a chunk shorter than k is dropped. This is invisible to the user beyond the ACGT-only, minimum-length-k constraints above. -- Kmers are always handled in canonical form (see [[theory-encoding|DNA encoding]]), so the tool is strand-agnostic throughout: a kmer and its reverse complement are always the same entry. +- Kmers are always handled in canonical form (see [[theory-encoding\|DNA encoding]]), so the tool is strand-agnostic throughout: a kmer and its reverse complement are always the same entry. ## Index dimensioning @@ -30,4 +30,4 @@ No CLI flag controls this directly; it is fully automatic at runtime. NUMA-aware ## Kmer filtering (`filter`) -[[usage-filter|`filter`]] evaluates predicates against the genome metadata matrix directly whenever every active filter can be expressed as a column-level test (e.g. "any outgroup column non-zero"), producing a per-slot keep/drop decision without touching kmer sequence data at all. If any active filter cannot be expressed this way, evaluation falls back to a per-kmer, row-level check. Either way, the result is always written as a single, freshly compacted layer (`unitigs.bin` and the MPHF are rebuilt from the surviving kmers), never as an additional layer on top of the source index. +[[usage-filter\|`filter`]] evaluates predicates against the genome metadata matrix directly whenever every active filter can be expressed as a column-level test (e.g. "any outgroup column non-zero"), producing a per-slot keep/drop decision without touching kmer sequence data at all. If any active filter cannot be expressed this way, evaluation falls back to a per-kmer, row-level check. Either way, the result is always written as a single, freshly compacted layer (`unitigs.bin` and the MPHF are rebuilt from the surviving kmers), never as an additional layer on top of the source index. diff --git a/formats-index_layout.md b/formats-index_layout.md index 12f5c6e..aa03071 100644 --- a/formats-index_layout.md +++ b/formats-index_layout.md @@ -2,9 +2,9 @@ ## Construction pipeline -Building an index ([[usage-index_command|`index`]]) proceeds through a fixed sequence of phases, each operating independently per partition (see [[theory-indexing_architecture|Partitioning and indexing architecture]]): +Building an index ([[usage-index_command\|`index`]]) proceeds through a fixed sequence of phases, each operating independently per partition (see [[theory-indexing_architecture\|Partitioning and indexing architecture]]): -1. **Scatter.** A single streaming pass over the input. Each sequence fragment is cut at non-ACGT bases, passed through the low-complexity entropy filter (see [[theory-entropy_filter|Low-complexity kmer filter]]), and any resulting segment shorter than k is dropped. Surviving segments are decomposed into super-kmers, canonicalized, and routed by `hash(minimizer) mod n_partitions` into one file per partition. +1. **Scatter.** A single streaming pass over the input. Each sequence fragment is cut at non-ACGT bases, passed through the low-complexity entropy filter (see [[theory-entropy_filter\|Low-complexity kmer filter]]), and any resulting segment shorter than k is dropped. Surviving segments are decomposed into super-kmers, canonicalized, and routed by `hash(minimizer) mod n_partitions` into one file per partition. 2. **Dereplication.** Within each partition, identical super-kmer sequences are merged and their occurrence counts summed. This count is per super-kmer, not per kmer — a kmer's true abundance is the sum of the counts of every super-kmer containing it. 3. **Exact counting.** Every kmer in every dereplicated super-kmer is enumerated and its exact total count computed. A per-genome kmer frequency spectrum is produced at this stage. 4. **Quorum filtering.** Kmers outside the `--min-abundance`/`--max-abundance` range are dropped, and super-kmers are recompacted around the surviving kmer set. @@ -19,10 +19,10 @@ Each partition's surviving kmers are mapped to a dense range of integer slots by ## Evidence: exact vs. approximate -Two verification modes are available, selected at build time (`index --approx`) and convertible afterwards ([[usage-convert|`convert`]]): +Two verification modes are available, selected at build time (`index --approx`) and convertible afterwards ([[usage-convert\|`convert`]]): - **Exact** (default): the hashed slot stores a pointer back into the partition's unitig data. At query time the kmer is reconstructed from that location and compared directly to the query. Zero false positives, at the cost of one extra random read per lookup. -- **Approximate** (`--approx`): the slot stores a short fingerprint (`--evidence-bits` bits) instead of a pointer; verification is a single fingerprint comparison. This trades a small, bounded false-positive rate ($1/2^b$ per kmer, reduced further to about $1/2^{b \cdot z}$ for a read requiring $z$ consecutive matching kmers via the `-z`/`--findere-z` parameter) for lower memory and disk usage, since no reconstruction index is needed. See [[usage-estimate|`estimate`]] to explore this trade-off before building. +- **Approximate** (`--approx`): the slot stores a short fingerprint (`--evidence-bits` bits) instead of a pointer; verification is a single fingerprint comparison. This trades a small, bounded false-positive rate ($1/2^b$ per kmer, reduced further to about $1/2^{b \cdot z}$ for a read requiring $z$ consecutive matching kmers via the `-z`/`--findere-z` parameter) for lower memory and disk usage, since no reconstruction index is needed. See [[usage-estimate\|`estimate`]] to explore this trade-off before building. ## On-disk layout @@ -49,8 +49,8 @@ Two verification modes are available, selected at build time (`index --approx`) `unitigs.bin` is the only file from which the indexed kmer content can be fully recovered; it is always retained. Every other file (MPHF, evidence, counts) is derived from it. -A **layer** corresponds to one increment of kmer content added to a partition — most commonly, one [[usage-merge|`merge`]] operation that introduces kmers not already present in the index. Genomes already present in the index simply gain new columns in the existing layers' count/presence data; only genuinely new kmer content is assembled into a new layer. Because of this, merging cost scales with the novel kmer content being added, not with the accumulated size of the index. A query against an index with several layers checks each layer's MPHF in turn. +A **layer** corresponds to one increment of kmer content added to a partition — most commonly, one [[usage-merge\|`merge`]] operation that introduces kmers not already present in the index. Genomes already present in the index simply gain new columns in the existing layers' count/presence data; only genuinely new kmer content is assembled into a new layer. Because of this, merging cost scales with the novel kmer content being added, not with the accumulated size of the index. A query against an index with several layers checks each layer's MPHF in turn. -Sources merged together must share the same kmer size, minimizer size, partition count, and evidence mode (including matching approximate-mode parameters); mismatches are rejected rather than silently reconciled — [[usage-convert|`convert`]] one of the sources first if needed. +Sources merged together must share the same kmer size, minimizer size, partition count, and evidence mode (including matching approximate-mode parameters); mismatches are rejected rather than silently reconciled — [[usage-convert\|`convert`]] one of the sources first if needed. `obikmer pack` consolidates a partition's per-column files (counts/presence) into a single file, reducing the number of file opens needed at query time. diff --git a/theory-indexing_architecture.md b/theory-indexing_architecture.md index 2b32669..6f01942 100644 --- a/theory-indexing_architecture.md +++ b/theory-indexing_architecture.md @@ -4,7 +4,7 @@ An index is split into a fixed number of **partitions**, each handling an indepe ## Routing -The canonical minimizer of a super-kmer (see [[theory-minimizer_selection|Minimizer selection]]) is hashed to produce a $p$-bit routing value that selects the destination partition: +The canonical minimizer of a super-kmer (see [[theory-minimizer_selection\|Minimizer selection]]) is hashed to produce a $p$-bit routing value that selects the destination partition: ``` canonical minimizer → hash(minimizer) → p-bit value → partition index @@ -12,7 +12,7 @@ canonical minimizer → hash(minimizer) → p-bit value → partition index The routing value is recomputed whenever it is needed (during construction and again at query time) rather than stored — it is not part of the on-disk super-kmer representation. -Within a partition, kmers are indexed as plain values via a minimal perfect hash function (see [[formats-index_layout|On-disk storage]]); the minimizer plays no further role once a super-kmer has reached its partition. +Within a partition, kmers are indexed as plain values via a minimal perfect hash function (see [[formats-index_layout\|On-disk storage]]); the minimizer plays no further role once a super-kmer has reached its partition. ## Why hashing is necessary diff --git a/theory-kmers_and_superkmers.md b/theory-kmers_and_superkmers.md index 695b6f7..4785147 100644 --- a/theory-kmers_and_superkmers.md +++ b/theory-kmers_and_superkmers.md @@ -5,11 +5,11 @@ A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k$, both enforced when a command starts (an invalid value exits immediately with an error): - $k \in [11, 31]$: long enough to be specific, short enough to fit in a single 64-bit word at 2 bits/base ($k \le 32$ is the hard limit; $k < 11$ gives insufficient specificity). -- $k$ **is odd**: an odd-length sequence can never equal its own reverse complement, so the two orientations of any kmer are always distinct. This is required for the canonical form (see [[theory-encoding|DNA encoding]]) to be well defined. +- $k$ **is odd**: an odd-length sequence can never equal its own reverse complement, so the two orientations of any kmer are always distinct. This is required for the canonical form (see [[theory-encoding\|DNA encoding]]) to be well defined. ## Super-kmers -A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [[theory-minimizer_selection|Minimizer selection]]). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary. +A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [[theory-minimizer_selection\|Minimizer selection]]). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary. For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately (@Zheng2020-ji; @Golan2025-xf): diff --git a/theory-minimizer_selection.md b/theory-minimizer_selection.md index 4d49c47..62a1393 100644 --- a/theory-minimizer_selection.md +++ b/theory-minimizer_selection.md @@ -4,7 +4,7 @@ A **minimizer** of a kmer window is the m-mer ($m < k$) that is smallest, among all $k - m + 1$ overlapping m-mers in the window, under a chosen ordering. The minimizer is always taken in canonical form (lexicographic minimum of forward and reverse complement) so that selection is strand-independent. -The minimizer partitions a sequence into super-kmers: maximal runs of overlapping kmers that share the same minimizer (see [[theory-kmers_and_superkmers|Kmers and super-kmers]]). +The minimizer partitions a sequence into super-kmers: maximal runs of overlapping kmers that share the same minimizer (see [[theory-kmers_and_superkmers\|Kmers and super-kmers]]). ## Hash-based ("random") minimizer @@ -39,4 +39,4 @@ The hash used to select a minimizer within a window (the minimum of several hash - **Selection** uses $H$ applied to every candidate m-mer in the window, keeping the minimum. - **Partition routing** recomputes $H$ on the single selected minimizer only, once its position is fixed. This is a hash of one specific value, not the minimum of several, so it is uniformly distributed and safe to use directly for routing. -See [[theory-indexing_architecture|Partitioning and indexing architecture]] for how the routing value is turned into a partition index. +See [[theory-indexing_architecture\|Partitioning and indexing architecture]] for how the routing value is turned into a partition index. diff --git a/usage-convert.md b/usage-convert.md index 43ceced..e87ca55 100644 --- a/usage-convert.md +++ b/usage-convert.md @@ -26,4 +26,4 @@ Exactly one of the first three is required: | `--fp FP` | Target false-positive rate per z-window (e.g. `0.01`); derives `b` or `z` when one of them isn't given directly | | `--block-size N` | Block size for exact evidence's on-disk index (unitigs per block). Ignored when converting to pure approximate evidence. Default `1` | -See [[usage-index_command#exact-vs-approximate-evidence|`index`]] for the exact/approximate trade-off and the underlying false-positive model, and [[usage-estimate|`estimate`]] to explore parameters beforehand. The index directory is locked for exclusive access during conversion. +See [[usage-index_command#exact-vs-approximate-evidence\|`index`]] for the exact/approximate trade-off and the underlying false-positive model, and [[usage-estimate\|`estimate`]] to explore parameters beforehand. The index directory is locked for exclusive access during conversion. diff --git a/usage-dump.md b/usage-dump.md index e70d8b5..eaef0c6 100644 --- a/usage-dump.md +++ b/usage-dump.md @@ -20,6 +20,6 @@ obikmer dump INDEX [OPTIONS] | `--debug` | off | Prefix each row with the partition and layer columns | | `--head N` | none | Limit output to the first N kmers | -`dump` also accepts the shared [[usage-filter#predicate-options|predicate options]] (`--ingroup`, `--outgroup`, `--min-count`, etc.) to restrict which kmers are dumped. +`dump` also accepts the shared [[usage-filter#predicate-options\|predicate options]] (`--ingroup`, `--outgroup`, `--min-count`, etc.) to restrict which kmers are dumped. Output is CSV on stdout. diff --git a/usage-filter.md b/usage-filter.md index addac0d..dc498a8 100644 --- a/usage-filter.md +++ b/usage-filter.md @@ -40,7 +40,7 @@ obikmer filter SOURCE -o OUTPUT [OPTIONS] | `--max-outgroup-frac` | `1.0` | Maximum fraction of outgroup genomes | | `--presence-threshold` | `0` | Minimum count for a genome to be considered a carrier of a kmer | -See [[usage-predicates|Genome predicates and taxonomy paths]] for the predicate syntax used by `--ingroup`/`--outgroup`. +See [[usage-predicates\|Genome predicates and taxonomy paths]] for the predicate syntax used by `--ingroup`/`--outgroup`. A negative `--min-count`/`--max-count` is interpreted as an offset from the group size — e.g. `--min-count=-1` means "all but one". diff --git a/usage-index_command.md b/usage-index_command.md index 41507d1..f12b93d 100644 --- a/usage-index_command.md +++ b/usage-index_command.md @@ -1,6 +1,6 @@ # index -Build a genome index from one or more sequence files. Construction proceeds in phases (scatter → dereplicate → count → layered MPHF), described in [[formats-index_layout|On-disk storage]]. +Build a genome index from one or more sequence files. Construction proceeds in phases (scatter → dereplicate → count → layered MPHF), described in [[formats-index_layout\|On-disk storage]]. ```bash obikmer index -o OUTPUT [OPTIONS] [INPUTS...] @@ -45,6 +45,6 @@ With `--approx`, evidence is stored as a compact **fingerprint** instead, tradin $$FP = \frac{1}{2^{b \cdot z}}$$ -where $b$ is `--evidence-bits` and $z$ is `--findere-z`. Any two of `-z`, `--evidence-bits`, `--fp` can be given and the third is derived; if none are given, defaults are $b=8$, $z=1$ ($FP \approx 1/256$). See [[usage-estimate|`estimate`]] to explore this trade-off before building an index, and [[usage-convert|`convert`]] to change an existing index's representation afterwards. +where $b$ is `--evidence-bits` and $z$ is `--findere-z`. Any two of `-z`, `--evidence-bits`, `--fp` can be given and the third is derived; if none are given, defaults are $b=8$, $z=1$ ($FP \approx 1/256$). See [[usage-estimate\|`estimate`]] to explore this trade-off before building an index, and [[usage-convert\|`convert`]] to change an existing index's representation afterwards. `z` must be strictly less than k: the effective indexed kmer length under approximate evidence is k−z+1. diff --git a/usage-predicates.md b/usage-predicates.md index cd30ba9..8064210 100644 --- a/usage-predicates.md +++ b/usage-predicates.md @@ -1,6 +1,6 @@ # Genome predicates and taxonomy paths -Several commands ([[usage-filter|`filter`]], [[usage-select|`select`]], [[usage-dump|`dump`]], [[usage-unitig|`unitig`]]) select or group genomes using the same predicate language over genome metadata (see [[usage-annotate|`annotate`]] for attaching metadata to a genome). +Several commands ([[usage-filter\|`filter`]], [[usage-select\|`select`]], [[usage-dump\|`dump`]], [[usage-unitig\|`unitig`]]) select or group genomes using the same predicate language over genome metadata (see [[usage-annotate\|`annotate`]] for attaching metadata to a genome). ## Predicate syntax diff --git a/usage-select.md b/usage-select.md index 3016425..3518733 100644 --- a/usage-select.md +++ b/usage-select.md @@ -1,6 +1,6 @@ # select -Project and/or aggregate the genome columns of an index into a new index. Where [[usage-filter|`filter`]] selects rows (kmers), `select` operates on columns (genomes): grouping several genomes into one aggregated column, reordering columns, or dropping some. +Project and/or aggregate the genome columns of an index into a new index. Where [[usage-filter\|`filter`]] selects rows (kmers), `select` operates on columns (genomes): grouping several genomes into one aggregated column, reordering columns, or dropping some. ```bash obikmer select SOURCE --output OUTPUT [OPTIONS] @@ -33,7 +33,7 @@ obikmer select SOURCE --output OUTPUT [OPTIONS] A `select` never changes the underlying kmer set — only the per-genome data (counts or presence) is rewritten, so an unaggregated pass-through column (a plain genome label in `--select`) is a cheap copy. -At least one output column must be defined; every name listed in `--select` must resolve to either a defined group or an existing genome label. See [[usage-predicates|Genome predicates and taxonomy paths]] for the predicate syntax used by `--group`. +At least one output column must be defined; every name listed in `--select` must resolve to either a defined group or an existing genome label. See [[usage-predicates\|Genome predicates and taxonomy paths]] for the predicate syntax used by `--group`. ## Disk usage diff --git a/usage-unitig.md b/usage-unitig.md index 1793ca0..23c78b7 100644 --- a/usage-unitig.md +++ b/usage-unitig.md @@ -14,6 +14,6 @@ obikmer unitig INDEX [OPTIONS] ## Options -`unitig` accepts the shared [[usage-filter#predicate-options|predicate options]] (`--ingroup`, `--outgroup`, `--min-count`, etc.) to restrict which kmers are included before the unitigs are enumerated. +`unitig` accepts the shared [[usage-filter#predicate-options\|predicate options]] (`--ingroup`, `--outgroup`, `--min-count`, etc.) to restrict which kmers are included before the unitigs are enumerated. Output is FASTA on stdout.