Implement super-kmers and hash-based minimizer selection

Introduces the concept of super-kmers as the primary unit of work, capped at 256 nucleotides. Also implements a new minimizer selection strategy based on a well-distributed hash function to ensure unbiased selection.
Eric Coissac committed 2026-09-12 13:01:51 +02:00
1 parent 331b9d330d
commit c1e139c597
2 files changed
+6 -2

No files matched your search

+3 -1
@@ -11,7 +11,7 @@ A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately (Golan & Shur 2025; Zheng *et al.* 2020):
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately ([Golan & Shur 2025](#ref-Golan2025-xf); [Zheng *et al.* 2020](#ref-Zheng2020-ji)):
$$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$
@@ -23,6 +23,8 @@ A **canonical super-kmer** is the lexicographic minimum of a super-kmer and its
Super-kmers are the unit of work used throughout construction and querying: sequences are decomposed into super-kmers first, and every downstream step (partition routing, deduplication, counting) operates on them rather than on individual kmers.
## Bibliography
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
<div id="ref-Golan2025-xf" class="csl-entry">
+3 -1
@@ -8,7 +8,7 @@ The minimizer partitions a sequence into super-kmers: maximal runs of overlappin
## Hash-based ("random") minimizer
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions (Golan & Shur 2025; Kille *et al.* 2023; Pan & Reinert 2024; Zheng *et al.* 2020, 2021).
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions ([Golan & Shur 2025](#ref-Golan2025-xf); [Kille *et al.* 2023](#ref-Kille2023-px); [Pan & Reinert 2024](#ref-Pan2024-hb); [Zheng *et al.* 2020](#ref-Zheng2020-ji), [2021](#ref-Zheng2021-cc)).
Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition.
@@ -56,6 +56,8 @@ $$\text{partition} = H(\text{minimizer}) \bmod 2^p$$
See [Partitioning and indexing architecture](theory-indexing_architecture) for more details.
## Bibliography
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
<div id="ref-Golan2025-xf" class="csl-entry">