Implement super-kmers and hash-based minimizer selection
Introduces the concept of super-kmers as the primary unit of work, capped at 256 nucleotides. Also implements a new minimizer selection strategy based on a well-distributed hash function to ensure unbiased selection.
1 parent
331b9d330d
commit
c1e139c597
2 files changed
+6
-2
No files matched your search
@@ -11,7 +11,7 @@ A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k
|
|||||||
|
|
||||||
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
|
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
|
||||||
|
|
||||||
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately (Golan & Shur 2025; Zheng *et al.* 2020):
|
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately ([Golan & Shur 2025](#ref-Golan2025-xf); [Zheng *et al.* 2020](#ref-Zheng2020-ji)):
|
||||||
|
|
||||||
$$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$
|
$$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$
|
||||||
|
|
||||||
@@ -23,6 +23,8 @@ A **canonical super-kmer** is the lexicographic minimum of a super-kmer and its
|
|||||||
|
|
||||||
Super-kmers are the unit of work used throughout construction and querying: sequences are decomposed into super-kmers first, and every downstream step (partition routing, deduplication, counting) operates on them rather than on individual kmers.
|
Super-kmers are the unit of work used throughout construction and querying: sequences are decomposed into super-kmers first, and every downstream step (partition routing, deduplication, counting) operates on them rather than on individual kmers.
|
||||||
|
|
||||||
|
## Bibliography
|
||||||
|
|
||||||
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
|
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
|
||||||
|
|
||||||
<div id="ref-Golan2025-xf" class="csl-entry">
|
<div id="ref-Golan2025-xf" class="csl-entry">
|
||||||
|
|||||||
@@ -8,7 +8,7 @@ The minimizer partitions a sequence into super-kmers: maximal runs of overlappin
|
|||||||
|
|
||||||
## Hash-based ("random") minimizer
|
## Hash-based ("random") minimizer
|
||||||
|
|
||||||
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions (Golan & Shur 2025; Kille *et al.* 2023; Pan & Reinert 2024; Zheng *et al.* 2020, 2021).
|
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions ([Golan & Shur 2025](#ref-Golan2025-xf); [Kille *et al.* 2023](#ref-Kille2023-px); [Pan & Reinert 2024](#ref-Pan2024-hb); [Zheng *et al.* 2020](#ref-Zheng2020-ji), [2021](#ref-Zheng2021-cc)).
|
||||||
|
|
||||||
Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition.
|
Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition.
|
||||||
|
|
||||||
@@ -56,6 +56,8 @@ $$\text{partition} = H(\text{minimizer}) \bmod 2^p$$
|
|||||||
|
|
||||||
See [Partitioning and indexing architecture](theory-indexing_architecture) for more details.
|
See [Partitioning and indexing architecture](theory-indexing_architecture) for more details.
|
||||||
|
|
||||||
|
## Bibliography
|
||||||
|
|
||||||
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
|
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
|
||||||
|
|
||||||
<div id="ref-Golan2025-xf" class="csl-entry">
|
<div id="ref-Golan2025-xf" class="csl-entry">
|
||||||
|
|||||||
Reference in new issue
Block a user