Introduce super-kmers and hash-based minimizer selection
Defines super-kmers as the fundamental unit of work, including the definition of canonical super-kmers. Implements a hash-based strategy for selecting minimizers, replacing or augmenting standard lexicographic ordering.
1 parent
5c1be9e23c
commit
331b9d330d
2 files changed
+54
-4
No files matched your search
@@ -11,7 +11,7 @@ A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k
|
|||||||
|
|
||||||
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
|
A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary.
|
||||||
|
|
||||||
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately (@Zheng2020-ji; @Golan2025-xf):
|
For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately (Golan & Shur 2025; Zheng *et al.* 2020):
|
||||||
|
|
||||||
$$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$
|
$$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$
|
||||||
|
|
||||||
@@ -22,3 +22,19 @@ For $k=31$, $m=13$ this is about 40 nucleotides; in practice super-kmers rarely
|
|||||||
A **canonical super-kmer** is the lexicographic minimum of a super-kmer and its reverse complement. When a read and its reverse complement are both encountered, they produce super-kmers that are reverse complements of each other; both reduce to the same canonical super-kmer, so a genomic region is represented once regardless of which strand was read.
|
A **canonical super-kmer** is the lexicographic minimum of a super-kmer and its reverse complement. When a read and its reverse complement are both encountered, they produce super-kmers that are reverse complements of each other; both reduce to the same canonical super-kmer, so a genomic region is represented once regardless of which strand was read.
|
||||||
|
|
||||||
Super-kmers are the unit of work used throughout construction and querying: sequences are decomposed into super-kmers first, and every downstream step (partition routing, deduplication, counting) operates on them rather than on individual kmers.
|
Super-kmers are the unit of work used throughout construction and querying: sequences are decomposed into super-kmers first, and every downstream step (partition routing, deduplication, counting) operates on them rather than on individual kmers.
|
||||||
|
|
||||||
|
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
|
||||||
|
|
||||||
|
<div id="ref-Golan2025-xf" class="csl-entry">
|
||||||
|
|
||||||
|
Golan, S. & Shur, A.M. (2025). [Expected density of random minimizers](https://doi.org/10.1007/978-3-031-82670-2\_25). In: *Lecture notes in computer science*, Lecture notes in computer science. Springer Nature Switzerland, Cham, pp. 347–360.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Zheng2020-ji" class="csl-entry">
|
||||||
|
|
||||||
|
Zheng, H., Kingsford, C. & Marçais, G. (2020). [Improved design and analysis of practical minimizers](https://doi.org/10.1093/bioinformatics/btaa472). *Bioinformatics (Oxford, England)*, 36, i119–i127.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
</div>
|
||||||
@@ -8,7 +8,7 @@ The minimizer partitions a sequence into super-kmers: maximal runs of overlappin
|
|||||||
|
|
||||||
## Hash-based ("random") minimizer
|
## Hash-based ("random") minimizer
|
||||||
|
|
||||||
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions (@Zheng2020-ji; @Zheng2021-cc; @Pan2024-hb; @Kille2023-px; @Golan2025-xf).
|
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions (Golan & Shur 2025; Kille *et al.* 2023; Pan & Reinert 2024; Zheng *et al.* 2020, 2021).
|
||||||
|
|
||||||
Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition.
|
Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition.
|
||||||
|
|
||||||
@@ -24,7 +24,7 @@ The hash function is a 64-bit mixing function (splitmix64-style finalizer) appli
|
|||||||
|
|
||||||
$$H(x) = \text{mix64}(x \oplus s)$$
|
$$H(x) = \text{mix64}(x \oplus s)$$
|
||||||
|
|
||||||
```
|
``` text
|
||||||
H(x):
|
H(x):
|
||||||
x ← x ⊕ s
|
x ← x ⊕ s
|
||||||
x ← x ⊕ (x >> 30)
|
x ← x ⊕ (x >> 30)
|
||||||
@@ -37,7 +37,7 @@ H(x):
|
|||||||
The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes; if one of them happened to be the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$, since $\text{mix64}(0) = 0$ is a fixed point — it would win far more windows than the composition-uniform behavior established above predicts, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. Exhaustive checks confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$:
|
The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes; if one of them happened to be the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$, since $\text{mix64}(0) = 0$ is a fixed point — it would win far more windows than the composition-uniform behavior established above predicts, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. Exhaustive checks confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$:
|
||||||
|
|
||||||
| $m$ | argmin (canonical) | decoded sequence | minimal period |
|
| $m$ | argmin (canonical) | decoded sequence | minimal period |
|
||||||
|---|---|---|---|
|
|-----|--------------------|-------------------|----------------|
|
||||||
| 3 | 16 | `CAA` | 3 |
|
| 3 | 16 | `CAA` | 3 |
|
||||||
| 5 | 78 | `ACATG` | 5 |
|
| 5 | 78 | `ACATG` | 5 |
|
||||||
| 7 | 5512 | `CCCGAGA` | 7 |
|
| 7 | 5512 | `CCCGAGA` | 7 |
|
||||||
@@ -55,3 +55,37 @@ A super-kmer's partition is the low $p$ bits of its minimizer's hash:
|
|||||||
$$\text{partition} = H(\text{minimizer}) \bmod 2^p$$
|
$$\text{partition} = H(\text{minimizer}) \bmod 2^p$$
|
||||||
|
|
||||||
See [Partitioning and indexing architecture](theory-indexing_architecture) for more details.
|
See [Partitioning and indexing architecture](theory-indexing_architecture) for more details.
|
||||||
|
|
||||||
|
<div id="refs" class="references csl-bib-body hanging-indent" data-entry-spacing="0">
|
||||||
|
|
||||||
|
<div id="ref-Golan2025-xf" class="csl-entry">
|
||||||
|
|
||||||
|
Golan, S. & Shur, A.M. (2025). [Expected density of random minimizers](https://doi.org/10.1007/978-3-031-82670-2\_25). In: *Lecture notes in computer science*, Lecture notes in computer science. Springer Nature Switzerland, Cham, pp. 347–360.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Kille2023-px" class="csl-entry">
|
||||||
|
|
||||||
|
Kille, B., Garrison, E., Treangen, T.J. & Phillippy, A.M. (2023). [Minmers are a generalization of minimizers that enable unbiased local jaccard estimation](https://doi.org/10.1093/bioinformatics/btad512). *Bioinformatics (Oxford, England)*, 39.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Pan2024-hb" class="csl-entry">
|
||||||
|
|
||||||
|
Pan, C. & Reinert, K. (2024). [A simple refined DNA minimizer operator enables 2-fold faster computation](https://doi.org/10.1093/bioinformatics/btae045). *Bioinformatics (Oxford, England)*, 40.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Zheng2020-ji" class="csl-entry">
|
||||||
|
|
||||||
|
Zheng, H., Kingsford, C. & Marçais, G. (2020). [Improved design and analysis of practical minimizers](https://doi.org/10.1093/bioinformatics/btaa472). *Bioinformatics (Oxford, England)*, 36, i119–i127.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Zheng2021-cc" class="csl-entry">
|
||||||
|
|
||||||
|
Zheng, H., Kingsford, C. & Marçais, G. (2021). [Sequence-specific minimizers via polar sets](https://doi.org/10.1093/bioinformatics/btab313). *Bioinformatics (Oxford, England)*, 37, i187–i195.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
</div>
|
||||||
Reference in new issue
Block a user