From c1e139c597334d61d12783baea90c2a7e0b8e10c Mon Sep 17 00:00:00 2001 From: Eric Coissac Date: Sat, 12 Sep 2026 13:00:56 +0200 Subject: [PATCH] Implement super-kmers and hash-based minimizer selection Introduces the concept of super-kmers as the primary unit of work, capped at 256 nucleotides. Also implements a new minimizer selection strategy based on a well-distributed hash function to ensure unbiased selection. --- theory-kmers_and_superkmers.md | 4 +++- theory-minimizer_selection.md | 4 +++- 2 files changed, 6 insertions(+), 2 deletions(-) diff --git a/theory-kmers_and_superkmers.md b/theory-kmers_and_superkmers.md index dacb1fd..7e01796 100644 --- a/theory-kmers_and_superkmers.md +++ b/theory-kmers_and_superkmers.md @@ -11,7 +11,7 @@ A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary. -For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately (Golan & Shur 2025; Zheng *et al.* 2020): +For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately ([Golan & Shur 2025](#ref-Golan2025-xf); [Zheng *et al.* 2020](#ref-Zheng2020-ji)): $$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$ @@ -23,6 +23,8 @@ A **canonical super-kmer** is the lexicographic minimum of a super-kmer and its Super-kmers are the unit of work used throughout construction and querying: sequences are decomposed into super-kmers first, and every downstream step (partition routing, deduplication, counting) operates on them rather than on individual kmers. +## Bibliography +
diff --git a/theory-minimizer_selection.md b/theory-minimizer_selection.md index 48104ba..f171b63 100644 --- a/theory-minimizer_selection.md +++ b/theory-minimizer_selection.md @@ -8,7 +8,7 @@ The minimizer partitions a sequence into super-kmers: maximal runs of overlappin ## Hash-based ("random") minimizer -`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions (Golan & Shur 2025; Kille *et al.* 2023; Pan & Reinert 2024; Zheng *et al.* 2020, 2021). +`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions ([Golan & Shur 2025](#ref-Golan2025-xf); [Kille *et al.* 2023](#ref-Kille2023-px); [Pan & Reinert 2024](#ref-Pan2024-hb); [Zheng *et al.* 2020](#ref-Zheng2020-ji), [2021](#ref-Zheng2021-cc)). Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition. @@ -56,6 +56,8 @@ $$\text{partition} = H(\text{minimizer}) \bmod 2^p$$ See [Partitioning and indexing architecture](theory-indexing_architecture) for more details. +## Bibliography +