Implement hash-based minimizer selection

Introduce hashing for balanced partitioning and dynamic routing architecture.
Eric Coissac committed 2026-09-12 12:52:48 +02:00
1 parent 60676795a1
commit 5c1be9e23c
2 files changed
+4 -10

No files matched your search

+2 -8
@@ -10,18 +10,12 @@ The canonical minimizer of a super-kmer (see [Minimizer selection](theory-minimi
canonical minimizer → hash(minimizer) → p-bit value → partition index
```
The routing value is recomputed whenever it is needed (during construction and again at query time) rather than stored — it is not part of the on-disk super-kmer representation.
Within a partition, kmers are indexed as plain values via a minimal perfect hash function (see [On-disk storage](formats-index_layout)); the minimizer plays no further role once a super-kmer has reached its partition.
## Why hashing is necessary
A canonical minimizer is an m-mer ($m \in \{9, 11, 13, 15\}$), and its distribution over all possible m-mer values is not uniform — as the lexicographic minimum of a window, small values are systematically over-represented (@Zheng2020-ji; @Zheng2021-cc; @Pan2024-hb; @Kille2023-px; @Golan2025-xf). Routing directly on the raw minimizer value would therefore produce badly unbalanced partitions.
Hashing the minimizer before routing redistributes this skewed distribution uniformly across partitions. This works reliably because the number of partition-index bits $p$ is chosen well below the number of bits available in the minimizer ($2m$): even with strong bias in the minimizer distribution, the hash has enough entropy margin to absorb it, provided the number of distinct minimizers actually observed is much larger than the number of partitions.
## Parameter guidance
Even though $H$ already makes minimizer values well-distributed (see [Minimizer selection](theory-minimizer_selection)), choosing $p$ well below the number of bits available in the minimizer ($2m$) leaves a comfortable entropy margin, provided the number of distinct minimizers actually observed is much larger than the number of partitions.
| Minimizer size $m$ | Minimizer bits ($2m$) | Typical partition-index bits $p$ | Partitions |
|----|-----------|-----------|------------|
| 9 | 18 | 6–8 | 64–256 |
+2 -2
@@ -8,7 +8,7 @@ The minimizer partitions a sequence into super-kmers: maximal runs of overlappin
## Hash-based ("random") minimizer
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions.
`obikmer` selects minimizers by hash order rather than plain lexicographic order. Ordering m-mers lexicographically on their 2-bit encoding systematically favors AT-rich m-mers (an all-A m-mer always encodes to 0), which causes low-complexity regions to dominate as minimizers and produces unbalanced partitions (@Zheng2020-ji; @Zheng2021-cc; @Pan2024-hb; @Kille2023-px; @Golan2025-xf).
Instead, a well-distributed hash function $H$ is applied to the canonical (lexicographically minimal) form of each m-mer, and the m-mer with the smallest $H$ value wins. Because $H$ is a bijection with good avalanche properties, every distinct m-mer in a window has an equal chance of holding the minimum hash value, independent of its nucleotide composition.
@@ -54,4 +54,4 @@ A super-kmer's partition is the low $p$ bits of its minimizer's hash:
$$\text{partition} = H(\text{minimizer}) \bmod 2^p$$
the same $H$ defined above, applied to the already-selected minimizer. See [Partitioning and indexing architecture](theory-indexing_architecture) for why $p$ is chosen well below $2m$ and how this keeps partitions balanced.
See [Partitioning and indexing architecture](theory-indexing_architecture) for more details.