Introduce hashing function to debias minimizer selection

This function ensures that minimizer selection is based on avalanche-mixed bits rather than nucleotide composition, preventing low-complexity sequences from being selected as winners.
Eric Coissac committed 2026-09-12 11:58:12 +02:00
1 parent 3a2b4b467d
commit 7960afdc8a
1 file changed
+13 -1
+13 -1
@@ -22,7 +22,19 @@ $$H(x) = \text{mix64}(x \oplus s), \quad s = \lfloor 2^{64}/\varphi \rfloor = \t
$$\begin{aligned} H(x) &: \\ x &\gets& x \oplus \texttt{0x9e3779b97f4a7c15} \\ x &\gets& x \oplus (x \gg 30) \\ x &\gets& x \times \texttt{0xbf58476d1ce4e5b9} \\ x &\gets& x \oplus (x \gg 27) \\ x &\gets& x \times \texttt{0x94d049bb133111eb} \\ &\text{return}& x \end{aligned}$$
The XOR seed avoids the finalizer's fixed point at 0 ($\text{mix64}(0) = 0$), which would otherwise make an all-A m-mer (canonical value 0) win every window comparison.
Hashing serves two distinct purposes. First, since a kmer's canonical form is a lexicographic minimum, ranking kmers by the raw value of their canonical form introduces bias into minimizer selection: it favors AT-rich composition, which in turn skews the distribution of kmers across partitions. $H$ removes this bias by making the winner depend on avalanche-mixed bits rather than on nucleotide composition, restoring a uniform distribution.
Second, independently of that general debiasing, the specific seed $s$ matters in its own right: low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes. If any of them were the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$ ($\text{mix64}(0) = 0$ is a fixed point) — it would win far more windows than a composition-uniform analysis would predict, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. The exhaustive checks below confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$. Consequently, a kmer or super-kmer containing a low-complexity substring has no increased chance of having that substring selected as its minimizer solely because of its low complexity.
| $m$ | argmin (canonical) | decoded sequence | minimal period |
|---|---|---|---|
| 3 | 16 | `CAA` | 3 |
| 5 | 78 | `ACATG` | 5 |
| 7 | 5512 | `CCCGAGA` | 7 |
| 9 | 108760 | `CGGGATCGA` | 9 |
| 11 | 179014 | `AAGGTGTCACG` | 11 |
| 13 | 33759044 | `GAAATACTTCACA` | 13 |
| 15 | 29869313 | `AACTACTTACCAAAC` | 15 |
## Partition routing is independent of minimizer selection