From 7960afdc8a0b125307e84150fa3215e5a3212df5 Mon Sep 17 00:00:00 2001 From: Eric Coissac Date: Sat, 12 Sep 2026 11:58:04 +0200 Subject: [PATCH] Introduce hashing function to debias minimizer selection This function ensures that minimizer selection is based on avalanche-mixed bits rather than nucleotide composition, preventing low-complexity sequences from being selected as winners. --- theory-minimizer_selection.md | 14 +++++++++++++- 1 file changed, 13 insertions(+), 1 deletion(-) diff --git a/theory-minimizer_selection.md b/theory-minimizer_selection.md index 6439fcb..4aa7483 100644 --- a/theory-minimizer_selection.md +++ b/theory-minimizer_selection.md @@ -22,7 +22,19 @@ $$H(x) = \text{mix64}(x \oplus s), \quad s = \lfloor 2^{64}/\varphi \rfloor = \t $$\begin{aligned} H(x) &: \\ x &\gets& x \oplus \texttt{0x9e3779b97f4a7c15} \\ x &\gets& x \oplus (x \gg 30) \\ x &\gets& x \times \texttt{0xbf58476d1ce4e5b9} \\ x &\gets& x \oplus (x \gg 27) \\ x &\gets& x \times \texttt{0x94d049bb133111eb} \\ &\text{return}& x \end{aligned}$$ -The XOR seed avoids the finalizer's fixed point at 0 ($\text{mix64}(0) = 0$), which would otherwise make an all-A m-mer (canonical value 0) win every window comparison. +Hashing serves two distinct purposes. First, since a kmer's canonical form is a lexicographic minimum, ranking kmers by the raw value of their canonical form introduces bias into minimizer selection: it favors AT-rich composition, which in turn skews the distribution of kmers across partitions. $H$ removes this bias by making the winner depend on avalanche-mixed bits rather than on nucleotide composition, restoring a uniform distribution. + +Second, independently of that general debiasing, the specific seed $s$ matters in its own right: low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes. If any of them were the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$ ($\text{mix64}(0) = 0$ is a fixed point) — it would win far more windows than a composition-uniform analysis would predict, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. The exhaustive checks below confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$. Consequently, a kmer or super-kmer containing a low-complexity substring has no increased chance of having that substring selected as its minimizer solely because of its low complexity. + +| $m$ | argmin (canonical) | decoded sequence | minimal period | +|---|---|---|---| +| 3 | 16 | `CAA` | 3 | +| 5 | 78 | `ACATG` | 5 | +| 7 | 5512 | `CCCGAGA` | 7 | +| 9 | 108760 | `CGGGATCGA` | 9 | +| 11 | 179014 | `AAGGTGTCACG` | 11 | +| 13 | 33759044 | `GAAATACTTCACA` | 13 | +| 15 | 29869313 | `AACTACTTACCAAAC` | 15 | ## Partition routing is independent of minimizer selection