Add hash function for robust minimizer selection

Introduces a specific 64-bit mixing hash function for minimizer selection. This function is designed to remove bias favoring AT-rich composition and accounts for low-complexity m-mers. Exhaustive checks confirm that the argmin is never a homopolymer or periodic repeat.
Eric Coissac committed 2026-09-12 12:17:02 +02:00
1 parent 7960afdc8a
commit 612b813190
1 file changed
+15 -5
+15 -5
@@ -16,15 +16,25 @@ The canonical form used as input to $H$ is still the lexicographic minimum of fo
### Hash function ### Hash function
The hash function is a 64-bit mixing function (splitmix64-style finalizer) applied to the m-mer XORed with a fixed non-zero seed: $s$ is a fixed non-zero seed:
$$H(x) = \text{mix64}(x \oplus s), \quad s = \lfloor 2^{64}/\varphi \rfloor = \texttt{0x9e3779b97f4a7c15}$$ $$s = \lfloor 2^{64}/\varphi \rfloor = \texttt{0x9e3779b97f4a7c15}$$
$$\begin{aligned} H(x) &: \\ x &\gets& x \oplus \texttt{0x9e3779b97f4a7c15} \\ x &\gets& x \oplus (x \gg 30) \\ x &\gets& x \times \texttt{0xbf58476d1ce4e5b9} \\ x &\gets& x \oplus (x \gg 27) \\ x &\gets& x \times \texttt{0x94d049bb133111eb} \\ &\text{return}& x \end{aligned}$$ The hash function is a 64-bit mixing function (splitmix64-style finalizer) applied to the m-mer XORed with $s$:
Hashing serves two distinct purposes. First, since a kmer's canonical form is a lexicographic minimum, ranking kmers by the raw value of their canonical form introduces bias into minimizer selection: it favors AT-rich composition, which in turn skews the distribution of kmers across partitions. $H$ removes this bias by making the winner depend on avalanche-mixed bits rather than on nucleotide composition, restoring a uniform distribution. $$H(x) = \text{mix64}(x \oplus s)$$
Second, independently of that general debiasing, the specific seed $s$ matters in its own right: low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes. If any of them were the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$ ($\text{mix64}(0) = 0$ is a fixed point) — it would win far more windows than a composition-uniform analysis would predict, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. The exhaustive checks below confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$. Consequently, a kmer or super-kmer containing a low-complexity substring has no increased chance of having that substring selected as its minimizer solely because of its low complexity. ```
H(x):
x ← x ⊕ s
x ← x ⊕ (x >> 30)
x ← x × 0xbf58476d1ce4e5b9
x ← x ⊕ (x >> 27)
x ← x × 0x94d049bb133111eb
return x ⊕ (x >> 31)
```
The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes; if one of them happened to be the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$, since $\text{mix64}(0) = 0$ is a fixed point — it would win far more windows than the composition-uniform behavior established above predicts, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. Exhaustive checks confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$:
| $m$ | argmin (canonical) | decoded sequence | minimal period | | $m$ | argmin (canonical) | decoded sequence | minimal period |
|---|---|---|---| |---|---|---|---|