From 612b813190b66aa2e5df30d13edca86ac9863c58 Mon Sep 17 00:00:00 2001 From: Eric Coissac Date: Sat, 12 Sep 2026 12:16:54 +0200 Subject: [PATCH] Add hash function for robust minimizer selection Introduces a specific 64-bit mixing hash function for minimizer selection. This function is designed to remove bias favoring AT-rich composition and accounts for low-complexity m-mers. Exhaustive checks confirm that the argmin is never a homopolymer or periodic repeat. --- theory-minimizer_selection.md | 20 +++++++++++++++----- 1 file changed, 15 insertions(+), 5 deletions(-) diff --git a/theory-minimizer_selection.md b/theory-minimizer_selection.md index 4aa7483..87f5500 100644 --- a/theory-minimizer_selection.md +++ b/theory-minimizer_selection.md @@ -16,15 +16,25 @@ The canonical form used as input to $H$ is still the lexicographic minimum of fo ### Hash function -The hash function is a 64-bit mixing function (splitmix64-style finalizer) applied to the m-mer XORed with a fixed non-zero seed: +$s$ is a fixed non-zero seed: -$$H(x) = \text{mix64}(x \oplus s), \quad s = \lfloor 2^{64}/\varphi \rfloor = \texttt{0x9e3779b97f4a7c15}$$ +$$s = \lfloor 2^{64}/\varphi \rfloor = \texttt{0x9e3779b97f4a7c15}$$ -$$\begin{aligned} H(x) &: \\ x &\gets& x \oplus \texttt{0x9e3779b97f4a7c15} \\ x &\gets& x \oplus (x \gg 30) \\ x &\gets& x \times \texttt{0xbf58476d1ce4e5b9} \\ x &\gets& x \oplus (x \gg 27) \\ x &\gets& x \times \texttt{0x94d049bb133111eb} \\ &\text{return}& x \end{aligned}$$ +The hash function is a 64-bit mixing function (splitmix64-style finalizer) applied to the m-mer XORed with $s$: -Hashing serves two distinct purposes. First, since a kmer's canonical form is a lexicographic minimum, ranking kmers by the raw value of their canonical form introduces bias into minimizer selection: it favors AT-rich composition, which in turn skews the distribution of kmers across partitions. $H$ removes this bias by making the winner depend on avalanche-mixed bits rather than on nucleotide composition, restoring a uniform distribution. +$$H(x) = \text{mix64}(x \oplus s)$$ -Second, independently of that general debiasing, the specific seed $s$ matters in its own right: low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes. If any of them were the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$ ($\text{mix64}(0) = 0$ is a fixed point) — it would win far more windows than a composition-uniform analysis would predict, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. The exhaustive checks below confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$. Consequently, a kmer or super-kmer containing a low-complexity substring has no increased chance of having that substring selected as its minimizer solely because of its low complexity. +``` +H(x): + x ← x ⊕ s + x ← x ⊕ (x >> 30) + x ← x × 0xbf58476d1ce4e5b9 + x ← x ⊕ (x >> 27) + x ← x × 0x94d049bb133111eb + return x ⊕ (x >> 31) +``` + +The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes; if one of them happened to be the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$, since $\text{mix64}(0) = 0$ is a fixed point — it would win far more windows than the composition-uniform behavior established above predicts, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. Exhaustive checks confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$: | $m$ | argmin (canonical) | decoded sequence | minimal period | |---|---|---|---|