Introduce hashing function to debias minimizer selection
This function ensures that minimizer selection is based on avalanche-mixed bits rather than nucleotide composition, preventing low-complexity sequences from being selected as winners.
1 parent
3a2b4b467d
commit
7960afdc8a
1 file changed
+13
-1
@@ -22,7 +22,19 @@ $$H(x) = \text{mix64}(x \oplus s), \quad s = \lfloor 2^{64}/\varphi \rfloor = \t
|
||||
|
||||
$$\begin{aligned} H(x) &: \\ x &\gets& x \oplus \texttt{0x9e3779b97f4a7c15} \\ x &\gets& x \oplus (x \gg 30) \\ x &\gets& x \times \texttt{0xbf58476d1ce4e5b9} \\ x &\gets& x \oplus (x \gg 27) \\ x &\gets& x \times \texttt{0x94d049bb133111eb} \\ &\text{return}& x \end{aligned}$$
|
||||
|
||||
The XOR seed avoids the finalizer's fixed point at 0 ($\text{mix64}(0) = 0$), which would otherwise make an all-A m-mer (canonical value 0) win every window comparison.
|
||||
Hashing serves two distinct purposes. First, since a kmer's canonical form is a lexicographic minimum, ranking kmers by the raw value of their canonical form introduces bias into minimizer selection: it favors AT-rich composition, which in turn skews the distribution of kmers across partitions. $H$ removes this bias by making the winner depend on avalanche-mixed bits rather than on nucleotide composition, restoring a uniform distribution.
|
||||
|
||||
Second, independently of that general debiasing, the specific seed $s$ matters in its own right: low-complexity m-mers (homopolymers, short tandem repeats) are disproportionately abundant in real genomes. If any of them were the perpetual argmin of $H$ — as the all-A m-mer is when $s = 0$ ($\text{mix64}(0) = 0$ is a fixed point) — it would win far more windows than a composition-uniform analysis would predict, not because $H$ favors it, but because that pathological input keeps recurring in real sequence data. The exhaustive checks below confirm that, with this seed, the argmin is never a homopolymer or any periodic repeat, for every tested $m$. Consequently, a kmer or super-kmer containing a low-complexity substring has no increased chance of having that substring selected as its minimizer solely because of its low complexity.
|
||||
|
||||
| $m$ | argmin (canonical) | decoded sequence | minimal period |
|
||||
|---|---|---|---|
|
||||
| 3 | 16 | `CAA` | 3 |
|
||||
| 5 | 78 | `ACATG` | 5 |
|
||||
| 7 | 5512 | `CCCGAGA` | 7 |
|
||||
| 9 | 108760 | `CGGGATCGA` | 9 |
|
||||
| 11 | 179014 | `AAGGTGTCACG` | 11 |
|
||||
| 13 | 33759044 | `GAAATACTTCACA` | 13 |
|
||||
| 15 | 29869313 | `AACTACTTACCAAAC` | 15 |
|
||||
|
||||
## Partition routing is independent of minimizer selection
|
||||
|
||||
|
||||
Reference in new issue
Block a user