From 60676795a17668abf168d905ba7be519f2dd40dc Mon Sep 17 00:00:00 2001 From: Eric Coissac Date: Sat, 12 Sep 2026 12:35:57 +0200 Subject: [PATCH] Clarify partition routing based on minimizer selection The text clarifies how partition routing is independent of minimizer selection and details the hash-based calculation for the partition index. --- theory-minimizer_selection.md | 9 ++++----- 1 file changed, 4 insertions(+), 5 deletions(-) diff --git a/theory-minimizer_selection.md b/theory-minimizer_selection.md index 06fda08..cde70ae 100644 --- a/theory-minimizer_selection.md +++ b/theory-minimizer_selection.md @@ -48,11 +48,10 @@ The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short t If the minimum period length is found to be $m$ as observed in the above table, then the sequence is aperiodic, as no shorter period can be identified. -## Partition routing is independent of minimizer selection +## Partition routing -The hash used to select a minimizer within a window (the minimum of several hash values) and the hash used to route a super-kmer to a storage partition are computed separately: +A super-kmer's partition is the low $p$ bits of its minimizer's hash: -- **Selection** uses $H$ applied to every candidate m-mer in the window, keeping the minimum. -- **Partition routing** recomputes $H$ on the single selected minimizer only, once its position is fixed. This is a hash of one specific value, not the minimum of several, so it is uniformly distributed and safe to use directly for routing. +$$\text{partition} = H(\text{minimizer}) \bmod 2^p$$ -See [Partitioning and indexing architecture](theory-indexing_architecture) for how the routing value is turned into a partition index. +the same $H$ defined above, applied to the already-selected minimizer. See [Partitioning and indexing architecture](theory-indexing_architecture) for why $p$ is chosen well below $2m$ and how this keeps partitions balanced.