Clarify partition routing based on minimizer selection

The text clarifies how partition routing is independent of minimizer selection and details the hash-based calculation for the partition index.
Eric Coissac committed 2026-09-12 12:36:05 +02:00
1 parent 413db85400
commit 60676795a1
1 file changed
+4 -5
+4 -5
@@ -48,11 +48,10 @@ The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short t
If the minimum period length is found to be $m$ as observed in the above table, then the sequence is aperiodic, as no shorter period can be identified. If the minimum period length is found to be $m$ as observed in the above table, then the sequence is aperiodic, as no shorter period can be identified.
## Partition routing is independent of minimizer selection ## Partition routing
The hash used to select a minimizer within a window (the minimum of several hash values) and the hash used to route a super-kmer to a storage partition are computed separately: A super-kmer's partition is the low $p$ bits of its minimizer's hash:
- **Selection** uses $H$ applied to every candidate m-mer in the window, keeping the minimum. $$\text{partition} = H(\text{minimizer}) \bmod 2^p$$
- **Partition routing** recomputes $H$ on the single selected minimizer only, once its position is fixed. This is a hash of one specific value, not the minimum of several, so it is uniformly distributed and safe to use directly for routing.
See [Partitioning and indexing architecture](theory-indexing_architecture) for how the routing value is turned into a partition index. the same $H$ defined above, applied to the already-selected minimizer. See [Partitioning and indexing architecture](theory-indexing_architecture) for why $p$ is chosen well below $2m$ and how this keeps partitions balanced.