Clarify partition routing based on minimizer selection
The text clarifies how partition routing is independent of minimizer selection and details the hash-based calculation for the partition index.
1 parent
413db85400
commit
60676795a1
1 file changed
+4
-5
@@ -48,11 +48,10 @@ The choice of $s$ is not arbitrary. Low-complexity m-mers (homopolymers, short t
|
|||||||
|
|
||||||
If the minimum period length is found to be $m$ as observed in the above table, then the sequence is aperiodic, as no shorter period can be identified.
|
If the minimum period length is found to be $m$ as observed in the above table, then the sequence is aperiodic, as no shorter period can be identified.
|
||||||
|
|
||||||
## Partition routing is independent of minimizer selection
|
## Partition routing
|
||||||
|
|
||||||
The hash used to select a minimizer within a window (the minimum of several hash values) and the hash used to route a super-kmer to a storage partition are computed separately:
|
A super-kmer's partition is the low $p$ bits of its minimizer's hash:
|
||||||
|
|
||||||
- **Selection** uses $H$ applied to every candidate m-mer in the window, keeping the minimum.
|
$$\text{partition} = H(\text{minimizer}) \bmod 2^p$$
|
||||||
- **Partition routing** recomputes $H$ on the single selected minimizer only, once its position is fixed. This is a hash of one specific value, not the minimum of several, so it is uniformly distributed and safe to use directly for routing.
|
|
||||||
|
|
||||||
See [Partitioning and indexing architecture](theory-indexing_architecture) for how the routing value is turned into a partition index.
|
the same $H$ defined above, applied to the already-selected minimizer. See [Partitioning and indexing architecture](theory-indexing_architecture) for why $p$ is chosen well below $2m$ and how this keeps partitions balanced.
|
||||||
Reference in new issue
Block a user