diff --git a/theory-kmers_and_superkmers.md b/theory-kmers_and_superkmers.md index dacb1fd..7e01796 100644 --- a/theory-kmers_and_superkmers.md +++ b/theory-kmers_and_superkmers.md @@ -11,7 +11,7 @@ A **kmer** is a DNA subsequence of fixed length $k$. Two constraints apply to $k A **super-kmer** is a maximal run of consecutive, overlapping kmers from a read that share the same canonical minimizer (see [Minimizer selection](theory-minimizer_selection)). Each kmer in the run overlaps the next by $k-1$ nucleotides. A super-kmer is capped at 256 nucleotides; a longer run is split at that boundary. -For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately (Golan & Shur 2025; Zheng *et al.* 2020): +For a random minimizer of length $m$ over kmers of length $k$, the expected length of a super-kmer is approximately ([Golan & Shur 2025](#ref-Golan2025-xf); [Zheng *et al.* 2020](#ref-Zheng2020-ji)): $$L_{\text{nt}} \approx \frac{k-m+2}{2} + k - 1$$ @@ -23,6 +23,8 @@ A **canonical super-kmer** is the lexicographic minimum of a super-kmer and its Super-kmers are the unit of work used throughout construction and querying: sequences are decomposed into super-kmers first, and every downstream step (partition routing, deduplication, counting) operates on them rather than on individual kmers. +## Bibliography +