{}
Introduces the definition of a minimizer for k-mer windows, defining the canonical form as the lexicographic minimum between the m-mer and its reverse complement. Selection is determined by a deterministic hash function that chooses the m-mer whose canonical form minimizes the hash value, including a warning about potential biases in the ordering.
This commit introduces new constraints for kmer size and minimizer selection, defines super-kmers, and adds extensive new command-line options for filtering, conversion, merging, and packing indices. Documentation across the codebase has also been updated.
Implement super-kmers and hash-based minimizer selection
Introduces the concept of super-kmers as the primary unit of work, capped at 256 nucleotides. Also implements a new minimizer selection strategy based on a well-distributed hash function to ensure unbiased selection.
Introduce super-kmers and hash-based minimizer selection
Defines super-kmers as the fundamental unit of work, including the definition of canonical super-kmers. Implements a hash-based strategy for selecting minimizers, replacing or augmenting standard lexicographic ordering.
Update documentation clarifying that partition routing is independent of minimizer selection and details how minimizer selection and super-kmer routing hashes are computed separately.
Introduces a specific 64-bit mixing hash function for minimizer selection. This function is designed to remove bias favoring AT-rich composition and accounts for low-complexity m-mers. Exhaustive checks confirm that the argmin is never a homopolymer or periodic repeat.
Introduce hashing function to debias minimizer selection
This function ensures that minimizer selection is based on avalanche-mixed bits rather than nucleotide composition, preventing low-complexity sequences from being selected as winners.
Gitea/Forgejo's [[page|label]] shortlink is post-processed on rendered HTML
text nodes: it cannot match when the label contains inline formatting (a
code span splits the surrounding text into separate DOM nodes), and even
plain labels had target/text swapped from the intended order. Native
Markdown links are real AST link nodes and have neither problem.
Refine indexing architecture and command usage documentation
Updates the core documentation for kmer structure, super-kmer routing, and indexing layout. Introduces new command-line options for filtering, dumping, and conversion, clarifying the Exact/Approx trade-offs and constraints for sequence processing.
Implement advanced indexing architecture and query features
Introduces a partitioned, NUMA-aware indexing system, supports kmer filtering, evidence conversion, index packing, and advanced query capabilities including estimation and phylogenetic analysis.