This commit introduces new constraints for kmer size and minimizer selection, defines super-kmers, and adds extensive new command-line options for filtering, conversion, merging, and packing indices. Documentation across the codebase has also been updated.
Implement super-kmers and hash-based minimizer selection
Introduces the concept of super-kmers as the primary unit of work, capped at 256 nucleotides. Also implements a new minimizer selection strategy based on a well-distributed hash function to ensure unbiased selection.
Introduce super-kmers and hash-based minimizer selection
Defines super-kmers as the fundamental unit of work, including the definition of canonical super-kmers. Implements a hash-based strategy for selecting minimizers, replacing or augmenting standard lexicographic ordering.
Gitea/Forgejo's [[page|label]] shortlink is post-processed on rendered HTML
text nodes: it cannot match when the label contains inline formatting (a
code span splits the surrounding text into separate DOM nodes), and even
plain labels had target/text swapped from the intended order. Native
Markdown links are real AST link nodes and have neither problem.
Refine indexing architecture and command usage documentation
Updates the core documentation for kmer structure, super-kmer routing, and indexing layout. Introduces new command-line options for filtering, dumping, and conversion, clarifying the Exact/Approx trade-offs and constraints for sequence processing.
Implement advanced indexing architecture and query features
Introduces a partitioned, NUMA-aware indexing system, supports kmer filtering, evidence conversion, index packing, and advanced query capabilities including estimation and phylogenetic analysis.