Refactor Kmer indexing and add phylogenetic analysis features
This commit introduces a complete overhaul of the Kmer indexing architecture, including new support for 2-bit encoding, canonical super-kmer handling, and a partitioned processing pipeline. Additionally, it adds comprehensive phylogenetic capabilities, including multiple distance metrics, SNP correction models, and models for rate heterogeneity and sampling design.
This commit introduces new constraints for kmer size and minimizer selection, defines super-kmers, and adds extensive new command-line options for filtering, conversion, merging, and packing indices. Documentation across the codebase has also been updated.
Gitea/Forgejo's [[page|label]] shortlink is post-processed on rendered HTML
text nodes: it cannot match when the label contains inline formatting (a
code span splits the surrounding text into separate DOM nodes), and even
plain labels had target/text swapped from the intended order. Native
Markdown links are real AST link nodes and have neither problem.
The documentation now specifies that obikmer is used for counting, indexing, querying, and comparing DNA sequences, and clarifies that the tool targets individual genome datasets and is exposed through a single binary with subcommands.
Refine indexing architecture and command usage documentation
Updates the core documentation for kmer structure, super-kmer routing, and indexing layout. Introduces new command-line options for filtering, dumping, and conversion, clarifying the Exact/Approx trade-offs and constraints for sequence processing.
Add documentation and define core tool constraints
Enhance the sidebar with navigation links and a Theory section. Introduce the Home file detailing the tool's functionality, including kmer size constraints, indexing architecture using super-kmers, supported input formats, and API subcommands.