Page:
theory phylogeny kmer_based distance_metrics
Pages
Home
architecture
formats index_layout
installation
theory encoding
theory entropy_filter
theory indexing_architecture
theory kmer_indexing encoding
theory kmer_indexing entropy_filter
theory kmer_indexing indexing_architecture
theory kmer_indexing kmers
theory kmer_indexing minimizer_selection
theory kmer_indexing superkmers
theory kmers_and_superkmers
theory minimizer_selection
theory phylogeny kmer_based distance_metrics
theory phylogeny kmer_based kmer_set_distances
theory phylogeny snp_based central_snp_model
theory phylogeny snp_based gamma_rate_heterogeneity
theory phylogeny snp_based sampling_design
theory phylogeny snp_based sankoff_model
usage annotate
usage convert
usage dump
usage estimate
usage filter
usage index_command
usage merge
usage pack
usage phylo
usage predicates
usage query
usage select
usage superkmer
usage unitig
usage utils
Clone
2
theory phylogeny kmer_based distance_metrics
Eric Coissac edited this page 2026-09-12 17:56:37 +02:00
Table of Contents
Whole-index distance metrics
These metrics operate on the kmer sets or counts stored in an index directly — no per-locus SNP calling, no sibling annex required. A/B denote the two genomes being compared; c_i^A / c_i^B are their raw counts at kmer i, p_i^A / p_i^B the corresponding relative frequencies (p_i = c_i / \sum_j c_j).
| Metric | Definition |
|---|---|
| jaccard | D = 1 - \dfrac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert} over the sets of kmers present in each genome |
| mash | derived from the Jaccard distance via D = -\dfrac{1}{k} \ln\!\left(\dfrac{2J}{1+J}\right) where J = 1 - D_{\text{jaccard}} and k is the index's kmer size; clamped to 1.0 when J \le 0 |
| hamming | number of kmer positions where presence differs between the two genomes (presence index only, not normalized): D = \sum_i \mathbb{1}[a_i \ne b_i] |
| bray-curtis | D = 1 - \dfrac{2 \sum_i \min(c_i^A, c_i^B)}{\sum_i c_i^A + \sum_i c_i^B} on raw per-kmer counts |
| relfreq-bray-curtis | the same formula computed on relative frequencies p_i instead of raw counts |
| euclidean | D = \sqrt{\sum_i (c_i^A - c_i^B)^2} on raw counts |
| relfreq-euclidean | the same formula on relative frequencies |
| hellinger | D = \dfrac{1}{\sqrt{2}} \sqrt{\sum_i \left(\sqrt{p_i^A} - \sqrt{p_i^B}\right)^2} on relative frequencies, bounded in [0, 1] |
| hellinger-euclidean | the unnormalized variant, D = \sqrt{2} \times D_{\text{hellinger}} |
hamming requires a presence/absence index; the others work on either index type.
See phylo for how to select a metric and the output formats.
Wiki sidebar
Theory
Kmer indexing
- DNA encoding
- Kmers
- Minimizer selection
- Super-kmers
- Partitioning and indexing architecture
- Low-complexity kmer filter
Phylogeny
Kmer-based
SNP-based
Usage
- superkmer
- index
- merge
- filter
- select
- query
- dump
- annotate
- phylo
- unitig
- estimate
- convert
- utils
- pack
- Predicates and taxonomy paths