2
theory phylogeny kmer_based distance_metrics
Eric Coissac edited this page 2026-09-12 17:56:37 +02:00

Whole-index distance metrics

These metrics operate on the kmer sets or counts stored in an index directly — no per-locus SNP calling, no sibling annex required. A/B denote the two genomes being compared; c_i^A / c_i^B are their raw counts at kmer i, p_i^A / p_i^B the corresponding relative frequencies (p_i = c_i / \sum_j c_j).

Metric Definition
jaccard D = 1 - \dfrac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert} over the sets of kmers present in each genome
mash derived from the Jaccard distance via D = -\dfrac{1}{k} \ln\!\left(\dfrac{2J}{1+J}\right) where J = 1 - D_{\text{jaccard}} and k is the index's kmer size; clamped to 1.0 when J \le 0
hamming number of kmer positions where presence differs between the two genomes (presence index only, not normalized): D = \sum_i \mathbb{1}[a_i \ne b_i]
bray-curtis D = 1 - \dfrac{2 \sum_i \min(c_i^A, c_i^B)}{\sum_i c_i^A + \sum_i c_i^B} on raw per-kmer counts
relfreq-bray-curtis the same formula computed on relative frequencies p_i instead of raw counts
euclidean D = \sqrt{\sum_i (c_i^A - c_i^B)^2} on raw counts
relfreq-euclidean the same formula on relative frequencies
hellinger D = \dfrac{1}{\sqrt{2}} \sqrt{\sum_i \left(\sqrt{p_i^A} - \sqrt{p_i^B}\right)^2} on relative frequencies, bounded in [0, 1]
hellinger-euclidean the unnormalized variant, D = \sqrt{2} \times D_{\text{hellinger}}

hamming requires a presence/absence index; the others work on either index type.

See phylo for how to select a metric and the output formats.