Introduce k-mer set distance metrics

Defines various distance metrics for k-mer sets, including presence/absence-based (Hamming), abundance-based (Euclidean), Hellinger, and Bray–Curtis dissimilarity.
Eric Coissac committed 2026-09-12 19:08:53 +02:00
1 parent 167d7155b0
commit 5110e73f69
1 file changed
+9 -9
+9 -9
@@ -4,7 +4,7 @@ The k-mer composition of a genome provides a representation of its sequence cont
Two complementary classes of comparison arise according to the information retained about k-mer occurrences ([Anderson et al. 2011](#ref-Anderson2011-rq)). **Presence/absence-based distances** consider only whether each distinct k-mer occurs in a genome, whereas **abundance-based distances** also incorporate its observed frequency. Abundance-based measures may be computed either from raw k-mer counts or from the corresponding relative frequencies. Two complementary classes of comparison arise according to the information retained about k-mer occurrences ([Anderson et al. 2011](#ref-Anderson2011-rq)). **Presence/absence-based distances** consider only whether each distinct k-mer occurs in a genome, whereas **abundance-based distances** also incorporate its observed frequency. Abundance-based measures may be computed either from raw k-mer counts or from the corresponding relative frequencies.
Let (A) and (B) denote the k-mer sets of two genomes. For each distinct k-mer (i), let (c\_i^A) and (c\_i^B) denote its counts in the two genomes, and let Let $A$ and $B$ denote the k-mer sets of two genomes. For each distinct k-mer $i$, let $c_i^A$ and $c_i^B$ denote its counts in the two genomes, and let
$$p_i^A=\frac{c_i^A}{\sum_j c_j^A}, \qquad p_i^B=\frac{c_i^B}{\sum_j c_j^B}$$ $$p_i^A=\frac{c_i^A}{\sum_j c_j^A}, \qquad p_i^B=\frac{c_i^B}{\sum_j c_j^B}$$
@@ -22,11 +22,11 @@ $$D = 1 - \frac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}.$$
Jaccard distance is a metric and therefore satisfies the triangle inequality ([Kosub 2019](#ref-Kosub2019-xt)). Jaccard distance is a metric and therefore satisfies the triangle inequality ([Kosub 2019](#ref-Kosub2019-xt)).
The **Hamming distance** is the (L^1) distance between the corresponding binary incidence vectors: The **Hamming distance** is the $L^1$ distance between the corresponding binary incidence vectors:
$$D = \sum_i \mathbb{1}[a_i \ne b_i],$$ $$D = \sum_i \mathbb{1}[a_i \ne b_i],$$
where (a\_i,b\_i\\in{0,1}) indicate the presence of k-mer (i) in the two genomes. It counts the number of k-mers for which the two incidence vectors differ, without normalization by the size of the k-mer space. As an (L^1) distance, Hamming distance is a metric. Unlike abundance-based measures, it cannot be recovered from k-mer counts alone once the presence/absence representation has been discarded. where $a_i,b_i\in\{0,1\}$ indicate the presence of k-mer $i$ in the two genomes. It counts the number of k-mers for which the two incidence vectors differ, without normalization by the size of the k-mer space. As an $L^1$ distance, Hamming distance is a metric. Unlike abundance-based measures, it cannot be recovered from k-mer counts alone once the presence/absence representation has been discarded.
**Mash** is not an independent set dissimilarity but a nonlinear transformation of the Jaccard similarity ([Ondov et al. 2016](#ref-Ondov2016-dq)). If **Mash** is not an independent set dissimilarity but a nonlinear transformation of the Jaccard similarity ([Ondov et al. 2016](#ref-Ondov2016-dq)). If
@@ -42,7 +42,7 @@ with the result clamped to $1.0$ when the Jaccard similarity reaches zero. This
When k-mer abundances are retained, genomes are represented by non-negative vectors of k-mer counts or frequencies. Distances can then distinguish genomes that contain the same k-mers but differ in their relative abundances. When k-mer abundances are retained, genomes are represented by non-negative vectors of k-mer counts or frequencies. Distances can then distinguish genomes that contain the same k-mers but differ in their relative abundances.
The **Euclidean distance** is the ordinary (L^2) distance between the corresponding vectors. On raw counts it is The **Euclidean distance** is the ordinary $L^2$ distance between the corresponding vectors. On raw counts it is
$$D = \sqrt{\sum_i (c_i^A-c_i^B)^2},$$ $$D = \sqrt{\sum_i (c_i^A-c_i^B)^2},$$
@@ -52,7 +52,7 @@ The **Hellinger distance** applies the Euclidean distance to the square-root-tra
$$D = \frac{1}{\sqrt{2}}\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2}.$$ $$D = \frac{1}{\sqrt{2}}\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2}.$$
It is bounded between 0 and 1 and is a metric because it is a positive rescaling of the Euclidean distance between the transformed vectors (\\sqrt{p^A}) and (\\sqrt{p^B}). Its unnormalized form, `hellinger-euclidean`, It is bounded between 0 and 1 and is a metric because it is a positive rescaling of the Euclidean distance between the transformed vectors $\sqrt{p^A}$ and $\sqrt{p^B}$. Its unnormalized form, `hellinger-euclidean`,
$$\begin{aligned} D &=\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2} \\ &=\sqrt{2}\,D_{\mathrm{Hellinger}}, \end{aligned}$$ $$\begin{aligned} D &=\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2} \\ &=\sqrt{2}\,D_{\mathrm{Hellinger}}, \end{aligned}$$
@@ -70,17 +70,17 @@ this can equivalently be written as
$$D = \frac{\sum_i|c_i^A-c_i^B|} {\sum_i c_i^A+\sum_i c_i^B}.$$ $$D = \frac{\sum_i|c_i^A-c_i^B|} {\sum_i c_i^A+\sum_i c_i^B}.$$
For arbitrary count vectors, Bray–Curtis is not a metric ([Gower & Legendre 1986](#ref-Gower1986-ri)): the pair-dependent normalization by the combined abundance can lead to violations of the triangle inequality. This distinguishes it from the (L^1)-based distances above, whose normalization factor is fixed independently of the pair being compared. For arbitrary count vectors, **Bray–Curtis** is not a metric ([Gower & Legendre 1986](#ref-Gower1986-ri)): the pair-dependent normalization by the combined abundance can lead to violations of the triangle inequality. This distinguishes it from the $L^1$-based distances above, whose normalization factor is fixed independently of the pair being compared.
The metric property is recovered when all vectors under comparison have the same total abundance (S). In that case, The metric property is recovered when all vectors under comparison have the same total abundance $S$. In that case,
$$\sum_i c_i^A=\sum_i c_i^B=S$$ $$\sum_i c_i^A=\sum_i c_i^B=S$$
for every pair, and Bray–Curtis reduces to for every pair, and **Bray–Curtis** reduces to
$$D = \frac{1}{2S} \sum_i|c_i^A-c_i^B|.$$ $$D = \frac{1}{2S} \sum_i|c_i^A-c_i^B|.$$
It is therefore a positive rescaling of the (L^1) distance and is a metric on this fixed-total space. It is therefore a positive rescaling of the $L^1$ distance and is a metric on this fixed-total space.
Relative frequencies satisfy this condition by construction, since Relative frequencies satisfy this condition by construction, since