From 5110e73f69933d5810180a2324e1a659602b85e0 Mon Sep 17 00:00:00 2001 From: Eric Coissac Date: Sat, 12 Sep 2026 19:08:45 +0200 Subject: [PATCH] Introduce k-mer set distance metrics MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Defines various distance metrics for k-mer sets, including presence/absence-based (Hamming), abundance-based (Euclidean), Hellinger, and Bray–Curtis dissimilarity. --- ...phylogeny-kmer_based-kmer_set_distances.md | 20 +++++++++---------- 1 file changed, 10 insertions(+), 10 deletions(-) diff --git a/theory-phylogeny-kmer_based-kmer_set_distances.md b/theory-phylogeny-kmer_based-kmer_set_distances.md index c23caf4..f8609aa 100644 --- a/theory-phylogeny-kmer_based-kmer_set_distances.md +++ b/theory-phylogeny-kmer_based-kmer_set_distances.md @@ -4,7 +4,7 @@ The k-mer composition of a genome provides a representation of its sequence cont Two complementary classes of comparison arise according to the information retained about k-mer occurrences ([Anderson et al. 2011](#ref-Anderson2011-rq)). **Presence/absence-based distances** consider only whether each distinct k-mer occurs in a genome, whereas **abundance-based distances** also incorporate its observed frequency. Abundance-based measures may be computed either from raw k-mer counts or from the corresponding relative frequencies. -Let (A) and (B) denote the k-mer sets of two genomes. For each distinct k-mer (i), let (c\_i^A) and (c\_i^B) denote its counts in the two genomes, and let +Let $A$ and $B$ denote the k-mer sets of two genomes. For each distinct k-mer $i$, let $c_i^A$ and $c_i^B$ denote its counts in the two genomes, and let $$p_i^A=\frac{c_i^A}{\sum_j c_j^A}, \qquad p_i^B=\frac{c_i^B}{\sum_j c_j^B}$$ @@ -22,11 +22,11 @@ $$D = 1 - \frac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}.$$ Jaccard distance is a metric and therefore satisfies the triangle inequality ([Kosub 2019](#ref-Kosub2019-xt)). -The **Hamming distance** is the (L^1) distance between the corresponding binary incidence vectors: +The **Hamming distance** is the $L^1$ distance between the corresponding binary incidence vectors: $$D = \sum_i \mathbb{1}[a_i \ne b_i],$$ -where (a\_i,b\_i\\in{0,1}) indicate the presence of k-mer (i) in the two genomes. It counts the number of k-mers for which the two incidence vectors differ, without normalization by the size of the k-mer space. As an (L^1) distance, Hamming distance is a metric. Unlike abundance-based measures, it cannot be recovered from k-mer counts alone once the presence/absence representation has been discarded. +where $a_i,b_i\in\{0,1\}$ indicate the presence of k-mer $i$ in the two genomes. It counts the number of k-mers for which the two incidence vectors differ, without normalization by the size of the k-mer space. As an $L^1$ distance, Hamming distance is a metric. Unlike abundance-based measures, it cannot be recovered from k-mer counts alone once the presence/absence representation has been discarded. **Mash** is not an independent set dissimilarity but a nonlinear transformation of the Jaccard similarity ([Ondov et al. 2016](#ref-Ondov2016-dq)). If @@ -42,7 +42,7 @@ with the result clamped to $1.0$ when the Jaccard similarity reaches zero. This When k-mer abundances are retained, genomes are represented by non-negative vectors of k-mer counts or frequencies. Distances can then distinguish genomes that contain the same k-mers but differ in their relative abundances. -The **Euclidean distance** is the ordinary (L^2) distance between the corresponding vectors. On raw counts it is +The **Euclidean distance** is the ordinary $L^2$ distance between the corresponding vectors. On raw counts it is $$D = \sqrt{\sum_i (c_i^A-c_i^B)^2},$$ @@ -52,7 +52,7 @@ The **Hellinger distance** applies the Euclidean distance to the square-root-tra $$D = \frac{1}{\sqrt{2}}\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2}.$$ -It is bounded between 0 and 1 and is a metric because it is a positive rescaling of the Euclidean distance between the transformed vectors (\\sqrt{p^A}) and (\\sqrt{p^B}). Its unnormalized form, `hellinger-euclidean`, +It is bounded between 0 and 1 and is a metric because it is a positive rescaling of the Euclidean distance between the transformed vectors $\sqrt{p^A}$ and $\sqrt{p^B}$. Its unnormalized form, `hellinger-euclidean`, $$\begin{aligned} D &=\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2} \\ &=\sqrt{2}\,D_{\mathrm{Hellinger}}, \end{aligned}$$ @@ -60,7 +60,7 @@ is likewise a metric, since multiplication by a positive constant preserves the The **Bray–Curtis dissimilarity** ([Bray & Curtis 1957](#ref-Bray1957-vv)) is defined from the amount of abundance shared between two vectors: -$$D = 1- \frac{2\sum_i\min(c_i^A,c_i^B)} {\sum_i c_i^A+\sum_i c_i^B}.$$ +$$D = 1-\frac{2\sum_i\min(c_i^A,c_i^B)}{\sum_i c_i^A+\sum_i c_i^B}.$$ Using @@ -70,17 +70,17 @@ this can equivalently be written as $$D = \frac{\sum_i|c_i^A-c_i^B|} {\sum_i c_i^A+\sum_i c_i^B}.$$ -For arbitrary count vectors, Bray–Curtis is not a metric ([Gower & Legendre 1986](#ref-Gower1986-ri)): the pair-dependent normalization by the combined abundance can lead to violations of the triangle inequality. This distinguishes it from the (L^1)-based distances above, whose normalization factor is fixed independently of the pair being compared. +For arbitrary count vectors, **Bray–Curtis** is not a metric ([Gower & Legendre 1986](#ref-Gower1986-ri)): the pair-dependent normalization by the combined abundance can lead to violations of the triangle inequality. This distinguishes it from the $L^1$-based distances above, whose normalization factor is fixed independently of the pair being compared. -The metric property is recovered when all vectors under comparison have the same total abundance (S). In that case, +The metric property is recovered when all vectors under comparison have the same total abundance $S$. In that case, $$\sum_i c_i^A=\sum_i c_i^B=S$$ -for every pair, and Bray–Curtis reduces to +for every pair, and **Bray–Curtis** reduces to $$D = \frac{1}{2S} \sum_i|c_i^A-c_i^B|.$$ -It is therefore a positive rescaling of the (L^1) distance and is a metric on this fixed-total space. +It is therefore a positive rescaling of the $L^1$ distance and is a metric on this fixed-total space. Relative frequencies satisfy this condition by construction, since