Introduce k-mer set distance metrics
Defines various distance metrics for k-mer sets, including presence/absence-based (Hamming), abundance-based (Euclidean), Hellinger, and Bray–Curtis dissimilarity.
1 parent
167d7155b0
commit
5110e73f69
1 file changed
+10
-10
@@ -4,7 +4,7 @@ The k-mer composition of a genome provides a representation of its sequence cont
|
|||||||
|
|
||||||
Two complementary classes of comparison arise according to the information retained about k-mer occurrences ([Anderson et al. 2011](#ref-Anderson2011-rq)). **Presence/absence-based distances** consider only whether each distinct k-mer occurs in a genome, whereas **abundance-based distances** also incorporate its observed frequency. Abundance-based measures may be computed either from raw k-mer counts or from the corresponding relative frequencies.
|
Two complementary classes of comparison arise according to the information retained about k-mer occurrences ([Anderson et al. 2011](#ref-Anderson2011-rq)). **Presence/absence-based distances** consider only whether each distinct k-mer occurs in a genome, whereas **abundance-based distances** also incorporate its observed frequency. Abundance-based measures may be computed either from raw k-mer counts or from the corresponding relative frequencies.
|
||||||
|
|
||||||
Let (A) and (B) denote the k-mer sets of two genomes. For each distinct k-mer (i), let (c\_i^A) and (c\_i^B) denote its counts in the two genomes, and let
|
Let $A$ and $B$ denote the k-mer sets of two genomes. For each distinct k-mer $i$, let $c_i^A$ and $c_i^B$ denote its counts in the two genomes, and let
|
||||||
|
|
||||||
$$p_i^A=\frac{c_i^A}{\sum_j c_j^A}, \qquad p_i^B=\frac{c_i^B}{\sum_j c_j^B}$$
|
$$p_i^A=\frac{c_i^A}{\sum_j c_j^A}, \qquad p_i^B=\frac{c_i^B}{\sum_j c_j^B}$$
|
||||||
|
|
||||||
@@ -22,11 +22,11 @@ $$D = 1 - \frac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}.$$
|
|||||||
|
|
||||||
Jaccard distance is a metric and therefore satisfies the triangle inequality ([Kosub 2019](#ref-Kosub2019-xt)).
|
Jaccard distance is a metric and therefore satisfies the triangle inequality ([Kosub 2019](#ref-Kosub2019-xt)).
|
||||||
|
|
||||||
The **Hamming distance** is the (L^1) distance between the corresponding binary incidence vectors:
|
The **Hamming distance** is the $L^1$ distance between the corresponding binary incidence vectors:
|
||||||
|
|
||||||
$$D = \sum_i \mathbb{1}[a_i \ne b_i],$$
|
$$D = \sum_i \mathbb{1}[a_i \ne b_i],$$
|
||||||
|
|
||||||
where (a\_i,b\_i\\in{0,1}) indicate the presence of k-mer (i) in the two genomes. It counts the number of k-mers for which the two incidence vectors differ, without normalization by the size of the k-mer space. As an (L^1) distance, Hamming distance is a metric. Unlike abundance-based measures, it cannot be recovered from k-mer counts alone once the presence/absence representation has been discarded.
|
where $a_i,b_i\in\{0,1\}$ indicate the presence of k-mer $i$ in the two genomes. It counts the number of k-mers for which the two incidence vectors differ, without normalization by the size of the k-mer space. As an $L^1$ distance, Hamming distance is a metric. Unlike abundance-based measures, it cannot be recovered from k-mer counts alone once the presence/absence representation has been discarded.
|
||||||
|
|
||||||
**Mash** is not an independent set dissimilarity but a nonlinear transformation of the Jaccard similarity ([Ondov et al. 2016](#ref-Ondov2016-dq)). If
|
**Mash** is not an independent set dissimilarity but a nonlinear transformation of the Jaccard similarity ([Ondov et al. 2016](#ref-Ondov2016-dq)). If
|
||||||
|
|
||||||
@@ -42,7 +42,7 @@ with the result clamped to $1.0$ when the Jaccard similarity reaches zero. This
|
|||||||
|
|
||||||
When k-mer abundances are retained, genomes are represented by non-negative vectors of k-mer counts or frequencies. Distances can then distinguish genomes that contain the same k-mers but differ in their relative abundances.
|
When k-mer abundances are retained, genomes are represented by non-negative vectors of k-mer counts or frequencies. Distances can then distinguish genomes that contain the same k-mers but differ in their relative abundances.
|
||||||
|
|
||||||
The **Euclidean distance** is the ordinary (L^2) distance between the corresponding vectors. On raw counts it is
|
The **Euclidean distance** is the ordinary $L^2$ distance between the corresponding vectors. On raw counts it is
|
||||||
|
|
||||||
$$D = \sqrt{\sum_i (c_i^A-c_i^B)^2},$$
|
$$D = \sqrt{\sum_i (c_i^A-c_i^B)^2},$$
|
||||||
|
|
||||||
@@ -52,7 +52,7 @@ The **Hellinger distance** applies the Euclidean distance to the square-root-tra
|
|||||||
|
|
||||||
$$D = \frac{1}{\sqrt{2}}\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2}.$$
|
$$D = \frac{1}{\sqrt{2}}\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2}.$$
|
||||||
|
|
||||||
It is bounded between 0 and 1 and is a metric because it is a positive rescaling of the Euclidean distance between the transformed vectors (\\sqrt{p^A}) and (\\sqrt{p^B}). Its unnormalized form, `hellinger-euclidean`,
|
It is bounded between 0 and 1 and is a metric because it is a positive rescaling of the Euclidean distance between the transformed vectors $\sqrt{p^A}$ and $\sqrt{p^B}$. Its unnormalized form, `hellinger-euclidean`,
|
||||||
|
|
||||||
$$\begin{aligned} D &=\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2} \\ &=\sqrt{2}\,D_{\mathrm{Hellinger}}, \end{aligned}$$
|
$$\begin{aligned} D &=\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2} \\ &=\sqrt{2}\,D_{\mathrm{Hellinger}}, \end{aligned}$$
|
||||||
|
|
||||||
@@ -60,7 +60,7 @@ is likewise a metric, since multiplication by a positive constant preserves the
|
|||||||
|
|
||||||
The **Bray–Curtis dissimilarity** ([Bray & Curtis 1957](#ref-Bray1957-vv)) is defined from the amount of abundance shared between two vectors:
|
The **Bray–Curtis dissimilarity** ([Bray & Curtis 1957](#ref-Bray1957-vv)) is defined from the amount of abundance shared between two vectors:
|
||||||
|
|
||||||
$$D = 1- \frac{2\sum_i\min(c_i^A,c_i^B)} {\sum_i c_i^A+\sum_i c_i^B}.$$
|
$$D = 1-\frac{2\sum_i\min(c_i^A,c_i^B)}{\sum_i c_i^A+\sum_i c_i^B}.$$
|
||||||
|
|
||||||
Using
|
Using
|
||||||
|
|
||||||
@@ -70,17 +70,17 @@ this can equivalently be written as
|
|||||||
|
|
||||||
$$D = \frac{\sum_i|c_i^A-c_i^B|} {\sum_i c_i^A+\sum_i c_i^B}.$$
|
$$D = \frac{\sum_i|c_i^A-c_i^B|} {\sum_i c_i^A+\sum_i c_i^B}.$$
|
||||||
|
|
||||||
For arbitrary count vectors, Bray–Curtis is not a metric ([Gower & Legendre 1986](#ref-Gower1986-ri)): the pair-dependent normalization by the combined abundance can lead to violations of the triangle inequality. This distinguishes it from the (L^1)-based distances above, whose normalization factor is fixed independently of the pair being compared.
|
For arbitrary count vectors, **Bray–Curtis** is not a metric ([Gower & Legendre 1986](#ref-Gower1986-ri)): the pair-dependent normalization by the combined abundance can lead to violations of the triangle inequality. This distinguishes it from the $L^1$-based distances above, whose normalization factor is fixed independently of the pair being compared.
|
||||||
|
|
||||||
The metric property is recovered when all vectors under comparison have the same total abundance (S). In that case,
|
The metric property is recovered when all vectors under comparison have the same total abundance $S$. In that case,
|
||||||
|
|
||||||
$$\sum_i c_i^A=\sum_i c_i^B=S$$
|
$$\sum_i c_i^A=\sum_i c_i^B=S$$
|
||||||
|
|
||||||
for every pair, and Bray–Curtis reduces to
|
for every pair, and **Bray–Curtis** reduces to
|
||||||
|
|
||||||
$$D = \frac{1}{2S} \sum_i|c_i^A-c_i^B|.$$
|
$$D = \frac{1}{2S} \sum_i|c_i^A-c_i^B|.$$
|
||||||
|
|
||||||
It is therefore a positive rescaling of the (L^1) distance and is a metric on this fixed-total space.
|
It is therefore a positive rescaling of the $L^1$ distance and is a metric on this fixed-total space.
|
||||||
|
|
||||||
Relative frequencies satisfy this condition by construction, since
|
Relative frequencies satisfy this condition by construction, since
|
||||||
|
|
||||||
|
|||||||
Reference in new issue
Block a user