Introduce k-mer set distance classification
Document the theoretical classification of k-mer set distances based on information retained and metric properties. Defines new presence/absence and abundance-based distances, including Jaccard, Hamming, Euclidean, and Hellinger. Clarifies metric properties of various dissimilarity measures.
1 parent
52858dfa4a
commit
23e6540dd7
1 file changed
+98
-30
@@ -1,71 +1,139 @@
|
|||||||
# Whole k-mer set distances
|
# Whole k-mer set distances
|
||||||
|
|
||||||
The k-mer composition of a genome provides a representation of its sequence content that can be compared directly between genomes. Distances in this section measure the dissimilarity between two genomes from their complete collection of k-mers.
|
The k-mer composition of a genome provides a representation of its sequence content that can be compared directly between genomes. The distances considered here quantify dissimilarity between genomes from their complete k-mer composition, independently of any particular representation, storage scheme, or indexing strategy.
|
||||||
|
|
||||||
Depending on the information retained about the k-mers, two complementary types of comparison can be considered. **Set-based distances** use only the presence or absence of each distinct k-mer, whereas **count-based distances** also account for how frequently each k-mer occurs. Count-based distances can in turn be expressed either from raw k-mer counts or from the corresponding relative frequencies.
|
Two complementary classes of comparison arise according to the information retained about k-mer occurrences ([Anderson et al. 2011](#ref-Anderson2011-rq)). **Presence/absence-based distances** consider only whether each distinct k-mer occurs in a genome, whereas **abundance-based distances** also incorporate its observed frequency. Abundance-based measures may be computed either from raw k-mer counts or from the corresponding relative frequencies.
|
||||||
|
|
||||||
Let $A$ and $B$ denote the k-mer sets of two genomes. For each distinct k-mer $i$, let $c_i^A$ and $c_i^B$ denote its counts in the two genomes, and let
|
Let (A) and (B) denote the k-mer sets of two genomes. For each distinct k-mer (i), let (c\_i^A) and (c\_i^B) denote its counts in the two genomes, and let
|
||||||
|
|
||||||
$$p_i^A=\frac{c_i^A}{\sum_j c_j^A}, \qquad p_i^B=\frac{c_i^B}{\sum_j c_j^B}$$
|
$$p_i^A=\frac{c_i^A}{\sum_j c_j^A}, \qquad p_i^B=\frac{c_i^B}{\sum_j c_j^B}$$
|
||||||
|
|
||||||
denote the corresponding relative frequencies.
|
denote the corresponding relative frequencies.
|
||||||
|
|
||||||
These distances can then be classified along two independent axes: **the information retained about k-mer occurrences** (presence/absence versus abundance) and **whether the resulting dissimilarity is a metric**, i.e. whether it satisfies the triangle inequality.
|
The resulting measures can be classified along two independent dimensions: the **information retained about k-mer occurrences** (presence/absence versus abundance) and the **metric properties** of the resulting dissimilarity, in particular whether the triangle inequality is satisfied.
|
||||||
|
|
||||||
## Presence/absence-based distances
|
## Presence/absence-based distances
|
||||||
|
|
||||||
When only presence or absence is retained, a genome is reduced to the set of k-mers it contains, and a distance can only compare which k-mers the two genomes share, not how often each occurs.
|
When only occurrence is retained, each genome is represented by its incidence vector over the k-mer space. The comparison therefore depends exclusively on the presence or absence of each distinct k-mer and is insensitive to its multiplicity.
|
||||||
|
|
||||||
Two distances built this way are true metrics, meaning they satisfy the triangle inequality — the distance from genome $A$ to genome $C$ can never exceed the sum of the distances from $A$ to $B$ and from $B$ to $C$, whichever genome $B$ is chosen as an intermediate. The **Jaccard distance** compares the two k-mer sets through the size of their intersection relative to their union:
|
The **Jaccard distance** compares two k-mer sets from the cardinalities of their intersection and union ([Jaccard 1901](#ref-Jaccard1901-mx)):
|
||||||
|
|
||||||
$$D = 1 - \frac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}$$
|
$$D = 1 - \frac{\lvert A \cap B \rvert}{\lvert A \cup B \rvert}.$$
|
||||||
|
|
||||||
The **Hamming distance** instead counts, k-mer by k-mer over the whole k-mer space, how many positions disagree on presence or absence between the two genomes:
|
Jaccard distance is a metric and therefore satisfies the triangle inequality ([Kosub 2019](#ref-Kosub2019-xt)).
|
||||||
|
|
||||||
$$D = \sum_i \mathbb{1}[a_i \ne b_i]$$
|
The **Hamming distance** is the (L^1) distance between the corresponding binary incidence vectors:
|
||||||
|
|
||||||
without normalizing by the number of positions compared. Because it is defined directly on presence/absence, Hamming distance cannot be computed from counts the way the other measures below can.
|
$$D = \sum_i \mathbb{1}[a_i \ne b_i],$$
|
||||||
|
|
||||||
A third distance, **Mash**, is derived from the Jaccard distance rather than computed independently. It approximates a per-site mutation rate between the two genomes from their Jaccard distance $J$ and the k-mer size $k$:
|
where (a\_i,b\_i\\in{0,1}) indicate the presence of k-mer (i) in the two genomes. It counts the number of k-mers for which the two incidence vectors differ, without normalization by the size of the k-mer space. As an (L^1) distance, Hamming distance is a metric. Unlike abundance-based measures, it cannot be recovered from k-mer counts alone once the presence/absence representation has been discarded.
|
||||||
|
|
||||||
$$D = -\frac{1}{k} \ln\!\left(\frac{2(1-J)}{2-J}\right)$$
|
**Mash** is not an independent set dissimilarity but a nonlinear transformation of the Jaccard similarity. If
|
||||||
|
|
||||||
(clamped to 1.0 when the Jaccard distance reaches its maximum). This transformation is monotonic — it never reorders which of two genome pairs is closer or farther — but a monotonic transformation of a metric is not guaranteed to remain a metric itself, and whether the Mash distance satisfies the triangle inequality has not been established here. It should be treated as a useful evolutionary approximation rather than a proven metric.
|
$$J = 1-D_{\mathrm{Jaccard}},$$
|
||||||
|
|
||||||
|
the Mash distance is defined as
|
||||||
|
|
||||||
|
$$D = -\frac{1}{k}\ln\left(\frac{2J}{1+J}\right) = -\frac{1}{k}\ln\left(\frac{2(1-J_{\mathrm{Jaccard}})} {2-J_{\mathrm{Jaccard}}}\right),$$
|
||||||
|
|
||||||
|
with the result clamped to (1.0) when the Jaccard similarity reaches zero. This transformation is monotone increasing with Jaccard distance and therefore preserves the ordering of pairwise dissimilarities. Monotonicity alone, however, does not guarantee preservation of the triangle inequality. Mash should consequently be regarded as an evolutionary distance estimator derived from k-mer similarity rather than assumed to be a metric.
|
||||||
|
|
||||||
## Abundance-based distances
|
## Abundance-based distances
|
||||||
|
|
||||||
When k-mer counts are retained, either as raw counts $c_i$ or as the relative frequencies $p_i$ introduced above, distances can additionally reflect how much more or less abundant a shared k-mer is between the two genomes, not merely whether it is present in both.
|
When k-mer abundances are retained, genomes are represented by non-negative vectors of k-mer counts or frequencies. Distances can then distinguish genomes that contain the same k-mers but differ in their relative abundances.
|
||||||
|
|
||||||
The **Euclidean distance**, computed on either raw counts or relative frequencies, is the ordinary straight-line distance between the two genomes' count vectors:
|
The **Euclidean distance** is the ordinary (L^2) distance between the corresponding vectors. On raw counts it is
|
||||||
|
|
||||||
$$D = \sqrt{\sum_i (c_i^A - c_i^B)^2}$$
|
$$D = \sqrt{\sum_i (c_i^A-c_i^B)^2},$$
|
||||||
|
|
||||||
Euclidean distance is always a metric, on counts as on frequencies (`relfreq-euclidean`).
|
and the same definition applied to relative frequencies gives `relfreq-euclidean`. In both cases, Euclidean distance is a metric.
|
||||||
|
|
||||||
The **Hellinger distance** is built the same way, but on the square roots of the relative frequencies rather than the frequencies themselves:
|
The **Hellinger distance** applies the Euclidean distance to the square-root-transformed relative frequencies ([Legendre & Gallagher 2001](#ref-Legendre2001-cq)):
|
||||||
|
|
||||||
$$D = \frac{1}{\sqrt{2}} \sqrt{\sum_i \left(\sqrt{p_i^A} - \sqrt{p_i^B}\right)^2}$$
|
$$D = \frac{1}{\sqrt{2}}\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2}.$$
|
||||||
|
|
||||||
and stays bounded between 0 and 1. Because it is, by construction, a rescaled Euclidean distance between the $\sqrt{p}$ vectors, it is itself a metric — as is `hellinger-euclidean`, its unnormalized variant $D = \sqrt{2} \times D_{\text{hellinger}}$, since multiplying a metric by a positive constant cannot break the triangle inequality.
|
It is bounded between 0 and 1 and is a metric because it is a positive rescaling of the Euclidean distance between the transformed vectors (\\sqrt{p^A}) and (\\sqrt{p^B}). Its unnormalized form, `hellinger-euclidean`,
|
||||||
|
|
||||||
The **Bray-Curtis dissimilarity** compares two count vectors through how much of their combined total is *not* shared:
|
$$\begin{aligned} D &=\sqrt{\sum_i\left(\sqrt{p_i^A}-\sqrt{p_i^B}\right)^2} \\ &=\sqrt{2}\,D_{\mathrm{Hellinger}}, \end{aligned}$$
|
||||||
|
|
||||||
$$D = 1 - \frac{2 \sum_i \min(c_i^A, c_i^B)}{\sum_i c_i^A + \sum_i c_i^B}$$
|
is likewise a metric, since multiplication by a positive constant preserves the triangle inequality.
|
||||||
|
|
||||||
Unlike the distances above, Bray-Curtis computed on raw counts is *not* a metric in general — it can violate the triangle inequality, a well-documented property of this measure rather than an approximation error. The reason becomes clear once the formula is rewritten, using $a+b-2\min(a,b) = \lvert a-b \rvert$, as
|
The **Bray–Curtis dissimilarity** ([Bray & Curtis 1957](#ref-Bray1957-vv)) is defined from the amount of abundance shared between two vectors:
|
||||||
|
|
||||||
$$D = \frac{\sum_i \lvert c_i^A - c_i^B \rvert}{\sum_i c_i^A + \sum_i c_i^B}$$
|
$$D = 1- \frac{2\sum_i\min(c_i^A,c_i^B)} {\sum_i c_i^A+\sum_i c_i^B}.$$
|
||||||
|
|
||||||
The denominator is the combined total abundance of both genomes, and it changes from one genome pair to the next whenever their total k-mer counts differ — it is precisely this shifting denominator that breaks the triangle inequality.
|
Using
|
||||||
|
|
||||||
That reasoning also shows exactly when Bray-Curtis *does* behave as a metric: whenever every count vector being compared shares the same fixed total $S$, the denominator becomes the constant $2S$ for every pair, and the formula reduces to
|
$$a+b-2\min(a,b)=|a-b|,$$
|
||||||
|
|
||||||
$$D = \frac{1}{2S}\sum_i \lvert c_i^A - c_i^B \rvert$$
|
this can equivalently be written as
|
||||||
|
|
||||||
a positive multiple of the $L^1$ distance, which is always a metric. Relative frequencies are exactly such a case: every genome's relative frequencies sum to $S = 1$ by construction, so **`relfreq-bray-curtis`** reduces to
|
$$D = \frac{\sum_i|c_i^A-c_i^B|} {\sum_i c_i^A+\sum_i c_i^B}.$$
|
||||||
|
|
||||||
$$D = \frac{1}{2}\sum_i \lvert p_i^A - p_i^B \rvert$$
|
For arbitrary count vectors, Bray–Curtis is not a metric ([Gower & Legendre 1986](#ref-Gower1986-ri)): the pair-dependent normalization by the combined abundance can lead to violations of the triangle inequality. This distinguishes it from the (L^1)-based distances above, whose normalization factor is fixed independently of the pair being compared.
|
||||||
|
|
||||||
the standard total variation distance between two probability distributions, a well-known metric. `bray-curtis` on raw counts and `relfreq-bray-curtis` on relative frequencies are therefore the same formula applied to two different kinds of vectors, and only one of the two — the one where every vector has the same total — is guaranteed to behave as a proper distance.
|
The metric property is recovered when all vectors under comparison have the same total abundance (S). In that case,
|
||||||
|
|
||||||
|
$$\sum_i c_i^A=\sum_i c_i^B=S$$
|
||||||
|
|
||||||
|
for every pair, and Bray–Curtis reduces to
|
||||||
|
|
||||||
|
$$D = \frac{1}{2S} \sum_i|c_i^A-c_i^B|.$$
|
||||||
|
|
||||||
|
It is therefore a positive rescaling of the (L^1) distance and is a metric on this fixed-total space.
|
||||||
|
|
||||||
|
Relative frequencies satisfy this condition by construction, since
|
||||||
|
|
||||||
|
$$\sum_i p_i^A=\sum_i p_i^B=1.$$
|
||||||
|
|
||||||
|
Consequently, `relfreq-bray-curtis` reduces to
|
||||||
|
|
||||||
|
$$D = \frac12\sum_i|p_i^A-p_i^B|,$$
|
||||||
|
|
||||||
|
which is the **total variation distance** between the two frequency distributions. It is therefore a metric.
|
||||||
|
|
||||||
|
Thus, `bray-curtis` and `relfreq-bray-curtis` have the same algebraic form but operate on different domains: the former compares count vectors whose total abundance may vary between genomes, whereas the latter compares normalized frequency vectors with a fixed total. This difference in normalization is sufficient to change the metric properties of the measure.
|
||||||
|
|
||||||
See [`phylo`](usage-phylo) for how to select a distance and the resulting output formats.
|
See [`phylo`](usage-phylo) for how to select a distance and the resulting output formats.
|
||||||
|
|
||||||
|
## Bibliography
|
||||||
|
|
||||||
|
<div id="refs" class="references csl-bib-body hanging-indent">
|
||||||
|
|
||||||
|
<div id="ref-Anderson2011-rq" class="csl-entry">
|
||||||
|
|
||||||
|
Anderson, M.J., Crist, T.O., Chase, J.M., Vellend, M., Inouye, B.D., Freestone, A.L., <span style="font-style: italic;">et al.</span> (2011). <a href="https://doi.org/10.1111/j.1461-0248.2010.01552.x">Navigating the multiple meanings of β diversity: a roadmap for the practicing ecologist</a>. <span style="font-style: italic;">Ecology Letters</span>, 14, 19–28.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Bray1957-vv" class="csl-entry">
|
||||||
|
|
||||||
|
Bray, J.R. & Curtis, J.T. (1957). <a href="https://doi.org/10.2307/1942268">An ordination of the upland forest communities of southern Wisconsin</a>. <span style="font-style: italic;">Ecological Monographs</span>, 27, 325–349.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Gower1986-ri" class="csl-entry">
|
||||||
|
|
||||||
|
Gower, J.C. & Legendre, P. (1986). <a href="https://doi.org/10.1007/bf01896809">Metric and Euclidean properties of dissimilarity coefficients</a>. <span style="font-style: italic;">Journal of Classification</span>, 3, 5–48.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Jaccard1901-mx" class="csl-entry">
|
||||||
|
|
||||||
|
Jaccard, P. (1901). <a href="https://doi.org/10.5169/seals-266450">Étude comparative de la distribution florale dans une portion des Alpes et du Jura</a>. <span style="font-style: italic;">Bulletin de la Société vaudoise des sciences naturelles</span>, 37, 547–579.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Kosub2019-xt" class="csl-entry">
|
||||||
|
|
||||||
|
Kosub, S. (2019). <a href="https://doi.org/10.1016/j.patrec.2018.12.007">A note on the triangle inequality for the Jaccard distance</a>. <span style="font-style: italic;">Pattern Recognition Letters</span>, 120, 36–38.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div id="ref-Legendre2001-cq" class="csl-entry">
|
||||||
|
|
||||||
|
Legendre, P. & Gallagher, E.D. (2001). <a href="https://doi.org/10.1007/s004420100716">Ecologically meaningful transformations for ordination of species data</a>. <span style="font-style: italic;">Oecologia</span>, 129, 271–280.
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
</div>
|
||||||
Reference in new issue
Block a user