Table of Contents
Central-position SNP model
Family
A family is the set of up to 4 kmers sharing identical flanking sequence and differing only at the central base — the odd kmer length guarantees a single, well-defined central position (see Kmers). A family is eligible for a genome pair (i,j) only if both genomes carry exactly one of its observed forms (single-copy, unambiguous) — this excludes multi-copy and absent loci from the comparison rather than folding them into an undifferentiated "not identical" bucket the way a whole-index distance would.
Conditioning on eligibility this way restricts every comparison to loci that are directly, positively confirmed comparable in both genomes: a locus never enters the statistic because of a genuinely absent homologous region, a diverged paralogous copy, or a genome-size asymmetry — only because a real single-copy substitution (or lack of one) was observed at flanks confirmed intact in both genomes.
Substitution corrections (snp-*)
For a genome pair, let L be its total number of eligible loci, p the raw proportion of substitutions among those loci, $P$/Q the transition/transversion proportions, $Q_1$/Q_2 Kimura's two transversion categories (A↔C & G↔T vs. A↔T & C↔G), $P_1$/P_2 the purine (A↔G) / pyrimidine (C↔T) transition proportions, and \pi_A,\pi_C,\pi_G,\pi_T the pair's pooled base frequencies.
snp-raw
d = p
snp-jc
d = -\frac{3}{4}\ln\!\left(1-\frac{4p}{3}\right)
snp-k2p
\begin{aligned} a_1 &= 1-2P-Q \\ a_2 &= 1-2Q \\ d &= -\frac{1}{2}\ln a_1-\frac{1}{4}\ln a_2 \end{aligned}
snp-k81
\begin{aligned} a_1 &= 1-2P-2Q_1 \\ a_2 &= 1-2P-2Q_2 \\ a_3 &= 1-2Q_1-2Q_2 \\ d &= -\frac{1}{4}\left(\ln a_1+\ln a_2+\ln a_3\right) \end{aligned}
snp-f81
\begin{aligned} E &= 1-\left(\pi_A^2+\pi_C^2+\pi_G^2+\pi_T^2\right) \\ d &= -E\ln\!\left(1-\frac{p}{E}\right) \end{aligned}
snp-t92
\begin{aligned} g &= \pi_C+\pi_G \\ w &= 2g(1-g) \\ a_1 &= 1-\frac{P}{w}-Q \\ a_2 &= 1-2Q \\ d &= -w\ln a_1-\frac{1}{2}(1-w)\ln a_2 \end{aligned}
snp-tn93
\begin{aligned} g_R &= \pi_A+\pi_G \\ g_Y &= \pi_C+\pi_T \\ k_1 &= \frac{2\pi_A\pi_G}{g_R} \\ k_2 &= \frac{2\pi_C\pi_T}{g_Y} \\ k_3 &= 2\left(g_Rg_Y-\frac{\pi_A\pi_G\,g_Y}{g_R}-\frac{\pi_C\pi_T\,g_R}{g_Y}\right) \\ w_1 &= 1-\frac{P_1}{k_1}-\frac{Q}{2g_R} \\ w_2 &= 1-\frac{P_2}{k_2}-\frac{Q}{2g_Y} \\ w_3 &= 1-\frac{Q}{2g_Rg_Y} \\ d &= -k_1\ln w_1-k_2\ln w_2-k_3\ln w_3 \end{aligned}
snp-tv — transversions only, deliberately uncorrected:
d = Q
A rate-heterogeneity correction applies to every formula above except snp-raw and snp-tv: each -\ln(x) term is replaced by \alpha\left(x^{-1/\alpha}-1\right) (same weight, same x) — see Gamma rate heterogeneity for the correction itself and how \alpha is estimated.
See phylo for how to select a snp-* value and the prerequisite index annex it requires.
Wiki sidebar
Theory
Kmer indexing
- DNA encoding
- Kmers
- Minimizer selection
- Super-kmers
- Partitioning and indexing architecture
- Low-complexity kmer filter
Phylogeny
Kmer-based
SNP-based
Usage
- superkmer
- index
- merge
- filter
- select
- query
- dump
- annotate
- phylo
- unitig
- estimate
- convert
- utils
- pack
- Predicates and taxonomy paths