2026-07-10 09:47:19 +02:00
|
|
|
# Central-position SNP distance (discussion)
|
|
|
|
|
|
|
|
|
|
Not implemented. Design discussion for a substitution-rate estimator that
|
|
|
|
|
observes SNPs directly from paired-genome k-mer comparison, as an alternative
|
|
|
|
|
to Mash's Poisson-Jaccard inference (see [obicompactvec](../implementation/obicompactvec.md)
|
|
|
|
|
for the implemented Jaccard/Mash distances).
|
|
|
|
|
|
|
|
|
|
## Motivation
|
|
|
|
|
|
2026-08-10 12:35:27 +02:00
|
|
|
**Primary intent: restrict the comparison to what is actually comparable.**
|
|
|
|
|
Mash's Jaccard is computed over the **union** of both genomes' k-mer content:
|
|
|
|
|
anything not identically shared is folded into a single undifferentiated
|
|
|
|
|
mass, whether the cause is a point substitution, a genuinely absent
|
|
|
|
|
homologous region (lineage-specific content, gene-family expansion, HGT,
|
|
|
|
|
genome-size asymmetry), or a diverged paralogous copy. The model then
|
|
|
|
|
back-infers a single mutation rate from that mass, silently attributing
|
|
|
|
|
non-homology to mutation. The central-SNP approach instead conditions every
|
|
|
|
|
comparison on local, positive evidence of homology: a locus only enters the
|
|
|
|
|
statistic if its `2m` flanking bases (`m = (k-1)/2`) are found intact in
|
|
|
|
|
*both* genomes — genuinely absent or non-homologous content is excluded from
|
|
|
|
|
the comparison entirely (neither numerator nor denominator), rather than
|
|
|
|
|
silently counted as divergence. This is a conditioning on comparability, not
|
|
|
|
|
just a richer summary statistic — see "Statistic and correspondence with
|
|
|
|
|
`shared`" below for how it plays out against genome-size asymmetry and
|
|
|
|
|
diverged gene families, and "Heterozygosity, ploidy, and consensus-assembly
|
|
|
|
|
inputs" for the corresponding paralogy/heterozygosity filter.
|
|
|
|
|
|
|
|
|
|
**Secondary benefit: access to the substitution's nature.** Because the
|
|
|
|
|
central base of an odd-k window is directly observable once the flanks are
|
|
|
|
|
confirmed conserved, this also yields more than a rate — the
|
|
|
|
|
transition/transversion split — enabling classical corrected distances
|
|
|
|
|
(Jukes-Cantor, Kimura 2-parameter, LogDet) that a single Jaccard scalar
|
|
|
|
|
cannot support.
|
2026-07-10 09:47:19 +02:00
|
|
|
|
|
|
|
|
## Statistic and correspondence with `shared`
|
|
|
|
|
|
|
|
|
|
A genomic position `p` is covered by `k` overlapping k-mer windows. Requiring
|
|
|
|
|
the substitution to sit at the window's **center** makes exactly one window
|
|
|
|
|
per SNP eligible — a 1:1 correspondence between SNP and center-neighbor k-mer
|
|
|
|
|
pair, avoiding the ~k-fold overcount of an any-position neighbor search.
|
|
|
|
|
|
2026-08-10 12:35:27 +02:00
|
|
|
A locus with a fully conserved `k`-window (flanks **and** center) is an
|
|
|
|
|
exact-shared k-mer at that locus; a locus with conserved flanks but a
|
|
|
|
|
substituted center is a "central SNP". Both count each locus exactly once, in
|
|
|
|
|
matching units:
|
2026-07-10 09:47:19 +02:00
|
|
|
|
|
|
|
|
```
|
|
|
|
|
p_hat[i,j] = SNP[i,j] / (SNP[i,j] + shared[i,j])
|
|
|
|
|
```
|
|
|
|
|
|
2026-08-10 12:35:27 +02:00
|
|
|
`p_hat` is `P(center substituted | 2m flanks conserved)`. `shared[i,j]` here
|
|
|
|
|
is **not** the general-purpose `shared_kmers` matrix used by Jaccard/Mash
|
|
|
|
|
(`--shared-kmers`, `BitPartials::partial_jaccard` /
|
|
|
|
|
`CountPartials::partial_threshold_jaccard`) — that matrix counts raw k-mer
|
|
|
|
|
identity with no per-genome copy-number constraint, whereas `p_hat`'s
|
|
|
|
|
denominator applies the eligibility rule defined below (raw or
|
|
|
|
|
paralogy-filtered). Both `SNP` and `shared` are accumulated by the same
|
|
|
|
|
sweep, from the same per-locus candidate set (source k-mer + 3 variants),
|
|
|
|
|
under the same eligibility rule — see "Locus eligibility" below, and
|
|
|
|
|
"Heterozygosity, ploidy, and consensus-assembly inputs" for why the
|
|
|
|
|
copy-number constraint matters and what it costs.
|
2026-07-10 09:47:19 +02:00
|
|
|
|
|
|
|
|
**Canonical invariance**: for odd k, the central position maps to itself under
|
|
|
|
|
reverse-complement (`m -> k-1-m = m`, base complemented). A transition maps to
|
|
|
|
|
a transition, a transversion to a transversion — the transition/transversion
|
|
|
|
|
split is well-defined in canonical space.
|
|
|
|
|
|
2026-08-10 17:21:04 +02:00
|
|
|
### Definitions: family, and the canonical form of a family
|
|
|
|
|
|
|
|
|
|
**Family.** The family of a k-mer `x` is the set of (up to) 4 k-mers sharing
|
|
|
|
|
`x`'s `2m` flanking bases, differing only at the central base `m`. Membership
|
|
|
|
|
is a property of the flank pattern, not of `x` itself: any of the 4 possible
|
|
|
|
|
central substitutions belongs to the same family.
|
|
|
|
|
|
|
|
|
|
**`central_canonical_neighbors()`** (`obikseq`, `CanonicalKmerOf::central_canonical_neighbors`)
|
|
|
|
|
generates all 4 members from any one of them (observed or not), each
|
|
|
|
|
independently canonicalised (`.canonical()`, i.e. `min(kmer, revcomp(kmer))`).
|
|
|
|
|
This independent canonicalisation is necessary because a central substitution
|
|
|
|
|
can flip which orientation is lexicographically smaller — two members of the
|
|
|
|
|
same family can end up canonicalised in *different* orientations. Despite
|
|
|
|
|
that, the **set** of 4 resulting canonical k-mers is invariant: calling
|
|
|
|
|
`central_canonical_neighbors()` on any member of a family — present in the
|
|
|
|
|
index or not — yields the same 4 values. This is relied upon throughout the
|
|
|
|
|
rest of this document.
|
|
|
|
|
|
|
|
|
|
**Canonical form of a family.** Because orientation can differ member to
|
|
|
|
|
member, "which of the 4 is the reference" cannot be defined relative to
|
|
|
|
|
*whichever member happened to be visited first*, nor relative to the
|
|
|
|
|
minorant (see below) — both are data-dependent (they depend on what is
|
|
|
|
|
actually observed), so using either as the reference would make the
|
|
|
|
|
reference itself vary depending on what happens to be present in a given
|
|
|
|
|
index. Instead: **the canonical form of a family is, by definition, the
|
|
|
|
|
member whose own central base — read in its own already-canonical
|
|
|
|
|
orientation — is `A`.** This is well-defined for every family, computed
|
|
|
|
|
purely from the flank pattern, whether or not that specific member (or any
|
|
|
|
|
member at all) is actually observed anywhere in the index. Concretely: call
|
|
|
|
|
`central_canonical_neighbors()` on any member (observed or not) to get the
|
|
|
|
|
family's 4 canonical forms; the one among them whose own centre nucleotide is
|
|
|
|
|
`A` is the family's canonical form. The other 3 (`C`, `G`, `T`) are labelled
|
|
|
|
|
relative to *that* fixed reference, not relative to the calling member's own
|
|
|
|
|
orientation.
|
|
|
|
|
|
|
|
|
|
**Consequence for the minorant.** With this fixed A-referenced labelling,
|
|
|
|
|
`minorant` (the smallest raw encoding among the family's *observed* members,
|
|
|
|
|
introduced further below) becomes directly computable rather than needing to
|
|
|
|
|
be tracked as extra state: regenerate the family's 4 canonical forms from
|
|
|
|
|
any member's own k-mer (cheap, no lookup), compare the raw encodings of
|
|
|
|
|
whichever are marked present, and take the smallest. No separate stored bit
|
|
|
|
|
is required — see Step 2b below, where this replaces the earlier
|
|
|
|
|
minorant-bit design.
|
|
|
|
|
|
2026-08-10 12:35:27 +02:00
|
|
|
## Locus eligibility: raw definition vs. paralogy filter
|
|
|
|
|
|
|
|
|
|
For each k-mer `x` observed in genome A (source, one MPHF slot; the 3
|
|
|
|
|
central-position variants generated as in the sweep below): check whether
|
|
|
|
|
A's locus (flanks fixed) is resolvable in genome B under one of the 4
|
|
|
|
|
central forms.
|
|
|
|
|
|
|
|
|
|
**Raw / no model.** The locus counts in the denominator iff at least one of
|
|
|
|
|
the 4 forms is present in B; it counts in the numerator iff the form found in
|
|
|
|
|
B differs from A's own. No constraint on A's or B's own copy number at this
|
|
|
|
|
locus. Open question, not resolved: what if **more than one** of the 4 forms
|
|
|
|
|
is present in B simultaneously (ambiguous target — count once arbitrarily,
|
|
|
|
|
count all, or drop)? The stringent filter below sidesteps the question by
|
|
|
|
|
construction rather than answering it.
|
|
|
|
|
|
|
|
|
|
**Stringent / paralogy-aware.** The locus counts only if exactly one of the
|
|
|
|
|
4 forms is present in A **and** exactly one is present in B (`count == 1` at
|
|
|
|
|
that slot too, when a count index is available, to also exclude same-allele
|
|
|
|
|
duplicates that presence alone cannot see). This drops the raw definition's
|
|
|
|
|
ambiguous-B case automatically, at the cost of also dropping heterozygous
|
|
|
|
|
sites indiscriminately alongside true duplications (see "Heterozygosity,
|
|
|
|
|
ploidy, and consensus-assembly inputs" below).
|
|
|
|
|
|
|
|
|
|
**Rejected: parsimony-based multiset pairing for multiplicity > 1.** Rather
|
|
|
|
|
than dropping ambiguous loci, pair identical alleles between A and B first
|
|
|
|
|
(0-mutation explanation preferred), then take `min(unmatched_A, unmatched_B)`
|
|
|
|
|
as inferred SNP pairs. Rejected on two grounds: (1) circularity — selecting
|
|
|
|
|
pairs by minimal apparent divergence, then measuring divergence on those same
|
|
|
|
|
pairs, deflates the estimate by construction, not a neutral heuristic; (2)
|
|
|
|
|
the discriminating signal is a single base among 4 possible values, and the
|
|
|
|
|
flanks are *already* guaranteed identical for every candidate by
|
|
|
|
|
construction (that is how the locus was selected) — no information remains
|
|
|
|
|
in a k-mer window to tell which copy in B truly corresponds to which copy in
|
|
|
|
|
A once multiplicity > 1 on either side. Any pairing rule invents a
|
|
|
|
|
correspondence the data cannot support. Multiplicity > 1 is treated as
|
|
|
|
|
non-identifiable, not as a puzzle to solve with a heuristic.
|
|
|
|
|
|
|
|
|
|
## Heterozygosity, ploidy, and consensus-assembly inputs
|
|
|
|
|
|
|
|
|
|
A within-genome multiplicity signal (more than one of the 4 central forms
|
|
|
|
|
present at a locus) is produced identically by two distinct causes:
|
|
|
|
|
paralogous duplication and diploid/polyploid heterozygosity. K-mer data alone
|
|
|
|
|
cannot distinguish them. The one real discriminator is sequencing depth
|
|
|
|
|
(heterozygous site: total depth of the present forms ~= the genome's
|
|
|
|
|
single-copy average; duplication: ~2x or more) — but that signal only exists
|
|
|
|
|
if genome "counts" are raw-read depth (FASTQ input), not occurrence counts in
|
|
|
|
|
an assembled FASTA, where per-locus depth is not preserved.
|
|
|
|
|
|
|
|
|
|
**Magnitude is taxon- and mating-system-dependent, not universal.**
|
|
|
|
|
Heterozygosity density: mammals ~1 site / 1-1.5 kb (~0.1%); highly
|
|
|
|
|
outcrossing plants (maize, poplar) reported an order of magnitude higher
|
|
|
|
|
(~1%); self-fertilising plants (*Arabidopsis thaliana*) near zero — but with
|
|
|
|
|
a documented failure mode where segmental duplication masquerades as
|
|
|
|
|
"pseudo-heterozygosity"; fungi split between haploid vegetative stages
|
|
|
|
|
(non-issue) and dikaryotic Basidiomycetes, where two long-diverged haploid
|
|
|
|
|
nuclei coexist without fusing. The estimator's target use case (closely
|
|
|
|
|
related genomes, k=31) is exactly where the stringent filter above costs the
|
|
|
|
|
least for low-heterozygosity taxa and the most for outcrossing/dikaryotic
|
|
|
|
|
ones — no universal threshold; this is a scope caveat to document, not a
|
|
|
|
|
problem to solve generically.
|
|
|
|
|
|
|
|
|
|
**Why assembled-consensus inputs don't make measured distances wrong.**
|
|
|
|
|
Phylogenetic inputs are near-universally assemblies, not raw reads, and
|
|
|
|
|
assemblers collapse heterozygous sites to one consensus allele per
|
|
|
|
|
position — effectively an arbitrary, largely uncorrelated-between-assemblies
|
|
|
|
|
choice at each het site. This does not inject unbounded noise: standard
|
|
|
|
|
population genetics gives `d_xy = d_a + (pi_A + pi_B)/2` — the expected
|
|
|
|
|
pairwise difference between a random allele of population A and a random
|
|
|
|
|
allele of population B equals the net (fixed) divergence `d_a` plus the
|
|
|
|
|
average of the two populations' own within-population diversity `pi`.
|
|
|
|
|
Consensus flattening realises exactly this random-allele draw, so the
|
|
|
|
|
measured genome-to-genome distance is a `d_xy`-like quantity, not `d_a` —
|
|
|
|
|
inflated by heterozygosity by a well-characterised additive term, not
|
|
|
|
|
distorted unpredictably. The term is negligible when `pi << d_xy` (the common
|
|
|
|
|
case for cross-species comparisons), and becomes material precisely in the
|
|
|
|
|
two cases already flagged above: very closely related genomes (this
|
|
|
|
|
estimator's explicit target) and highly heterozygous outcrossing organisms,
|
|
|
|
|
where `pi` and `d_xy` are the same order of magnitude.
|
|
|
|
|
|
|
|
|
|
Caveat: this assumes the flattening is uncorrelated with the phylogenetic
|
|
|
|
|
signal — plausible for de novo assembly, not guaranteed for reference-guided
|
|
|
|
|
assembly biased toward one allele (e.g. the reference's) at each het site,
|
|
|
|
|
which would turn the noise term into a systematic bias toward the reference
|
|
|
|
|
lineage. Not evaluated here.
|
|
|
|
|
|
|
|
|
|
**Forward-looking implication, not part of the current design.** The
|
|
|
|
|
multiplicity > 1 signal discarded by the stringent filter is a crude
|
|
|
|
|
per-genome proxy for `pi` (under low background paralogy). If a `pi_hat` per
|
|
|
|
|
genome were tallied alongside `SnpTally`, a `d_a` correction
|
|
|
|
|
(`p_hat - mean(pi_hat_i, pi_hat_j)/2`, roughly) could recover an estimate
|
|
|
|
|
closer to net divergence instead of `d_xy` — a possible extension, not
|
|
|
|
|
scoped here.
|
|
|
|
|
|
2026-07-10 09:47:19 +02:00
|
|
|
## Sufficient statistic: 4x4 base-pair tally
|
|
|
|
|
|
|
|
|
|
Tabulating the joint distribution of `(center_i, center_j)` over conserved-flank
|
|
|
|
|
loci, per genome pair, is sufficient for every downstream correction:
|
|
|
|
|
|
|
|
|
|
| Estimator | Input | Formula |
|
|
|
|
|
|---|---|---|
|
|
|
|
|
| Raw p-distance | total off-diagonal / total | `p = SNP / (SNP + shared)` |
|
|
|
|
|
| Jukes-Cantor | p | `d = -3/4 * ln(1 - 4p/3)` |
|
|
|
|
|
| Kimura 2-parameter | transition rate P, transversion rate Q | `d = 1/2 ln(1/(1-2P-Q)) + 1/4 ln(1/(1-2Q))` |
|
|
|
|
|
| LogDet/paralinear | full 4x4 + base-composition margins | `d ~= -1/4 ln det(F)`, robust to non-stationary base composition |
|
|
|
|
|
|
2026-08-10 12:35:27 +02:00
|
|
|
JC/K2P need only the total and the transition/transversion split (the
|
|
|
|
|
diagonal collapses to a single "shared" total). LogDet needs the full 4x4,
|
|
|
|
|
already populated at no extra cost (see Step 1/2 below).
|
2026-07-10 09:47:19 +02:00
|
|
|
|
|
|
|
|
Memory for the 4x4 tally: `n^2 * 16` counters. Trivial for the project's
|
|
|
|
|
genome-scale use case (tens to hundreds of genomes); ~13 GB at n=10^4 — outside
|
|
|
|
|
scope but worth flagging if n grows.
|
|
|
|
|
|
|
|
|
|
## Biases (properties of the estimator, not defects)
|
|
|
|
|
|
|
|
|
|
1. **Conserved-flank ascertainment bias.** Only SNPs with intact `2m`-base
|
|
|
|
|
flanks are visible; window-intact probability decays as `(1-p)^{2m}`. For
|
|
|
|
|
k=31 (2m=30): 0.74 at p=1%, 0.21 at p=5%, 0.04 at p=10%. This estimator
|
|
|
|
|
targets **closely related genomes**. Under rate heterogeneity across sites
|
|
|
|
|
(universal in practice), conserved flanks correlate with slow centers, so
|
|
|
|
|
`p_hat` underestimates the genome-wide average rate — it specifically
|
|
|
|
|
estimates the substitution rate of **conserved regions**.
|
2026-08-10 12:35:27 +02:00
|
|
|
Two distinct factors are at play here, not one: `P(centre of a given
|
|
|
|
|
window is a SNP) = p` exactly, **independent of k** — a direct restatement
|
|
|
|
|
of the raw per-site rate via the bijective window<->centre-position
|
|
|
|
|
correspondence (Statistic section above), not a k-dependent quantity.
|
|
|
|
|
`(1-p)^{2m}` is the *separate*, genuinely k-dependent ascertainment factor
|
|
|
|
|
(are the flanks also intact). The two multiply:
|
|
|
|
|
`P(usable window showing a central SNP) = p * (1-p)^{2m}` — e.g. at
|
|
|
|
|
p=1/31 (~3.2%), k=31: `p * (1-p)^30 ~= 0.0323 * 0.374 ~= 1.2%`, i.e. about
|
|
|
|
|
1 window in 83, not 1 in 31 (which is only the centre-mutated fraction,
|
|
|
|
|
before requiring intact flanks).
|
2026-07-10 09:47:19 +02:00
|
|
|
2. **Bias toward isolated SNPs.** Two SNPs within k of each other disqualify
|
|
|
|
|
each other's flanks. Hypervariable regions are invisible by construction.
|
|
|
|
|
3. **Indels are invisible.** A frameshift destroys k-mer matches in a block;
|
|
|
|
|
this channel captures substitutions only. Indel divergence shows up as lost
|
|
|
|
|
shared k-mers (lower Jaccard/Mash), not as SNP signal.
|
|
|
|
|
4. **k-dependent specificity.** "A k-mer match implies common ancestry" is
|
|
|
|
|
quantitative. For a 3 Gbp genome, expected random flank-30 collisions
|
|
|
|
|
(k=31): `(3e9)^2 / 4^30 ~= 8` — negligible. At k=21: `(3e9)^2 / 4^20 ~= 2e6`
|
|
|
|
|
— no longer negligible. k=31 is safe; k<=21 is marginal to unreliable for
|
|
|
|
|
large genomes. The large k that guarantees homology is the same k that
|
|
|
|
|
shrinks the detectable-divergence window — an inherent tension.
|
|
|
|
|
|
|
|
|
|
## Implementation: avoid materializing a de Bruijn graph
|
|
|
|
|
|
|
|
|
|
A central-SNP pair is topologically a simple bubble in the colored de Bruijn
|
|
|
|
|
graph (source/sink k-mer shared, two length-k branches differing only at the
|
|
|
|
|
midpoint). Classical bubble-calling (Cortex/discoSNP-style) finds these, but
|
|
|
|
|
requires the graph — nodes plus adjacency for ~10^9 colored k-mers — resident
|
|
|
|
|
in memory. **Rejected**: prohibitive RAM for this project's scale.
|
|
|
|
|
|
|
|
|
|
A naive per-pair generalisation of variant lookup across n genomes (query each
|
|
|
|
|
non-shared k-mer's 3 central variants against every counterpart genome's
|
|
|
|
|
index) costs `O(n^2 . N . 3)` random lookups, with the same k-mer's 3 variants
|
|
|
|
|
regenerated and requeried once per counterpart genome — pure redundant work.
|
|
|
|
|
**Rejected** as the basis for an n-genome design.
|
|
|
|
|
|
|
|
|
|
## Implementation: sequential per-partition sweep (no scratch, no graph)
|
|
|
|
|
|
|
|
|
|
`KmerIndex::distance()` already opens every partition's `presence_store`/
|
|
|
|
|
`count_store` simultaneously, memory-mapped, into one `LayeredStore`
|
|
|
|
|
(`distance.rs:73-77`). "Querying another partition" is therefore not a new
|
|
|
|
|
I/O pattern to design — it is the same O(1) MPHF+evidence lookup the `query`
|
|
|
|
|
command already performs at scale. This lets the SNP tally be computed with
|
|
|
|
|
**no scratch files and no auxiliary graph**, by sweeping partitions once each
|
|
|
|
|
as a source:
|
|
|
|
|
|
|
|
|
|
1. For source partition `p`, enumerate its **distinct** k-mers (one per MPHF
|
|
|
|
|
slot; each already carries its full multi-genome presence/count vector —
|
|
|
|
|
no need to explode per (k-mer, genome) occurrence).
|
|
|
|
|
2. For each, generate the 3 central-substitution variants and **canonicalise
|
|
|
|
|
each independently** (`min(kmer, revcomp)`, exactly as any normal query) —
|
|
|
|
|
this avoids the orientation edge case a masked-flank grouping would have
|
|
|
|
|
(a substitution that flips canonical orientation is handled correctly
|
|
|
|
|
because each variant is canonicalised on its own, not inferred from a
|
|
|
|
|
fixed-orientation flank key).
|
|
|
|
|
3. Compute each variant's target partition `q` via its minimizer; batch/sort
|
|
|
|
|
the partition's outgoing variant queries by `q` for locality.
|
|
|
|
|
4. Look up each variant in `q`'s already-mmap'd MPHF+evidence; on a hit,
|
|
|
|
|
combine the source's presence vector (base `a`) with the variant's
|
|
|
|
|
presence vector (base `b`): for every `i` carrying `a` and `j` carrying
|
|
|
|
|
`b`, `tally[i,j][a,b] += 1`.
|
|
|
|
|
|
|
|
|
|
**Deduplication needs no persisted state.** Sweeping partitions in a fixed
|
|
|
|
|
order `p = 0, 1, ..., P-1` and only acting on a variant when its target
|
|
|
|
|
partition `q >= p` guarantees each unordered SNP pair is counted exactly
|
|
|
|
|
once: a pair with `q < p` was already resolved earlier, when `q` was itself
|
|
|
|
|
the source partition and `p` (being `>= q`) was a valid forward target. No
|
|
|
|
|
cross-partition flag array is needed — the sweep order *is* the
|
|
|
|
|
deduplication rule. Within the same partition (`q == p`), a lightweight
|
|
|
|
|
transient tie-break suffices: either a `#slots(p)`-bit scratch flag reset per
|
|
|
|
|
partition, or simply comparing the two k-mers' raw `u64` encodings and only
|
|
|
|
|
counting when `kmer_source < kmer_variant` — no storage at all.
|
|
|
|
|
|
|
|
|
|
This is a distinct computation stage, not a `partial_*` in the existing
|
|
|
|
|
additive-by-partition sense: step 3-4 read across partition boundaries by
|
|
|
|
|
construction, unlike the row-local `partial_jaccard`/`partial_threshold_jaccard`
|
|
|
|
|
primitives. But it requires no new index files, permanent or scratch:
|
|
|
|
|
`unitigs.bin`, `mphf.bin`, `evidence.bin`, and the presence/count columns are
|
|
|
|
|
read as-is, and the only extra memory is the current partition's small
|
|
|
|
|
outgoing-query batch (`#kmers(p) * 3`, released once `p` is done) plus the
|
|
|
|
|
persistent `tally` accumulator (`n^2 * 16` counters, see above).
|
|
|
|
|
|
|
|
|
|
**Outer loop (over source partitions `p`) must stay sequential.** Two
|
|
|
|
|
independent reasons, not just one: (a) memory — the bounded-footprint claim
|
|
|
|
|
above only holds with one partition's outgoing-query batch in flight; running
|
|
|
|
|
`T` source partitions concurrently multiplies that batch by `T`, exactly the
|
|
|
|
|
blowup the design avoids; (b) correctness — the `q >= p` deduplication rule
|
|
|
|
|
requires partitions to be claimed as sources in a fixed order; running `p1 <
|
|
|
|
|
p2` concurrently gives no guarantee `p1` has finished claiming its `q >= p1`
|
|
|
|
|
targets before `p2` starts claiming its own, breaking the "counted exactly
|
|
|
|
|
once" property.
|
|
|
|
|
|
|
|
|
|
**Inner loop (target-partition lookups for a fixed `p`) parallelises safely.**
|
|
|
|
|
Each lookup is an O(1) read against an already-mmap'd structure, independent
|
|
|
|
|
of the others, with no growing allocation — no memory blowup, no ordering
|
|
|
|
|
dependency between different `q`. The only shared mutable state is `tally`;
|
|
|
|
|
give each worker thread a **thread-local partial tally** (fixed `n^2 * 16`
|
|
|
|
|
size, independent of partition size) and merge into the global `tally` once
|
|
|
|
|
`p`'s inner loop completes — the same reduce-then-merge pattern Rayon already
|
|
|
|
|
uses elsewhere in this codebase to open partitions in parallel. Extra memory:
|
|
|
|
|
`#threads * n^2 * 16`, negligible (~512 MB at n~1000, 32 threads) and
|
|
|
|
|
unrelated to partition size.
|
|
|
|
|
|
|
|
|
|
**Cost**: `3 * N_distinct` MPHF lookups total across the whole index (each
|
|
|
|
|
partition swept once as source) — the same order of magnitude and the same
|
|
|
|
|
operation as running `query` over the index's entire k-mer content against
|
|
|
|
|
itself, three times. This is the tool's already-optimized regime, not a new
|
|
|
|
|
I/O profile to validate.
|
|
|
|
|
|
|
|
|
|
## Cheaper: subsampling
|
|
|
|
|
|
|
|
|
|
Since the target is a ratio, restricting the source-partition sweep to a
|
|
|
|
|
bottom-`s` hash sketch (only enumerate k-mers with `hash < threshold` as
|
|
|
|
|
sources) divides the lookup count by the sampling factor without biasing
|
|
|
|
|
`p_hat`. Mash-like tradeoff: rate estimated from a sample, not the full
|
|
|
|
|
k-mer set.
|
|
|
|
|
|
|
|
|
|
## Recommendation
|
|
|
|
|
|
|
|
|
|
Sequential per-partition sweep (Route D): reuse the already-mmap'd
|
|
|
|
|
per-partition MPHF/evidence/presence structures for O(1) variant lookups,
|
|
|
|
|
dedup via the fixed sweep-order rule (`q >= p`, plus an in-partition
|
2026-08-10 12:35:27 +02:00
|
|
|
tie-break), no scratch files, no graph materialisation. Both the SNP
|
|
|
|
|
(off-diagonal) and shared (diagonal) counts are accumulated by this same
|
|
|
|
|
sweep, under the locus-eligibility rule chosen (raw or paralogy-filtered) —
|
|
|
|
|
not reused from the general-purpose `shared_kmers` matrix, whose raw-identity
|
|
|
|
|
definition does not apply the same copy-number constraint. Distances (p, JC,
|
|
|
|
|
K2P, LogDet) as finalisations of the resulting 4x4 tally, mirroring the
|
|
|
|
|
`partial_* -> *_dist_matrix` pattern used for Jaccard/Mash/Bray-Curtis/etc.
|
2026-07-10 09:47:19 +02:00
|
|
|
|
|
|
|
|
## Detailed implementation plan
|
|
|
|
|
|
|
|
|
|
Grounded in the current codebase. File/type references are anchors, not
|
|
|
|
|
prescriptions; adjust to reality when implementing.
|
|
|
|
|
|
|
|
|
|
### Step 0 — new low-level primitives (`obikseq`)
|
|
|
|
|
|
|
|
|
|
Two helpers do not yet exist and are prerequisites:
|
|
|
|
|
|
|
|
|
|
1. **Central neighbours.** `CanonicalKmerOf<L>` already exposes
|
|
|
|
|
`left_canonical_neighbors()` / `right_canonical_neighbors()`
|
|
|
|
|
(`obikseq/src/kmer.rs`), each returning the 4 canonicalised neighbours at
|
|
|
|
|
an end position. Add `central_canonical_neighbors()` returning the 4
|
|
|
|
|
variants at position `m = (k-1)/2` (each independently canonicalised via
|
|
|
|
|
`.canonical()`). The 3 that differ from the source are the query variants;
|
|
|
|
|
skip the identity. Building on `nucleotide(i)` / the raw 2-bit layout keeps
|
|
|
|
|
it O(1).
|
|
|
|
|
2. **Lone-k-mer minimiser.** Routing a *synthetic* variant to its partition
|
|
|
|
|
needs its minimiser, but `RollingStat` (`obiskbuilder/src/rolling_stat.rs`)
|
|
|
|
|
only computes minimisers incrementally along a sequence. Add a standalone
|
|
|
|
|
`minimizer(kmer) -> Minimizer` that scans the `k-m+1` m-mer windows
|
|
|
|
|
(`PackedSeq::mmer`, `obikseq/src/packed_seq.rs`), canonicalises each, and
|
|
|
|
|
takes the min by `seq_hash()` — the same selection `RollingStat` performs,
|
|
|
|
|
evaluated once. Partition index is then
|
|
|
|
|
`(minimizer.seq_hash() & (n_partitions - 1)) as usize`, exactly as
|
|
|
|
|
`QueryBatch::from_records` (`obikmer/src/cmd/query.rs:142`); `n_partitions`
|
|
|
|
|
is a power of two so the mask is valid.
|
|
|
|
|
|
|
|
|
|
### Step 1 — the tally accumulator (`obikindex`)
|
|
|
|
|
|
|
|
|
|
A `SnpTally` holding, per genome pair, the 4x4 joint count of central bases:
|
|
|
|
|
`n * n * 4 * 4` `u64` (or a packed lower-triangular form since it is
|
|
|
|
|
symmetric). Provide `merge(&mut self, other: &SnpTally)` for the thread-local
|
|
|
|
|
reduce, and accessors yielding, per pair `(i,j)`: total off-diagonal (SNP),
|
2026-08-10 12:35:27 +02:00
|
|
|
diagonal (shared, i.e. `p_hat`'s denominator minus SNP), transition count
|
|
|
|
|
`P`, transversion count `Q`. The diagonal is always populated — it is not an
|
|
|
|
|
optional LogDet-only extra, since `p_hat`'s denominator is no longer sourced
|
|
|
|
|
from the external `shared_kmers` matrix (see "Locus eligibility" and
|
|
|
|
|
"Statistic and correspondence with `shared`" above): the source k-mer's own
|
|
|
|
|
presence/count vector, already in hand when it is enumerated, supplies the
|
|
|
|
|
diagonal entry directly, at no extra lookup cost.
|
2026-07-10 09:47:19 +02:00
|
|
|
|
|
|
|
|
### Step 2 — the sweep (`obikindex`, new `snp.rs`)
|
|
|
|
|
|
|
|
|
|
Mirror `distance.rs`: open the presence or count store per partition. But
|
|
|
|
|
instead of a per-partition `partial_*`, run the sequential source sweep:
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
for p in 0..n_partitions: # OUTER — sequential
|
|
|
|
|
open source partition p's layers (QueryLayer-style, obikpartitionner)
|
|
|
|
|
enumerate distinct canonical k-mers of p (one per MPHF slot) with their
|
|
|
|
|
presence/count vectors # column-major, as query stage 2
|
|
|
|
|
par_iter over these source k-mers: # INNER — rayon, thread-local tally
|
2026-08-10 12:35:27 +02:00
|
|
|
apply eligibility rule to the source's own vector (raw: none; # diagonal
|
|
|
|
|
stringent: exactly one of the 4 forms present in each genome) # gate
|
|
|
|
|
for i in genomes eligible with source base a:
|
|
|
|
|
for j in genomes eligible with source base a:
|
|
|
|
|
thread_tally[i,j][a,a] += 1 # diagonal — no extra lookup
|
2026-07-10 09:47:19 +02:00
|
|
|
for each of the 3 central variants:
|
|
|
|
|
q = partition_of(variant)
|
|
|
|
|
if q < p: continue # dedup: forward targets only
|
|
|
|
|
if q == p and variant <= source.raw(): continue # in-partition tie-break
|
|
|
|
|
slot = layers[q].find_slot(variant) # MphfLayer::find, mmap'd
|
|
|
|
|
if hit:
|
|
|
|
|
vb = variant presence/count vector
|
2026-08-10 12:35:27 +02:00
|
|
|
apply eligibility rule to vb (as above)
|
|
|
|
|
for i in eligible genomes with source base a:
|
|
|
|
|
for j in eligible genomes with variant base b:
|
2026-07-10 09:47:19 +02:00
|
|
|
thread_tally[i,j][a,b] += 1
|
|
|
|
|
merge thread-local tallies into global SnpTally
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
The inner lookup is precisely `QueryLayer::find_slot` +
|
|
|
|
|
`col_value(g, slot)` (`obikpartitionner/src/query_layer.rs`) — reuse or factor
|
|
|
|
|
out that path rather than reimplementing MPHF access. Enumerating "all distinct
|
|
|
|
|
k-mers of a partition with their vectors" is the `dump`/`query` stage-2
|
|
|
|
|
column-major scan already implemented in `dump_layer.rs` /
|
|
|
|
|
`query_partition_with`; factor a reusable iterator if none fits.
|
|
|
|
|
|
|
|
|
|
`presence_threshold` applies exactly as elsewhere: a genome "carries base b"
|
|
|
|
|
iff its count at that slot is `>= presence_threshold` (trivially `>= 1` for
|
|
|
|
|
presence indexes).
|
|
|
|
|
|
2026-08-10 12:35:27 +02:00
|
|
|
### Open problem (unresolved, session end — not yet fully convinced)
|
|
|
|
|
|
|
|
|
|
The `q >= p` / tie-break dedup rule in Step 2's pseudocode above is **flawed**
|
|
|
|
|
for the stringent (paralogy-filtered) eligibility rule: it only ever brings
|
|
|
|
|
two family members into view at once (the source and one looked-up variant),
|
|
|
|
|
never all four simultaneously, and which subset gets compared depends on
|
|
|
|
|
partition sweep order. "Exactly one of the 4 forms present in genome A" is a
|
|
|
|
|
whole-family property and cannot be decided correctly from a sequence of
|
|
|
|
|
pairwise, order-dependent glimpses — the pseudocode above needs revision, not
|
|
|
|
|
just the eligibility gate bolted onto it as written.
|
|
|
|
|
|
|
|
|
|
Direction discussed, **not yet settled**:
|
|
|
|
|
|
|
|
|
|
1. **Every distinct source k-mer looks up all 3 variants unconditionally**
|
|
|
|
|
(drop the `q < p` skip entirely) so that every observed family member
|
|
|
|
|
independently gathers all 4 vectors (its own + whichever of the 3
|
|
|
|
|
variants exist) at once — a whole-family, order-independent view, computed
|
|
|
|
|
redundantly once per observed member. Same total lookup order of
|
|
|
|
|
magnitude as already budgeted (`3 * N_distinct`), just organised
|
|
|
|
|
differently (no lookup actually skipped, versus the original rule which
|
|
|
|
|
skipped roughly half).
|
|
|
|
|
2. **Tie-break after gathering, not before**: only the member whose own
|
|
|
|
|
canonical encoding is the smallest *among the members actually observed*
|
|
|
|
|
(now known, since all were just looked up) writes to `SnpTally`; the
|
|
|
|
|
others silently discard their redundant computation. Deterministic,
|
|
|
|
|
order-independent — as a side effect this also removes the "outer loop
|
|
|
|
|
must stay sequential" constraint from the cost/parallelism discussion
|
|
|
|
|
above, since no step depends on partition processing order any more.
|
|
|
|
|
3. **Proposed optimisation**: precompute, once at index build time, a
|
|
|
|
|
compact global (not per-genome) annex per MPHF slot — the count of
|
|
|
|
|
*other* family members observed anywhere in the dataset (0-3). Slots with
|
|
|
|
|
count 0 (majority under low divergence and few genomes, but see the
|
|
|
|
|
scaling caveat below) need no cross-lookup at all: eligibility reduces to
|
|
|
|
|
a local `count == 1` check at that single slot, and only slots with count
|
|
|
|
|
>= 1 enter the 3-lookup sweep machinery above. Revised (see Step 2b
|
|
|
|
|
below): minorant status *is* stored alongside the count after all, on 3
|
|
|
|
|
bits rather than 2 — it comes for free from the same lookups needed to
|
|
|
|
|
count siblings, and storing it lets the sweep discard non-minorant slots
|
|
|
|
|
without re-fetching anything.
|
|
|
|
|
|
|
|
|
|
**Minorant/sibling-count relationship, worked out precisely.** "Minorant" is
|
|
|
|
|
a one-way implication from sibling count, not an equivalence: `0 siblings
|
|
|
|
|
=> minorant` (trivially — with no other observed member, the k-mer is by
|
|
|
|
|
definition the smallest of the observed set, itself alone), and its
|
|
|
|
|
contrapositive `not minorant => >= 1 sibling`. The converse does not hold:
|
|
|
|
|
being the minorant says nothing about sibling count — a minorant can have 0,
|
|
|
|
|
1, 2 or 3 siblings, all with larger encodings than itself. Consequence: this
|
|
|
|
|
confirms, as a logical necessity rather than a heuristic, that a 0-sibling
|
|
|
|
|
slot can always write its diagonal contribution with zero ambiguity and no
|
|
|
|
|
lookup (it is unconditionally its own minorant) — but it gives no shortcut
|
|
|
|
|
for the >= 1-sibling case, where minorant status still requires the actual
|
|
|
|
|
comparison of gathered encodings; sibling count alone never determines it.
|
|
|
|
|
|
|
|
|
|
**When to compute the annex, and cache invalidation.** Sibling count is a
|
|
|
|
|
property of the whole set of columns (genomes/groups) currently in the
|
|
|
|
|
index, not of any single genome — it cannot be computed correctly at
|
|
|
|
|
mono-genome build time (a family may gain siblings, or its minorant may
|
|
|
|
|
change, once more genomes are merged in later). Computing it eagerly at
|
|
|
|
|
every `merge` would also waste work on intermediate merged states nobody
|
|
|
|
|
ever queries. Instead: compute it lazily, on first `distance` call against a
|
|
|
|
|
given index, and persist the result alongside that index for subsequent
|
|
|
|
|
calls — the same lazy-derived-cache pattern `PersistentBitMatrix` already
|
|
|
|
|
uses for `Columnar` -> `Packed`. This requires no explicit invalidation for
|
|
|
|
|
`merge` or `filter` (`obikindex/src/merge.rs`, `obikmer/src/cmd/filter.rs`):
|
|
|
|
|
both only ever write to a fresh `--output` directory, never mutate an input
|
|
|
|
|
index in place, so a re-merged/re-filtered index is simply a new state with
|
|
|
|
|
no annex yet. `select --in-place` (`select_layer.rs:139-235`) is the
|
|
|
|
|
exception: it aggregates genome columns into groups (Any/All/None/Sum/Min/
|
|
|
|
|
Max) by mutating the existing index's files without changing its location.
|
|
|
|
|
It does not remove k-mer rows, but it can still change eligibility and
|
|
|
|
|
sibling counts derived from those rows (e.g. a `Sum` over several
|
|
|
|
|
single-copy genomes can read as multi-copy at the group level). Because it
|
|
|
|
|
mutates in place, **`select --in-place` must explicitly invalidate (delete
|
|
|
|
|
or mark stale) any cached sibling-count annex for that index** — the one
|
|
|
|
|
operation in the current pipeline where this doesn't happen for free.
|
|
|
|
|
|
|
|
|
|
Not yet convinced this is the right shape, and Step 2's pseudocode above has
|
|
|
|
|
not been rewritten to match — flagged for the next pass rather than resolved
|
|
|
|
|
here.
|
|
|
|
|
|
|
|
|
|
### Step 2b — sibling-count / minorant annex (consolidated plan)
|
|
|
|
|
|
2026-08-10 17:21:04 +02:00
|
|
|
Scope: only the precursor annex — not the SNP tally itself, whose Step 2
|
|
|
|
|
sweep remains unresolved above. This piece is simpler than the sweep,
|
|
|
|
|
because it writes to an independent per-slot value, not a shared
|
|
|
|
|
cross-k-mer accumulator, so it needs no dedup/ownership logic at all at this
|
|
|
|
|
stage.
|
|
|
|
|
|
|
|
|
|
**Revised annex encoding — 4-bit presence mask, not 3-bit (minorant +
|
|
|
|
|
count).** Superseded after settling the "canonical form of a family"
|
|
|
|
|
definition above. The 3-bit design (1 minorant bit + 2-bit sibling count,
|
|
|
|
|
§ below, kept for the historical record) had two problems: it discards
|
|
|
|
|
*which* variants are present (only how many), so any future consumer
|
|
|
|
|
(the SNP sweep, or a stats pass — see below) that needs to know which bases
|
|
|
|
|
exist still has to regenerate and blindly re-query all 3 candidates; and
|
|
|
|
|
the minorant bit's meaning was tied to whichever member was visited, not to
|
|
|
|
|
a fixed reference. Storing instead a **4-bit mask** — one bit per base
|
|
|
|
|
(A/C/G/T), set iff that member of the family (labelled relative to the
|
|
|
|
|
family's fixed canonical form, i.e. the member with `A` at the centre — see
|
|
|
|
|
above) is observed anywhere in the index — fixes both:
|
|
|
|
|
- **Sibling count is derived, not stored**: `siblings = popcount(mask) - 1`.
|
|
|
|
|
- **Minorant is derived, not stored**: regenerate the family's 4 canonical
|
|
|
|
|
forms from the slot's own k-mer (cheap, no lookup — see above), compare
|
|
|
|
|
the raw encodings of whichever bits are set in the mask, take the
|
|
|
|
|
smallest.
|
|
|
|
|
- **A future consumer knows exactly which variants to (re-)query** —
|
|
|
|
|
`popcount(mask) - 1` lookups instead of always 3, and it knows *which*
|
|
|
|
|
3 (or fewer) to issue, not just how many hits to expect.
|
|
|
|
|
- The all-zero value (no base present at all) is still logically
|
|
|
|
|
unreachable as a real result — the slot's *own* base is always present in
|
|
|
|
|
its own family — so it remains available as a free "not yet computed"
|
|
|
|
|
sentinel, exactly as before.
|
2026-08-10 12:35:27 +02:00
|
|
|
|
|
|
|
|
1. **Primitive.** Reuse `central_canonical_neighbors()` from Step 0
|
2026-08-10 17:21:04 +02:00
|
|
|
unchanged — the 3 canonicalised central-substitution variants of a k-mer
|
|
|
|
|
(plus the identity, i.e. all 4 members of the family — see "Definitions"
|
|
|
|
|
above).
|
|
|
|
|
2. **New annex type** (`obicompactvec`, alongside `bitmatrix.rs`): a 4-bit-
|
|
|
|
|
per-slot packed array (the presence mask above), one per partition — same
|
|
|
|
|
on-disk shape family as `PersistentBitMatrix`'s `Packed` variant, but
|
|
|
|
|
simpler (no per-genome columns, a single derived read-only value per
|
|
|
|
|
slot).
|
|
|
|
|
<details><summary>Superseded 3-bit design (historical)</summary>
|
|
|
|
|
3 bits, storing minorant status alongside sibling count directly, since
|
|
|
|
|
it came for free from the same lookups (point 3 below) — 5 real states
|
2026-08-10 12:35:27 +02:00
|
|
|
(not-minorant; minorant with 0/1/2/3 siblings) fit in 3 bits (8 states,
|
2026-08-10 17:21:04 +02:00
|
|
|
3 unused). This let the future SNP sweep discard a non-minorant slot
|
|
|
|
|
instantly, with no lookup at all. The otherwise-unreachable combination
|
|
|
|
|
"not-minorant + 0 siblings" (0 siblings always implies minorant) doubled
|
|
|
|
|
as the "not yet computed" sentinel. Replaced by the 4-bit mask above,
|
|
|
|
|
which subsumes this benefit (minorant still derivable, now for free at
|
|
|
|
|
read time rather than stored) while also fixing the "which variant"
|
|
|
|
|
blindness.
|
|
|
|
|
</details>
|
2026-08-10 12:35:27 +02:00
|
|
|
3. **Computation pass** (`obikindex`, new `siblings.rs`): **one
|
|
|
|
|
`obipipeline` run per layer, iterated sequentially over the index's
|
|
|
|
|
layers** — settled after two false starts, worth recording both.
|
|
|
|
|
- *False start 1*: "fully parallel over every partition/slot at once,
|
|
|
|
|
no ordering at all". Correctness is fine with this (sibling count and
|
|
|
|
|
minorant are order-independent, unlike the old `q >= p` dedup they
|
|
|
|
|
replace), but it reintroduces, at a larger scale, exactly the
|
|
|
|
|
memory-blowup the original Step 2 sweep's sequential-outer-loop
|
|
|
|
|
constraint existed to prevent: scattering every source partition at
|
|
|
|
|
once multiplies the in-flight outgoing-query volume by the number of
|
|
|
|
|
partitions.
|
|
|
|
|
- *False start 2*: push the layer loop itself into the pipeline (source
|
|
|
|
|
= the index's layers, a first `Flat` stage expands each layer into
|
|
|
|
|
its k-mers). `obipipeline`'s scheduler already bounds memory on its
|
|
|
|
|
own — it dispatches every item through a **shared** worker pool at
|
|
|
|
|
each stage boundary (`scheduler.rs:217-372`, `dispatch()` into a
|
|
|
|
|
common `worker_tx` queue, any free worker picks up any pending item;
|
|
|
|
|
not "one worker owns a chunk end to end"), with a biased `Select`
|
|
|
|
|
that prioritises draining items already advanced in the chain over
|
|
|
|
|
admitting new source items (`scheduler.rs:271-282`: stage results
|
|
|
|
|
outrank the source, "vider le pipeline en priorité" / "dernier
|
|
|
|
|
recours" for new data) — so bounded channel `capacity` plus this
|
|
|
|
|
drain-first bias already caps in-flight work without any external
|
|
|
|
|
sequential discipline. Correct, but it means k-mers from several
|
|
|
|
|
layers can be completing concurrently, so the sink would need to
|
|
|
|
|
track several open per-layer annex-file writers at once — real,
|
|
|
|
|
avoidable complexity.
|
|
|
|
|
- **Settled design**: keep the layer loop external and sequential —
|
|
|
|
|
not for memory (the pipeline's own `capacity`/priority mechanism
|
|
|
|
|
already provides that, for free, regardless), but so each pipeline
|
|
|
|
|
run's sink targets exactly one layer's annex file, no concurrent
|
|
|
|
|
multi-writer bookkeeping. Per layer: source = that layer's distinct
|
|
|
|
|
k-mers; a `Flat` (1->N) stage generates the 3 central variants of a
|
|
|
|
|
k-mer, each tagged with its origin (local slot); a transform stage
|
|
|
|
|
routes each variant to its target partition (unchanged per-k-mer
|
|
|
|
|
minimiser); a transform stage performs the lookup (existence-only —
|
|
|
|
|
`find_slot` hit/miss, cheaper than the SNP sweep's full column
|
|
|
|
|
fetch); a final stage/sink folds each answer into its origin's
|
|
|
|
|
running state (below) and, once a layer's k-mers are all resolved,
|
|
|
|
|
flushes the completed array to that layer's annex file. Many small,
|
|
|
|
|
single-purpose stages on purpose, to let the scheduler interleave
|
|
|
|
|
them finely across many in-flight items — this deliberately does
|
|
|
|
|
**not** mirror how `obipipeline` is used elsewhere today: `query.rs`'s
|
|
|
|
|
`process_chunk` lumps parse+route+query+serialise into one closure
|
|
|
|
|
(`query.rs:325,743-758`), and `scatter.rs` only pipelines file-
|
|
|
|
|
reading/superkmer construction, routing partitions afterwards in a
|
|
|
|
|
plain sequential loop (`KmerPartition::write_batch`,
|
|
|
|
|
`partition.rs:140`) — both under-use the fine-grained scheduling the
|
|
|
|
|
mechanism offers, so they are not precedents to copy, only existing
|
|
|
|
|
(and arguably improvable, out of scope here) usages. Cross-partition
|
|
|
|
|
lookups (querying another layer's MPHF for a variant) remain
|
|
|
|
|
necessary as before — only the *output* side is kept single-layer.
|
|
|
|
|
- **Reconciliation**: processed at the granularity of one *answer batch
|
|
|
|
|
per destination partition*, not one source k-mer at a time — this is a
|
|
|
|
|
proper shuffle, not a per-k-mer wait. Each source partition `p` holds a
|
|
|
|
|
small array of running states `(minorant = true, siblings = 0)`, one
|
|
|
|
|
per local slot, initialised at scatter time and **persisting across
|
|
|
|
|
however many destination-partition batches answer it** (up to 3, one
|
|
|
|
|
per variant, not necessarily all from the same `q`). Every scattered
|
|
|
|
|
query carries an origin tag (source partition + local slot) so its
|
|
|
|
|
answer can be routed back. When target partition `q` returns its batch
|
|
|
|
|
(all answers for every query that named `q`, regardless of which source
|
|
|
|
|
k-mer or which source partition they came from), that batch is walked
|
|
|
|
|
once, locally, and each answer updates — via its origin tag — the
|
|
|
|
|
matching entry in *its* source partition's array: a miss changes
|
|
|
|
|
nothing; a hit does `siblings += 1`, and if the found sibling's own
|
|
|
|
|
encoding is smaller than the source's, `minorant = false`. A given
|
|
|
|
|
source k-mer's state is final only once every destination batch
|
|
|
|
|
concerning it has been folded in; its partition's array is flushed to
|
|
|
|
|
the persistent annex once complete. Commutative per entry, so the order
|
|
|
|
|
in which destination batches arrive and get folded in doesn't matter.
|
|
|
|
|
|
|
|
|
|
**Open optimisation, not adopted yet — real tradeoff, not a free win.**
|
|
|
|
|
Since looking up sibling `y` from `x`'s visit already yields everything
|
|
|
|
|
needed to fill `y`'s own annex entry too, one visit per *family* could in
|
|
|
|
|
principle replace one visit per *observed family member* — cutting this
|
|
|
|
|
pass's cost roughly by the average family size instead of paying
|
|
|
|
|
`3 * N_distinct` regardless. But it means threads processing different
|
|
|
|
|
source k-mers can end up writing the *same* sibling's slot concurrently —
|
|
|
|
|
the fully independent, ownership-free parallelism of the plan above is
|
|
|
|
|
deliberately traded away for this gain. It stays safe only because the
|
|
|
|
|
computed value for a given slot is deterministic regardless of who
|
|
|
|
|
computes it, so redundant concurrent writes converge to the same
|
|
|
|
|
value — correct as long as each write is atomic, no locking needed — but
|
|
|
|
|
it is a real design complexity increase over "every member redoes its
|
|
|
|
|
own 3 lookups independently," not a strict improvement to adopt by
|
|
|
|
|
default.
|
|
|
|
|
4. **Trigger and caching** (`obikindex::KmerIndex`/`distance.rs`): compute
|
|
|
|
|
lazily on first `distance` call for an SNP-family metric against a given
|
|
|
|
|
index; check for an existing annex file first (mirrors
|
|
|
|
|
`PersistentBitMatrix::open()`'s auto-detect-and-fall-back,
|
|
|
|
|
`bitmatrix.rs:264-287`); if absent, run step 3 and persist; if present,
|
|
|
|
|
mmap and reuse.
|
|
|
|
|
5. **Invalidation.** `merge` and `filter` always write to a fresh `--output`
|
|
|
|
|
directory (`obikindex/src/merge.rs`, `obikmer/src/cmd/filter.rs`) so a
|
|
|
|
|
re-merged/re-filtered index simply has no annex yet — nothing to
|
|
|
|
|
invalidate. `select --in-place` (`select_layer.rs:139-235`) mutates
|
|
|
|
|
columns of an existing index without changing its location, which can
|
|
|
|
|
change sibling counts without removing rows — it must explicitly delete
|
|
|
|
|
any cached annex for that index as part of its in-place rewrite.
|
|
|
|
|
6. **Testing**: hand-built tiny indexes with known sibling counts (0-3);
|
|
|
|
|
order-independence (recompute twice on a static index, identical
|
|
|
|
|
result, given the fully-parallel no-ownership design); invalidation
|
|
|
|
|
(annex absent/correctly recomputed after `select --in-place`); once
|
|
|
|
|
Step 2's sweep is fixed, a regression check that sibling_count == 0
|
|
|
|
|
slots are never looked up cross-partition during the sweep.
|
|
|
|
|
|
|
|
|
|
Cost: `3 * N_distinct` existence-only lookups, computed once per index
|
|
|
|
|
state and amortised over every subsequent `distance` call that reuses the
|
|
|
|
|
cached annex — cheaper per-lookup than the sweep itself (hit/miss only, no
|
|
|
|
|
column fetch).
|
|
|
|
|
|
2026-07-10 09:47:19 +02:00
|
|
|
### Step 3 — finalisation (`obikindex`)
|
|
|
|
|
|
2026-08-10 12:35:27 +02:00
|
|
|
From the global `SnpTally` alone (diagonal and off-diagonal both populated by
|
|
|
|
|
the sweep, see Step 1/2 — no dependency on the external `shared_kmers`
|
|
|
|
|
matrix), derive n x n distance matrices, each a pure function of the
|
|
|
|
|
accumulated counts (same shape as `jaccard_to_mash`):
|
2026-07-10 09:47:19 +02:00
|
|
|
|
|
|
|
|
- `p_hat[i,j] = SNP / (SNP + shared)`
|
|
|
|
|
- Jukes-Cantor, Kimura-2P (from `P`, `Q`), optionally LogDet (needs the
|
|
|
|
|
diagonal + base-composition margins).
|
|
|
|
|
|
|
|
|
|
Guard the singularities (`p >= 3/4` for JC, `1-2P-Q <= 0` or `1-2Q <= 0` for
|
|
|
|
|
K2P) by clamping to a max distance, as `jaccard_to_mash` clamps `J <= 0`.
|
|
|
|
|
|
|
|
|
|
### Step 4 — surfacing (`obikindex` + `obikmer` CLI)
|
|
|
|
|
|
|
|
|
|
These metrics do not fit `DistanceMetric`'s current `LayeredStore`-partial
|
|
|
|
|
dispatch (they need the cross-partition sweep and produce a different
|
|
|
|
|
intermediate). Two options, to decide:
|
|
|
|
|
|
|
|
|
|
- **(a)** New `DistanceMetric` variants (`Pdistance`, `JukesCantor`,
|
|
|
|
|
`Kimura2P`, `LogDet`) whose `KmerIndex::distance` arm calls the sweep
|
|
|
|
|
(`snp.rs`) instead of the partial path, still returning `DistanceOutput`.
|
|
|
|
|
Keeps one CLI surface (`--metric jukes-cantor`), at the cost of a branch in
|
|
|
|
|
`distance()` that ignores the `LayeredStore` it built.
|
|
|
|
|
- **(b)** A dedicated pathway (`KmerIndex::snp_distance`) and a distinct CLI
|
|
|
|
|
entry, if mixing a cross-partition sweep into the partition-local `distance`
|
|
|
|
|
command is judged architecturally muddy.
|
|
|
|
|
|
|
|
|
|
Recommendation: (a) for user ergonomics (all pairwise distances under
|
|
|
|
|
`distance`, all feeding NJ/UPGMA/`--shared-kmers` unchanged), but compute the
|
|
|
|
|
sweep lazily only when an SNP-family metric is requested, so the existing
|
|
|
|
|
metrics keep their partition-local fast path untouched.
|
|
|
|
|
|
|
|
|
|
### Step 5 — subsampling flag
|
|
|
|
|
|
|
|
|
|
Add `--snp-sample <fraction>` (or a bottom-`s` hash threshold): restrict the
|
|
|
|
|
source-k-mer enumeration in Step 2 to `seq_hash(kmer) < threshold`. Divides
|
|
|
|
|
lookups proportionally; `p_hat` is unbiased. Off by default (exact).
|
|
|
|
|
|
|
|
|
|
### Testing
|
|
|
|
|
|
|
|
|
|
- **Primitive unit tests**: `central_canonical_neighbors` on hand-checked
|
|
|
|
|
k-mers incl. palindrome-boundary cases; lone-k-mer `minimizer` against
|
|
|
|
|
`RollingStat`'s incremental result on the same k-mer.
|
|
|
|
|
- **End-to-end tiny index**: two 1-genome indexes differing by a handful of
|
|
|
|
|
known isolated SNPs (transitions and transversions placed by hand), assert
|
|
|
|
|
exact `SNP`, `P`, `Q` counts and the resulting JC/K2P values.
|
|
|
|
|
- **Dedup invariant**: assert the tally is identical regardless of genome/
|
|
|
|
|
partition order and that no pair is double-counted (compare against a
|
|
|
|
|
brute-force all-pairs reference on a small index).
|
|
|
|
|
- **Subsampling**: `p_hat` within sampling error of the exact run.
|
|
|
|
|
|
|
|
|
|
### Suggested phasing
|
|
|
|
|
|
|
|
|
|
1. Step 0 primitives + their unit tests (self-contained, no distance wiring).
|
|
|
|
|
This also unblocks the long-declared-but-unimplemented `query --mismatch`
|
|
|
|
|
(`obikmer/src/cmd/query.rs:676`, currently a warning), which needs the same
|
|
|
|
|
neighbour + routing machinery.
|
|
|
|
|
2. `SnpTally` + finalisation math with a brute-force (non-swept) reference
|
|
|
|
|
backend, validated on a tiny index.
|
|
|
|
|
3. The real per-partition sweep (Step 2) behind the same finalisation; assert
|
|
|
|
|
it matches the brute-force backend.
|
|
|
|
|
4. CLI surfacing (Step 4a) and NJ/UPGMA integration (already generic over the
|
|
|
|
|
matrix).
|
|
|
|
|
5. Subsampling (Step 5).
|
|
|
|
|
|
|
|
|
|
## References
|
|
|
|
|
|
|
|
|
|
The Mash mutation-rate model this discussion contrasts with:
|
|
|
|
|
[@Mash-distances-doc; @Fan2015-mash-formula].
|