Compare commits

...
3 Commits
Author SHA1 Message Date
Eric Coissac e6f0ca472c ci: disable numa feature, bump obikmer, and document Sankoff costs
Release / create-release (push) Successful in 2m28s
ci.yml / build (pull_request) Failing after 3h0m41s
Release / build-linux-x86_64 (push) Successful in 8m31s
Release / build-macos-arm64 (push) Successful in 1m53s
Disable the `numa` default feature in CI build and test steps to prevent container environment deadlocks, and add comments explaining the cache key salt bump (`v2`) to mitigate incremental compilation corruption. Document a 16-state Sankoff cost matrix derived from set-edit distances, including substitution, gain/loss, and context-disappearance costs compatible with TNT's interface. Bump `obikmer` crate version to 1.1.43.
2026-08-11 18:26:54 +02:00
coissac 442f7a9e4c Merge pull request 'chore: update ci cache, document distance metrics, and bump version' (#64) from push-wpxsvyylwmsq into main
Reviewed-on: #64
2026-08-11 15:17:42 +00:00
Eric Coissac a63692b8c4 chore: update ci cache, document distance metrics, and bump version
Release / create-release (push) Successful in 2m26s
Release / build-linux-x86_64 (push) Successful in 8m43s
Release / build-macos-arm64 (push) Successful in 2m7s
ci.yml / build (pull_request) Canceled after 59m15s
Updated CI workflow cache keys with a `v2` salt and `Cargo.lock` hash to prevent stale incremental compilation caches and deadlocks, while updating restore keys and documenting interrupted job state. Introduced a 3-way ordinal distance metric framework that replaces ambiguous IUPAC encoding with explicit k-mer scoring, bridging pairwise methods to character-based phylogenetics via Sankoff parsimony. Bumped the `obikmer` crate version to 1.1.42.
2026-08-11 17:12:28 +02:00
4 changed files with 284 additions and 6 deletions
+12 -4
View File
@@ -25,11 +25,19 @@ jobs:
~/.cargo/registry
~/.cargo/git
src/target
key: ${{ runner.os }}-cargo-${{ hashFiles('src/Cargo.lock') }}
restore-keys: ${{ runner.os }}-cargo-
key: ${{ runner.os }}-cargo-v2-${{ hashFiles('src/Cargo.lock') }}
restore-keys: ${{ runner.os }}-cargo-v2-
# Both `obikmer` and `obikindex` default to the `numa` feature
# (hwloc-based topology detection + CPU pinning), which is only useful
# on bare-metal multi-socket indexing hosts. Under this runner's
# container/cgroup setup it deadlocks at startup — confirmed live
# (2026-08-11): the same test binary hangs indefinitely with `numa` on
# and passes instantly, repeatedly, with it off, on the same
# container. Disable it for CI; it has nothing to do with test
# correctness.
- name: Build
run: cargo build --release
run: cargo build --release --no-default-features
- name: Test
run: cargo test --release
run: cargo test --release --no-default-features
+270
View File
@@ -203,6 +203,276 @@ implication is that columns should be allowed partial coverage (>=2 resolved
genomes, not unanimous) rather than requiring every genome to be net
single-copy at that locus.
## Context, detectability, and a 3-way ordinal distance per pair
Empirical follow-up to the pseudo-alignment idea above: `obikmer distance
--snp` was run on a real 20-genome benchmark index and the resulting FASTA
fed to `raxml-ng`. Two problems surfaced, both traced back to conflating
distinct notions under one symbol.
**"Context", precisely.** Sharing a central base between two genomes is not
just sharing a nucleotide — it is sharing a **context**: the `2m` flanking
bases, identical, which is a homology claim about that flanked window
(guaranteed non-coincidental by k-specificity, Bias 4 above), *not* a claim
about orthology or paralogy of the copy each genome carries. `A` opposite `C`
= same context, divergent centre. `A` opposite nothing = **this context is
not observed in one of the two genomes** — informative, not neutral.
**Why the IUPAC/DNA encoding used for the first `--snp` test was wrong.**
Feeding IUPAC-coded ambiguity into a standard DNA model (`raxml-ng --model
GTR+G`) is a semantic mismatch: Felsenstein-pruning ML treats an ambiguous
tip as "exactly one true state, unknown which" (a uniform partial-likelihood
vector over compatible bases), not "these states are simultaneously
present". The two encodings look identical (same IUPAC letters) but the
software reads them backwards from what was intended — this invalidates the
literal branch lengths from that first experiment (topology-level groupings
by genus were still informative, see the worked example further down).
**Why `-` (absence) must not be scored as similarity, but also must not be
scored as a shared character between two absences.** Two genomes both
lacking a context are not observed to resemble each other at that locus —
neither is observed to differ from the other either. It is a symmetric
non-observation, uninformative for that pair, and should contribute nothing
(not a small positive nor a small negative signal) to their distance. A
genome carrying a state (`A`) against one carrying none is a different case
entirely: informative, and should not be scored as neutral "missing data"
the way a generic DNA/ML pipeline would.
**Detectability vs existence — a deliberate simplification, accepted.**
"Context not observed" conflates two different biological events: (1) true
loss of the locus, (2) the locus still exists but a mutation/indel *outside*
the centre, anywhere in the `2m` flanks, broke k-mer recognition. The design
adopts a rigorist stance on purpose: any flank-breaking mutation counts as
"this context no longer exists", full stop — because both causes (1) and (2)
independently require *at least* one more mutational event than a lone
central substitution would. This licenses treating "context absent in one of
the two genomes" as a **lower bound** on distance strictly greater than a
plain central SNP, without needing to know which of the two causes applies.
Coarser than a true event count, and accepted as such (fine substitution-type
modelling, e.g. transition/transversion weighting, is a secondary
refinement, not required for this to be useful).
**Resulting ordinal distance between two genomes at one family/context:**
| Comparison | Distance | Meaning |
|---|---|---|
| same centre (`A`/`A`) | `0` | identical |
| different centre, both single-copy (`A`/`C`) | `1` | plain central SNP |
| one genome has a state, the other has none | `>1` (lower bound) | context undetectable in one genome — at least one extra mutational event, of unknown type |
| neither genome has any state (`∅`/`∅`) | excluded | symmetric non-observation, not comparable, contributes nothing |
This is a direct extension of `KmerIndex::raw_snp_distance` (`obikindex/src/siblings.rs`),
which today only implements the `0`/`1` rows and silently drops everything
else (including the informative `>1` row) rather than scoring it.
**Open, not yet resolved:**
- Calibrating `>1` to a real number for tools expecting continuous distances
(NJ/UPGMA, ML branch lengths), rather than an arbitrary placeholder.
Natural route: estimate `p_hat` from the resolved (`0`/`1`) sites first,
then use the already-derived ascertainment formula (`P(usable window
showing a central SNP) = p * (1-p)^(2m)`, Bias 1 above) to derive a
model-consistent value for the `>1` bucket instead of guessing a constant.
- Where multi-copy/ambiguous states (the IUPAC case: a genome carrying more
than one form) fit into this ordinal scheme — plausibly also `>1` by the
same "at least one extra event" argument (a second form appearing is a
gain, itself an event), but not yet worked out.
**Practical alternative validated for the pseudo-alignment output itself**
(orthogonal to the ordinal-distance question above, useful regardless of how
`>1` ends up calibrated): re-encode each family as 4 independent binary
presence/absence characters (`A`,`C`,`G`,`T` columns) instead of one IUPAC
column, feed to a `BIN`-type model instead of `DNA`. `∅` becomes an explicit
`0000` state (identity with another `0000`, not missing data) rather than a
gap — removes the semantic mismatch above by construction. Known cost,
accepted for now: a plain central substitution (`A` -> `C`) becomes 2 binary
flips (`1000` -> `0100`), overweighting substitutions relative to true
gain/loss events, and the 4 sub-characters of one family are not
statistically independent the way a generic `BIN` model assumes. A proper
fix (single 16-state alphabet, i.e. the powerset of `{A,C,G,T}`, with a
substitution-rate structure that respects the subset lattice rather than a
fully general 16x16 GTR-analogue) is very likely not expressible in
`raxml-ng`'s `MULTI` datatype as-is (Mk or fully-general rates only) and a
fully general 16-state rate matrix is almost certainly unidentifiable here
(states of cardinality >=3 are ~2% of sites in the benchmark run). Treated as
a longer-term research question, not a near-term implementation target.
## Sankoff parsimony as the resolution of the 16-state model problem
The "longer-term research question" just above (a 16-state alphabet — the
powerset of `{A,C,G,T}` — with a substitution structure that respects the
subset lattice) turns out to have a near-term answer, once the *unification*
question below is worked through.
**Distance methods (NJ/UPGMA/ME) vs. character methods (parsimony/ML): not a
deep philosophical divide, but a real practical distinction for this
project.** Historically "phenetic" (characters -> distances -> tree) and
"cladistic" (characters -> tree directly) approaches were presented as
opposed schools; the modern view is a mathematical continuity, not a
dichotomy — Minimum Evolution (ME: find the tree minimizing total branch
length from a distance matrix) and Maximum Parsimony (MP: find the tree
minimizing total character-state changes) are both instances of "minimise a
global explanatory cost", and coincide under simple encodings (see Farris
1983, "The Logical Basis of Phylogenetic Analysis"; the MP/ME connection is
developed in the Minimum Evolution / Balanced Minimum Evolution literature,
e.g. Nei and colleagues — citations not independently re-verified here, flag
before quoting further). NJ's own agglomeration step already uses the whole
distance matrix jointly (the Q-matrix), not just the pair being merged — an
earlier claim in this discussion that distance methods are "blind" to
cross-taxon structure at every stage was too strong.
What *does* remain a real, structural distinction for this project: in a
character method, a given character's cost is **re-evaluated per candidate
topology** during tree search (the same family can cost 1 change under one
topology, 2 under another). In a pairwise-distance pipeline
(`raw_snp_distance` as it exists today), each family's contribution to
`d(i,j)` is computed **once**, independent of any candidate topology, before
NJ/UPGMA ever runs — so a question like "does this shared `∅` look like a
synapomorphy under topology T" can never be posed in that pipeline, for any
T. That question is only answerable by a method that tests candidate
topologies and re-scores characters under each — i.e. a character method.
**Sankoff parsimony directly resolves the `∅`/gain-loss/substitution
question, without the identifiability problem of a fitted 16-state model.**
Sankoff's algorithm generalises Fitch parsimony to an arbitrary
user-supplied cost matrix between states (`obikseq`/`obikindex` would treat
each family as a `2^{4}`-state character, state = subset of `{A,C,G,T}`
observed in that genome, `∅` included as a real state, not a gap). The
previous 16-state idea failed specifically because *fitting* a full 16x16
rate matrix by ML is unidentifiable at this data volume; Sankoff sidesteps
that because the cost matrix is **fixed a priori from domain knowledge**, not
estimated — e.g. `c({A},{C}) = 1` (a substitution), `c({A},{A,C}) = 1` (a
gain), `c({A,C},{A}) = 1` (a loss), `c({A,C},{G,T}) = 2` (two changes) — no
estimation, no overparameterisation. This reframes gain/loss and central
substitution as two cost categories with independently chosen weights,
exactly the "two families of parameters" (`mu_substitution`, `mu_gain/loss`)
floated earlier in this discussion, now with an actual algorithmic home.
**Caveat, not blocking for this project's scope.** Sankoff is still
parsimony: in principle exposed to Felsenstein's statistical-inconsistency
result under long-branch attraction (already invoked earlier against
`D_F = min(a,b)`) — parsimony and ML only provably coincide in the
short-branch regime. This is not a practical concern here because it is
exactly this estimator's declared target (closely related genomes, short
branches) — the regime where parsimony's known failure mode does not apply —
but worth stating explicitly as a scope guard rather than leaving it
implicit.
**Cheapest next experiment: don't write a Sankoff tree-search engine, use
one that exists.** The hard part of a from-scratch implementation is not the
Sankoff DP itself (a straightforward dynamic program over a *fixed* tree) but
the topology search (SPR/NNI with incremental re-scoring) that comes for
free with `raxml-ng` on the ML side. **TNT** (Tree analysis using New
Technology, free, standard in morphological cladistics) already implements
Sankoff parsimony with a custom cost matrix plus topology search — the
family-state matrix (already close to what `--snp` produces, minus the
IUPAC/DNA-model mismatch) could be fed there directly, no new code required,
before considering a bespoke engine.
**The concrete comparison this unlocks:** run both pipelines on the same
family data —
`k-mer families -> D_ij -> NJ/ME/UPGMA` (phenetic, what exists today) vs.
`k-mer families -> characters -> argmin_T Sankoff-cost(T)` (cladistic, via
TNT) — and compare the resulting topologies. Agreement would validate that
the pairwise-distance projection preserves the phylogenetic signal;
disagreement would pinpoint exactly what the projection to a single number
per pair loses. Not yet run.
### A concrete Sankoff cost matrix for the 16-state alphabet
**Set-edit-distance formula.** For two states `X, Y ⊆ {A,C,G,T}`, split into
`seulement_X = X \ Y` (size `a`) and `seulement_Y = Y \ X` (size `b`).
Elements present in both cost nothing. Pair up to `min(a,b)` of the
remaining elements as **substitutions** (cheaper than treating them as an
unrelated loss + gain whenever `c_sub < 2*c_gl`, which any sane parameter
choice satisfies); whatever is left over after pairing is a pure
**gain/loss**:
```
cost(X,Y) = min(a,b)*c_sub + |a-b|*c_gl
```
Worked examples (`c_sub = c_gl = 1`): `{A}->{C}` = 1 (one substitution);
`{A}->{A,C}` = 1 (one gain, no substitution pair available since nothing is
only-in-Y that matches an only-in-X element after the shared `A` is
excluded); `{A,C}->{G,T}` = 2 (two substitution pairs, `A/C` vs `G/T`, both
same-size sets share nothing); `{A,C,G}->{A,C,T}` = 1 (`G`/`T` is the only
mismatched pair, `A,C` shared).
**`∅` is a flat-cost special case, not a instance of the general formula.**
Applying the formula naively to e.g. `{A,C,G} -> ∅` would charge `3*c_gl`
(three independent losses). Reject that: total context disappearance
(flank-breaking, or true structural loss) is plausibly **one** event, not
`|X|` of them, so:
```
cost(X, ∅) = cost(∅, X) = c_ctx (constant, independent of |X|)
cost(∅, ∅) = 0
```
`c_ctx` should not be guessed — it is exactly the value already derived for
the `>1` bucket in "Context, detectability, and a 3-way ordinal distance per
pair" above (`(2m*p_hat)/(1-(1-p_hat)^(2m)) + p_hat`), reused rather than
invented.
**Better construction method: shortest path in a small state graph, not the
closed-form formula directly.** Build a graph on the 16 states with two edge
types — substitution edges between same-cardinality sets differing by one
element (weight `c_sub`, or `c_ts`/`c_tv` if split further below), and
gain/loss edges between sets whose cardinality differs by one (weight
`c_gl`) — then define `cost(X,Y)` as shortest-path distance in that graph,
precomputed once (16 nodes, trivial) into a dense 16x16 matrix before
feeding it to Sankoff/TNT. Verified equivalent to the closed-form formula
above in the uniform-cost case (checked by hand on `{A}->{C,G,T}`: both give
`c_sub + 2*c_gl`). The graph construction is not just a reformulation for
its own sake: it is the version that generalises correctly once substitution
costs stop being uniform (next point) — the closed-form's `min(a,b)`
counting silently assumes *any* pairing costs the same, which breaks the
moment `c_sub` depends on which two bases are involved.
**Transition/transversion refinement.** Split `c_sub` into `c_ts` (A<->G or
C<->T) and `c_tv` (the other four pairs) — already well-defined in canonical
space (see "Canonical invariance" above: a transition maps to a transition,
a transversion to a transversion, regardless of orientation). This turns
"pick `min(a,b)` substitution pairs" into a genuine (tiny, <=4 elements per
side, trivially enumerable) minimum-cost bipartite matching problem instead
of a plain count — the state-graph shortest-path construction handles this
automatically, no separate logic needed. Calibration is not a guess either:
the "Sufficient statistic: 4x4 base-pair tally" section below already plans
to collect the joint `(centre_i, centre_j)` distribution over resolved (`0`
or `1`) sites — that tally directly gives the empirical Ts/Tv ratio, from
which `c_ts`/`c_tv` follow (e.g. `cost ∝ -log(observed rate)`, the standard
generalised-parsimony step-weighting heuristic), reusing a statistic already
planned rather than adding a new one.
**`gamma`/`mu` (i.e. `c_gl`/`c_sub`) sensitivity sweep before calibration.**
No strong prior on whether gain/loss events should cost more or less than a
point substitution — duplication/deletion rates are not a priori equal to
point-mutation rates, but the direction isn't obvious, and setting it too
high risks the parsimony search *eliminating* exactly the
heterozygosity/paralogy signal the design is meant to tolerate (see
"Heterozygosity, ploidy..." below). Cheap first step: run the topology at a
handful of ratios (`0.5, 1, 2, 5`) and check whether it's stable — a robust
topology across that range is far more trustworthy than one built on a
single, unvalidated guess. Empirical calibration of `c_gl` from the
family-size distribution already available (`sibling_annex_stats`) is a
natural follow-up once the sensitivity sweep shows the topology is worth
refining further.
**Caveat carried over from the `D_F = min(a,b)` rejection earlier:** this
cost matrix is only valid as **Sankoff step-cost input**, re-evaluated for
every branch of every candidate topology during tree search. Reusing
`cost(leaf_A, leaf_B)` directly as a standalone pairwise distance (bypassing
the tree) would reintroduce the exact circularity already rejected — the
`min(a,b)` pairing here is a locally-defined edit distance between two
states, not a claim about the true evolutionary history between two
specific genomes.
**Feasibility confirmed:** TNT's `costs` command accepts custom step
matrices for multistate characters, so this whole construction (16x16
matrix derived from the state graph, `c_ts`/`c_tv`/`c_gl`/`c_ctx` as
tunable parameters) is directly usable there — no new tooling required
before testing it.
## Heterozygosity, ploidy, and consensus-assembly inputs
A within-genome multiplicity signal (more than one of the 4 central forms
+1 -1
View File
@@ -1715,7 +1715,7 @@ dependencies = [
[[package]]
name = "obikmer"
version = "1.1.41"
version = "1.1.43"
dependencies = [
"clap",
"csv",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "obikmer"
version = "1.1.41"
version = "1.1.43"
edition = "2024"
[[bin]]