1
theory phylogeny snp_based sampling_design
Eric Coissac edited this page 2026-09-12 17:45:11 +02:00

Sampling design

Exhaustive computation over every non-monomorphic family in an index (see Central-position SNP model) is the default, but can be bounded to approximately N families instead. Sampling is designed around two properties: staying representative of the whole index, and letting the user favor families that actually carry signal.

Proportional sampling

Families are drawn in proportion to how many candidate families each part of the index actually holds, rather than uniformly across index partitions — this keeps the sample's composition representative of the whole index regardless of how unevenly candidate families happen to be distributed across partitions.

Per-family informativeness: Shannon entropy

A family's informativeness is measured directly rather than assumed uniform. For a family, entropy is computed over its observed states across genomes:

Quantity Definition
entropy15 Shannon entropy (bits) over the 16 possible states (the 15 non-empty subsets of {A,C,G,T}), genomes absent from the family excluded from the count
entropy4 the same, reduced to the 4 plain bases

A family where every genome carries the same single state has entropy 0 (uninformative); one where genomes are spread evenly across several states has higher entropy (more informative for distinguishing genomes).

Entropy-biased sampling

By default, sampling draws families uniformly. Weighting instead by informativeness biases the draw toward a target entropy \mu with a Gaussian curve of width \sigma: a family's probability of being kept scales with \exp\!\left(-\dfrac{(\text{entropy15} - \mu)^2}{2\sigma^2}\right) — a soft preference, not a hard cutoff, so no family is categorically excluded purely for having low or high entropy.

See phylo for how to bound and bias the sample, and for session/checkpointing behavior.