Consolidate compare_sparse as example and clean up project artifacts

Restructure the project by moving the standalone compare_sparse utility into an example directory, removing Sankoff parameter configurations and benchmark scripts, updating version control ignores, and expanding the test suite with diagnostic checks and performance benchmarks.
This commit is contained in:
Eric Coissac
2026-08-20 12:48:42 +02:00
parent dc3d82f8db
commit dbd8af376c
13 changed files with 12 additions and 1184 deletions
Vendored
BIN
View File
Binary file not shown.
+1
View File
@@ -1,4 +1,5 @@
.venv/
.DS_Store
src/target
data-stress
*.fasta
-43
View File
@@ -1,43 +0,0 @@
Voici la version corrigée :
---
**Bug** : dans `base_pair_tally`, toutes les transitions/comptes depuis/vers A valent 0 dans `_sankoff_params.yaml`, alors que C/G/T sont corrects.
**Contexte** : obikmer, pipeline phylogénétique `--sankoff`. L’index est construit sur 20 génomes bactériens. Même symptôme sur un jeu de 100 génomes de plantes : A est toujours à 0.
**Fichier clé** : `src/obikphylo/src/siblings/sankoff_bundle.rs` (Pass A + Pass B).
**Ce qui a été vérifié** :
- Le fichier de sortie `_sankoff_params.yaml` montre bien `composition_transitions` avec A à 0 partout.
- L’index contient bien des familles avec A (`mask.has(0) == true`), et même des familles où A co-existe avec d’autres bases (`mask == 0b0011` par ex.).
- Un k-mer propriétaire de famille avec `mask == 0b0001` (A seul) a été identifié : forward `GAACAAGAGATCTCGATCTTGTCTACAAGGA`, revcomp `TCCTTGTAGACAAGATCGAGATCTCTTGTTC`.
- Le diagnostic CLI sur l’index réel donne :
- Pass A : `a_pairs=623342 a_snp=623342 a_shared=0 a_both_a=0`
- Pass B : `families_with_a=22965521 a_single_form_genomes=22913238 a_included_pairs=0 a_same_incremented=0 bp_same=[0, 96389, 222720, 277909] bp_counts[0]=[0, 0, 0, 0]`
**Interprétation** : A est fréquemment en `single_form` (mask == 1) chez certains génomes, mais **jamais simultanément** chez deux génomes différents dans la même famille. Donc toutes les paires “avec A” sont 100% SNP → ratio = 1.0 > `ratio_ceiling=0.5` → toutes exclues par le filtre `included`. C’est pourquoi `bp_same[0]` et `bp_counts[0][*]` restent à 0.
**Point crucial** : le bug n’apparaît **que sur l’index compacté sparse**. Sur le même index avant compaction (matrice dense `matrix.pbmx`), `--sankoff` produit des tallies corrects pour A. Dès qu’on compacte avec `pack --sparse`, A disparaît.
**Vérifications supplémentaires (diagnostic sparse)** :
- La compaction `pack --sparse` produit une matrice `PersistentSparseBitMatrix` dont le contenu est **strictement identique** à la matrice dense d'origine : vérification exhaustive coordonnée par coordonnée sur **1 804 774 880 cellules** (512 partitions × 2 layers), **zéro différence**.
- `fill_row` et `fill_sub_matrix` (les deux chemins de lecture utilisés par le pipeline phylogénétique) restituent les mêmes bits sur dense et sparse.
- **Conclusion** : le bug n'est **pas** dans la compaction sparse elle-même, ni dans les chemins de lecture individuels. La structure stocke correctement A, C, G, T.
**Conséquence logique** :
Si les matrices sont identiques mais que le résultat final diffère, le bug se situe dans l'**intersection** des informations — c'est-à-dire dans le code qui **combine** les lectures des deux matrices (ou qui transforme les résultats bruts en tallies). Deux endroits possibles :
1. **Le scan `sankoff_bundle`** (`family_scan.rs` + `sankoff_bundle.rs`) : la boucle qui lit les matrices, construit `genome_mask`, et accumule `bp_counts` / `same`. C'est l'étape d'intersection proprement dite.
2. **La conversion des tallies en YAML** (`obikmer/src/cmd/phylo/sankoff.rs`) : moins probable, mais possible si quelque chose sélectionne/filtre les transitions avant écriture.
**Hypothèse la plus probable** : bug dans la résolution cross-partition lors de la construction de l'annex sibling (`build_sibling_annex`). A (bit 0) serait systématiquement manquant ou mal résolu quand on interroge les variants d'une famille depuis une partition différente. À vérifier dans `src/obikphylo/src/siblings/build.rs` et `src/obikphylo/src/siblings/cache.rs` (`PartitionCache::find` / `find_presence_batch`).
**Prochaine étape logique** :
1. Inspecter `build_sibling_annex` pour voir si les variants avec base A sont bien générés et bien recherchés dans `cache.find`.
2. Vérifier `PartitionCache::find` et `resolve_layer_hits` pour un éventuel biais contre le bit 0.
3. Si besoin, ajouter un diagnostic ciblé (compteurs par base) **uniquement** dans `cache.rs` ou `build.rs`, pas dans `sankoff_bundle.rs` qui est déjà propre.
**Contraintes** :
- Ne pas modifier `sankoff_bundle.rs` davantage.
- Ne pas toucher à git.
- Faire des diagnostics minimaux et ciblés.
@@ -1,167 +0,0 @@
ratio_ceiling: 0.5
cardinality_transitions:
- from: 0
to: 0
count: 115955295
probability: 0.8297007290571883
- from: 0
to: 1
count: 12901663
probability: 0.0923159153460836
- from: 0
to: 2
count: 10519367
probability: 0.07526975347801175
- from: 0
to: 3
count: 339782
probability: 0.0024312591600108434
- from: 0
to: 4
count: 39459
probability: 0.00028234295870548727
- from: 1
to: 0
count: 12901663
probability: 0.9007741539906029
- from: 1
to: 1
count: 973491
probability: 0.0679676357956696
- from: 1
to: 2
count: 444884
probability: 0.031061112720426456
- from: 1
to: 3
count: 2739
probability: 0.00019123274323474897
- from: 1
to: 4
count: 84
probability: 5.864750066344985e-6
- from: 2
to: 0
count: 10519367
probability: 0.9590597257562327
- from: 2
to: 1
count: 444884
probability: 0.04056045644508228
- from: 2
to: 2
count: 3796
probability: 0.00034608458084699006
- from: 2
to: 3
count: 327
probability: 0.000029812870900149037
- from: 2
to: 4
count: 43
probability: 3.920346937940087e-6
- from: 3
to: 0
count: 339782
probability: 0.9908029486551426
- from: 3
to: 1
count: 2739
probability: 0.007986913010007698
- from: 3
to: 2
count: 327
probability: 0.0009535306879417734
- from: 3
to: 3
count: 58
probability: 0.00016912776728019222
- from: 3
to: 4
count: 30
probability: 0.00008747987962768563
- from: 4
to: 0
count: 39459
probability: 0.9957604663487016
- from: 4
to: 1
count: 84
probability: 0.0021197668256491787
- from: 4
to: 2
count: 43
probability: 0.0010851187321775557
- from: 4
to: 3
count: 30
probability: 0.0007570595805889924
- from: 4
to: 4
count: 11
probability: 0.0002775885128826305
composition_transitions:
- from: 'A'
to: 'A'
count: 169043
probability: 0.49442957633191476
- from: 'A'
to: 'C'
count: 27672
probability: 0.08093712982055894
- from: 'A'
to: 'G'
count: 124160
probability: 0.36315242983957063
- from: 'A'
to: 'T'
count: 21020
probability: 0.06148086400795566
- from: 'C'
to: 'A'
count: 27672
probability: 0.08651690662664728
- from: 'C'
to: 'C'
count: 145183
probability: 0.4539167409213838
- from: 'C'
to: 'G'
count: 23425
probability: 0.07323859994684925
- from: 'C'
to: 'T'
count: 123565
probability: 0.3863277525051197
- from: 'G'
to: 'A'
count: 124160
probability: 0.3846082361177367
- from: 'G'
to: 'C'
count: 23425
probability: 0.07256320820761906
- from: 'G'
to: 'G'
count: 148125
probability: 0.4588441927749658
- from: 'G'
to: 'T'
count: 27112
probability: 0.08398436289967846
- from: 'T'
to: 'A'
count: 21020
probability: 0.06258131551760583
- from: 'T'
to: 'C'
count: 123565
probability: 0.3678810776371534
- from: 'T'
to: 'G'
count: 27112
probability: 0.08071858355439246
- from: 'T'
to: 'T'
count: 164186
probability: 0.4888190232908483
-18
View File
@@ -1,18 +0,0 @@
#!/usr/bin/awk -f
/^>/ {
print
next
}
{
gsub(/[Aa]/, "a")
gsub(/[Gg]/, "A")
gsub(/a/, "G")
gsub(/[Cc]/, "c")
gsub(/[Tt]/, "C")
gsub(/c/, "T")
print
}
+1 -10
View File
@@ -493,16 +493,6 @@ dependencies = [
"impl-tools",
]
[[package]]
name = "compare_sparse"
version = "0.1.0"
dependencies = [
"anyhow",
"fastrand",
"obicompactvec",
"obikindex",
]
[[package]]
name = "console"
version = "0.15.11"
@@ -1771,6 +1761,7 @@ dependencies = [
name = "obikindex"
version = "0.1.0"
dependencies = [
"anyhow",
"crossbeam-channel",
"hwlocality",
"indicatif 0.17.11",
+1 -1
View File
@@ -1,5 +1,5 @@
[workspace]
resolver = "3"
members = ["obikseq", "obiread", "obiskbuilder", "obifastwrite", "obikmer","obikrope","obipipeline", "obikpartitionner","obiskio","obidebruinj","obilayeredmap", "obicompactvec", "obisys", "obikindex", "obitaxonomy", "obikentropy", "obikphylo", "compare_sparse"]
members = ["obikseq", "obiread", "obiskbuilder", "obifastwrite", "obikmer","obikrope","obipipeline", "obikpartitionner","obiskio","obidebruinj","obilayeredmap", "obicompactvec", "obisys", "obikindex", "obitaxonomy", "obikentropy", "obikphylo"]
[profile.release]
debug = 1
-130
View File
@@ -1,130 +0,0 @@
use std::path::PathBuf;
use obicompactvec::{BinaryMatrix, PersistentBitMatrix, PersistentSparseBitMatrix};
use obikindex::KmerIndex;
fn main() -> anyhow::Result<()> {
let orig_root = std::env::args().nth(1).expect("usage: compare_sparse <orig_dense> <sparse>");
let sparse_root = std::env::args().nth(2).expect("usage: compare_sparse <orig_dense> <sparse>");
let orig = KmerIndex::open(&orig_root)?;
let sparse = KmerIndex::open(&sparse_root)?;
let n_parts = orig.n_partitions();
let n_layers = orig.n_layers_per_partition()?;
let n_genomes = orig.meta().genomes.len();
println!("index orig: {} partitions, {} layers, {} genomes", n_parts, n_layers, n_genomes);
println!("index sparse:{} partitions, {} layers, {} genomes", sparse.n_partitions(), sparse.n_layers_per_partition()?, sparse.meta().genomes.len());
let mut total_layers = 0usize;
let mut mismatches = 0usize;
let mut slot_checked = 0usize;
for part in 0..n_parts {
let part_dir_orig = orig.partition().part_dir(part);
let part_dir_sparse = sparse.partition().part_dir(part);
if !part_dir_orig.join("index").exists() || !part_dir_sparse.join("index").exists() {
continue;
}
for layer in 0..n_layers {
let layer_dir_orig = part_dir_orig.join("index").join(format!("layer_{layer}"));
let layer_dir_sparse = part_dir_sparse.join("index").join(format!("layer_{layer}"));
if !layer_dir_orig.exists() || !layer_dir_sparse.exists() {
continue;
}
let presence_orig = layer_dir_orig.join("presence");
let presence_sparse = layer_dir_sparse.join("presence");
if !presence_orig.exists() || !presence_sparse.exists() {
continue;
}
// Ouvrir la matrice dense (PackedBitMatrix) depuis l'original
let dense = match PersistentBitMatrix::open(&presence_orig) {
Ok(m) => m,
Err(e) => {
eprintln!("ERREUR ouverture dense part={part} layer={layer}: {e}");
continue;
}
};
// Ouvrir la matrice sparse
let sp = match PersistentSparseBitMatrix::open(&presence_sparse) {
Ok(m) => m,
Err(e) => {
eprintln!("ERREUR ouverture sparse part={part} layer={layer}: {e}");
continue;
}
};
if dense.n_cols() != sp.n_cols() {
eprintln!("ERREUR n_cols différent part={part} layer={layer}: dense={} sparse={}", dense.n_cols(), sp.n_cols());
continue;
}
let n = dense.n();
if n != sp.n() {
eprintln!("ERREUR n différent part={part} layer={layer}: dense={} sparse={}", n, sp.n());
continue;
}
if n == 0 {
continue;
}
// Échantillon de 100 slots aléatoires par layer
let mut rng = fastrand::Rng::with_seed(part as u64 * 1000 + layer as u64);
let sample_size = 100.min(n);
let mut checked_any = false;
for _ in 0..sample_size {
let slot = rng.usize(0..n);
let mut dense_row = vec![0u32; dense.n_cols()];
let mut sparse_row = vec![0u32; sp.n_cols()];
dense.fill_row(slot, &mut dense_row);
sp.fill_row(slot, &mut sparse_row);
if dense_row != sparse_row {
// Premier mismatch pour ce layer : diagnostiquer la colonne A
if !checked_any {
let a_col = 0; // A est la colonne 0
let mut a_dense = 0u32;
let mut a_sparse = 0u32;
// scanner tous les slots pour compter les différences sur la colonne A
let mut diff_a = 0usize;
for s in 0..n {
let mut d = vec![0u32; dense.n_cols()];
let mut s2 = vec![0u32; sp.n_cols()];
dense.fill_row(s, &mut d);
sp.fill_row(s, &mut s2);
if d[a_col] != s2[a_col] {
diff_a += 1;
}
}
eprintln!("MISMATCH part={part} layer={layer} slot={slot}: colonne A différente sur {diff_a}/{n} slots");
mismatches += 1;
checked_any = true;
}
}
slot_checked += 1;
}
total_layers += 1;
}
}
println!("Vérifié {} layers, {} slots échantillonnés", total_layers, slot_checked);
if mismatches == 0 {
println!("OK : aucune différence détectée.");
} else {
println!("MISMATCHES détectés dans {} layers", mismatches);
}
Ok(())
}
-14
View File
@@ -1,14 +0,0 @@
[package]
name = "compare_sparse"
version = "0.1.0"
edition = "2021"
[[bin]]
name = "compare_sparse"
path = "src/main.rs"
[dependencies]
obikindex = { path = "../obikindex", default-features = false }
obicompactvec = { path = "../obicompactvec" }
anyhow = "1"
fastrand = "2"
+1
View File
@@ -24,6 +24,7 @@ hwlocality = { version = "1.0.0-alpha.11", features = ["vendored"], option
obiread = { path = "../obiread" }
tempfile = "3"
tracing-subscriber = { version = "0.3", features = ["fmt", "env-filter"] }
anyhow = "1"
[features]
default = ["numa"]
@@ -1,9 +1,14 @@
//! Diagnostic tool: verify that a sparse-packed presence index (`pack --sparse`)
//! is bit-for-bit identical to the dense index it was compacted from.
//!
//! Usage: cargo run --example compare_sparse -p obikindex -- <sparse_root> <dense_root>
use obicompactvec::{PersistentBitMatrix, PersistentSparseBitMatrix};
use obikindex::KmerIndex;
fn main() -> anyhow::Result<()> {
let sparse_root = std::env::args().nth(1).expect("usage: compare_full <sparse> <orig_dense>");
let dense_root = std::env::args().nth(2).expect("usage: compare_full <sparse> <orig_dense>");
let sparse_root = std::env::args().nth(1).expect("usage: compare_sparse <sparse_root> <dense_root>");
let dense_root = std::env::args().nth(2).expect("usage: compare_sparse <sparse_root> <dense_root>");
let sparse = KmerIndex::open(&sparse_root)?;
let dense = KmerIndex::open(&dense_root)?;
+1 -632
View File
@@ -15,7 +15,7 @@ use super::cardinality::CardinalityExt;
use super::distance::DistanceExt;
use super::entropy::ShannonEntropyExt;
use super::entropy_annex::{EntropyAnnex, ENTROPY_ANNEX_FILE_NAME};
use super::helpers::{central_base, is_minorant};
use super::helpers::is_minorant;
use super::sankoff_bundle::SankoffBundleExt;
use super::stats::SiblingStatsExt;
use super::subsample::EntropyBias;
@@ -550,634 +550,3 @@ fn entropy_annex_builds_on_demand_and_biases_selection() {
assert!(find_layer(&selections_far).is_none(), "mu far from the true entropy must never select it, in any layer");
}
#[test]
#[ignore]
fn diag_real_index_layer_distribution() {
let idx = KmerIndex::open("/Users/coissac/Sync/travail/__MOI__/obikmer/benchmark/global_index_presence")
.expect("open real index");
let layer_dirs = super::family_scan::sibling_layer_dirs(&idx).expect("layer dirs");
let counts = super::subsample::non_monomorphic_counts(&layer_dirs).expect("counts");
let total: u64 = counts.iter().sum();
let n_zero = counts.iter().filter(|&&c| c == 0).count();
let max = counts.iter().max().unwrap();
let min_nonzero = counts.iter().filter(|&&c| c > 0).min().unwrap_or(&0);
println!("n_layers={} total={} n_zero_layers={} max={} min_nonzero={}",
counts.len(), total, n_zero, max, min_nonzero);
let n = 100_000usize;
let selections = super::subsample::compute_selections(&idx, &layer_dirs, Some(n), None).expect("selections");
let selected_total: usize = selections.iter().map(|s| match s {
None => 0, // shouldn't happen when total_count > n
Some(set) => set.len(),
}).sum();
println!("requested={n} selected_total={selected_total}");
let mut sorted_counts = counts.clone();
sorted_counts.sort_unstable_by(|a, b| b.cmp(a));
println!("top10 counts: {:?}", &sorted_counts[..10.min(sorted_counts.len())]);
}
#[test]
#[ignore]
fn diag_real_index_base_a_representation() {
// Diagnostic for a user-reported observation, reproduced on two
// unrelated real datasets (a plant and this bacterial index):
// `composition_transitions`'s row/column for base 'A' is entirely
// zero, even though 'A' appears at real (~6-22%) frequency in the
// pseudo-alignment itself. A minimal synthetic reproduction
// (`base_pair_tally_accumulates_base_a_diagnostic`) showed the
// pairwise tally mechanism itself works correctly for base A in
// isolation — so this checks one level lower, purely structural
// (annex bits only, no cross-partition/genome-level resolution): does
// base A even show up as *present* in minorant families' own
// `FamilyMask`s at all, at a rate proportional to the other 3 bases?
// If yes, the anomaly is specific to the pairwise/`single_form`
// resolution stage; if `has_base[0]` is itself near-zero here, the
// anomaly originates earlier, in `build_sibling_annex`/`central_base`.
let idx = KmerIndex::open("/Users/coissac/Sync/travail/__MOI__/obikmer/benchmark/global_index_presence")
.expect("open real index");
let layer_dirs = super::family_scan::sibling_layer_dirs(&idx).expect("layer dirs");
let mut has_base = [0u64; 4];
let mut minorant_total = 0u64;
let mut minorant_central_a = 0u64;
for layer_dir in &layer_dirs {
let annex = SiblingAnnex::open(&layer_dir.join(ANNEX_FILE_NAME)).unwrap();
for slot in 0..annex.len() {
let Some(mask) = annex.get(slot) else { continue };
if !mask.is_minorant() {
continue;
}
minorant_total += 1;
for b in 0..4u8 {
if mask.has(b) {
has_base[b as usize] += 1;
}
}
}
// Cross-check: for a sample of this layer's minorants, is the
// slot's *own* central base ever actually 'A' (bit 0)? — reads
// the real k-mer via the layer's MPHF, not just the mask.
let meta = PartitionMeta::load(layer_dir.parent().unwrap()).unwrap();
let mphf = MphfLayer::open(layer_dir, &meta.mode).unwrap();
for (order, kmer) in mphf.enumerate_kmers().take(2_000_000) {
let Some(mask) = annex.get(order) else { continue };
if mask.is_minorant() && central_base(kmer, idx.kmer_size()) == 0 {
minorant_central_a += 1;
}
}
}
println!(
"minorant_total={minorant_total} has_base(A,C,G,T)={has_base:?} minorant_central_a_sampled={minorant_central_a}"
);
}
#[test]
#[ignore]
fn diag_real_index_sankoff_bundle_base_a() {
// Structural check (`diag_real_index_base_a_representation`) shows
// base A fully, proportionally represented at the family level (~23M
// minorants, the *highest* of the four bases) — so the anomaly must
// be downstream, in `sankoff_bundle`'s actual pairwise resolution.
// Reproduces through the real cross-partition code path (not the
// minimal synthetic fixture, which showed the mechanism working in
// isolation), bounded by `--subsample` to stay fast.
let idx = KmerIndex::open("/Users/coissac/Sync/travail/__MOI__/obikmer/benchmark/global_index_presence")
.expect("open real index");
let n_genomes = idx.meta().genomes.len();
let exclude_mask = vec![false; n_genomes];
let bundle = idx.sankoff_bundle(Some(1_000_000), None, 0.5, &exclude_mask).expect("sankoff_bundle");
println!("base_pair_tally.same={:?}", bundle.base_pair_tally.same);
println!("base_pair_tally.counts={:?}", bundle.base_pair_tally.counts);
let alignment_a_count: usize = bundle.alignment.sequences.iter()
.map(|seq| seq.iter().filter(|&&b| b == b'A').count())
.sum();
println!("alignment raw 'A' byte count={alignment_a_count}");
}
#[test]
#[ignore]
fn diag_real_index_genome_mask_for_base_a_families() {
// Directly inspects the raw `genome_mask` array `sankoff_bundle`'s
// closures see, for the first few variable families where base A is
// present — to check whether A ever co-occurs as `single_form` in two
// *different* genomes at the same family at all (the precondition for
// `bp_same`/`bp_counts` to ever increment for base A), rather than
// reasoning about it further.
use std::sync::Arc;
let idx = KmerIndex::open("/Users/coissac/Sync/travail/__MOI__/obikmer/benchmark/global_index_presence")
.expect("open real index");
let n_parts = idx.n_partitions();
let n_genomes = idx.meta().genomes.len();
let with_counts = idx.meta().config.with_counts;
let k = idx.kmer_size();
let n_bits = n_parts.trailing_zeros() as usize;
let partition = obikpartitionner::KmerPartition::open_with_config(
idx.root_path(), idx.kmer_size(), idx.minimizer_size(), n_bits,
).unwrap();
let cache = Arc::new(super::cache::PartitionCache::build(&partition, n_parts, with_counts).unwrap());
let layer_dirs = super::family_scan::sibling_layer_dirs(&idx).unwrap();
let mut printed = 0usize;
for layer_dir in &layer_dirs {
super::family_scan::scan_layer_families(
layer_dir, n_parts, n_genomes, with_counts, k, &cache, &super::family_scan::Selection::All,
|_family_idx, mask, genome_mask| {
if printed >= 15 || !mask.has(0) || mask.family_size() < 2 {
return;
}
let single_a: Vec<usize> = (0..n_genomes).filter(|&g| genome_mask[g] == 1).collect();
let others: Vec<(usize, u8)> = (0..n_genomes)
.filter(|&g| genome_mask[g] != 0 && genome_mask[g] != 1)
.map(|g| (g, genome_mask[g]))
.collect();
println!(
"family bits={:#06b} size={} single_A_genomes={:?} other_nonzero_genomes={:?}",
mask.bits(), mask.family_size(), single_a, others,
);
printed += 1;
},
).unwrap();
if printed >= 15 {
break;
}
}
}
#[test]
#[ignore]
fn diag_plant_index_presence_matrix_sparsity() {
let idx = KmerIndex::open("/Users/coissac/travail/obiskim/data/phyloalps/phyloskims_sal_vac")
.expect("open real plant index");
let layer_dirs = super::family_scan::sibling_layer_dirs(&idx).expect("layer dirs");
println!("n_layers={}", layer_dirs.len());
// Sample a spread of layers rather than just the first few (partition
// order, not necessarily representative) — every 37th layer, capped at 20.
let sample: Vec<&std::path::PathBuf> = layer_dirs.iter().step_by(37).take(20).collect();
let mut total_ones: u64 = 0;
let mut total_cells: u64 = 0;
for layer_dir in &sample {
let mat = obicompactvec::PersistentBitMatrix::open(layer_dir).expect("open presence matrix");
let n = mat.n() as u64;
let n_cols = mat.n_cols() as u64;
let ones: u64 = mat.count_ones().iter().sum();
let cells = n * n_cols;
let density = ones as f64 / cells as f64;
println!(
"{}: n_slots={n} n_cols={n_cols} ones={ones} cells={cells} density={:.6} sparsity={:.6}",
layer_dir.display(), density, 1.0 - density,
);
total_ones += ones;
total_cells += cells;
}
let overall_density = total_ones as f64 / total_cells as f64;
println!(
"OVERALL sampled_layers={} total_ones={total_ones} total_cells={total_cells} density={:.6} sparsity={:.6}",
sample.len(), overall_density, 1.0 - overall_density,
);
}
#[test]
#[ignore]
fn diag_plant_index_multigenome_set_duplication() {
use std::collections::HashMap;
let idx = KmerIndex::open("/Users/coissac/travail/obiskim/data/phyloalps/phyloskims_sal_vac")
.expect("open real plant index");
let layer_dirs = super::family_scan::sibling_layer_dirs(&idx).expect("layer dirs");
// Same spread-sampling strategy as the sparsity diagnostic — every
// 37th layer, capped at 8 (this one is heavier per layer: builds a
// hash map of every distinct multi-genome set).
let sample: Vec<&std::path::PathBuf> = layer_dirs.iter().step_by(37).take(8).collect();
let mut total_rows: u64 = 0;
let mut total_multi_rows: u64 = 0;
let mut total_distinct_multi_sets: u64 = 0;
let mut total_singleton_rows: u64 = 0;
for layer_dir in &sample {
let mat = obicompactvec::PersistentBitMatrix::open(layer_dir).expect("open presence matrix");
let n = mat.n();
let n_cols = mat.n_cols();
let mut seen: HashMap<Vec<u32>, u32> = HashMap::new();
let mut singleton = 0u64;
let mut multi = 0u64;
let mut buf = vec![0u32; n_cols];
for slot in 0..n {
mat.fill_row(slot, &mut buf);
let set: Vec<u32> = (0..n_cols).filter(|&c| buf[c] != 0).map(|c| c as u32).collect();
match set.len() {
0 => {}
1 => singleton += 1,
_ => {
multi += 1;
*seen.entry(set).or_insert(0) += 1;
}
}
}
let distinct = seen.len() as u64;
println!(
"{}: n_slots={n} n_cols={n_cols} singleton_rows={singleton} multi_rows={multi} distinct_multi_sets={distinct} dedup_ratio={:.4}",
layer_dir.display(),
if multi > 0 { distinct as f64 / multi as f64 } else { 0.0 },
);
total_rows += n as u64;
total_singleton_rows += singleton;
total_multi_rows += multi;
total_distinct_multi_sets += distinct;
}
println!(
"OVERALL sampled_layers={} total_rows={total_rows} singleton_rows={total_singleton_rows} multi_rows={total_multi_rows} distinct_multi_sets={total_distinct_multi_sets} dedup_ratio={:.4}",
sample.len(),
if total_multi_rows > 0 { total_distinct_multi_sets as f64 / total_multi_rows as f64 } else { 0.0 },
);
}
#[test]
#[ignore]
fn diag_plant_index_cardinality_distribution() {
let idx = KmerIndex::open("/Users/coissac/travail/obiskim/data/phyloalps/phyloskims_sal_vac")
.expect("open real plant index");
let layer_dirs = super::family_scan::sibling_layer_dirs(&idx).expect("layer dirs");
let sample: Vec<&std::path::PathBuf> = layer_dirs.iter().step_by(37).take(8).collect();
let mut hist: Vec<u64> = Vec::new(); // hist[k] = number of rows with cardinality k
let mut total_rows: u64 = 0;
for layer_dir in &sample {
let mat = obicompactvec::PersistentBitMatrix::open(layer_dir).expect("open presence matrix");
let n = mat.n();
let n_cols = mat.n_cols();
if hist.len() < n_cols + 1 {
hist.resize(n_cols + 1, 0);
}
let mut buf = vec![0u32; n_cols];
for slot in 0..n {
mat.fill_row(slot, &mut buf);
let card = buf.iter().filter(|&&v| v != 0).count();
hist[card] += 1;
}
total_rows += n as u64;
}
println!("total_rows={total_rows}");
let mut cum: u64 = 0;
for (card, &count) in hist.iter().enumerate() {
if count == 0 { continue; }
cum += count;
println!(
"cardinality={card:3} count={count:10} frac={:.6} cumulative_frac={:.6}",
count as f64 / total_rows as f64,
cum as f64 / total_rows as f64,
);
}
}
#[test]
#[ignore]
fn bench_sparse_matrix_inmemory_construction_ram() {
use std::collections::HashMap;
use std::time::Instant;
let idx = KmerIndex::open("/Users/coissac/travail/obiskim/data/phyloalps/phyloskims_sal_vac")
.expect("open real plant index");
let layer_dirs = super::family_scan::sibling_layer_dirs(&idx).expect("layer dirs");
// A `layer_1` (the larger, merged layer) — pick the one already
// profiled: part_00018/index/layer_1, ~30.2M rows.
let layer_dir = layer_dirs.iter()
.find(|p| p.to_string_lossy().contains("part_00018") && p.to_string_lossy().contains("layer_1"))
.expect("expected layer not found — layer_dirs order may have changed");
let mat = obicompactvec::PersistentBitMatrix::open(layer_dir).expect("open presence matrix");
let n = mat.n();
let n_cols = mat.n_cols();
println!("layer={} n_slots={n} n_cols={n_cols}", layer_dir.display());
let rss_before = obisys::peak_rss_bytes();
let t0 = Instant::now();
// Plan design: is_multi flag (1 bit/row) + two SEPARATE arrays
// (singleton-only, multi-only), each sized to its own value range —
// not one array covering the whole dict_id space. In-memory prototype
// (u32 per entry, not yet bit-packed — the fixed-width primitive
// isn't built yet); this benchmark measures RAM + the split's real
// final size, not the construction-time representation's size.
let mut is_multi: Vec<bool> = Vec::with_capacity(n);
let mut singleton_array: Vec<u32> = Vec::new();
let mut multi_array: Vec<u32> = Vec::new();
let mut dedup: HashMap<Vec<u32>, u32> = HashMap::new();
let mut next_dict_id: u32 = 0;
let mut dict_values_bytes: u64 = 0; // running varint-encoded size estimate
let mut buf = vec![0u32; n_cols];
for slot in 0..n {
mat.fill_row(slot, &mut buf);
let set: Vec<u32> = (0..n_cols).filter(|&c| buf[c] != 0).map(|c| c as u32).collect();
match set.len() {
0 => { is_multi.push(false); singleton_array.push(0); } // shouldn't happen on a real built index; placeholder
1 => {
is_multi.push(false);
singleton_array.push(set[0]);
}
_ => {
is_multi.push(true);
let id = if let Some(&id) = dedup.get(&set) {
id
} else {
let id = next_dict_id;
next_dict_id += 1;
// varint size estimate: 1 byte per index < 128 (always
// true here, n_cols=91), matching the plan's encoding.
dict_values_bytes += set.iter().map(|&v| if v < 128 { 1 } else { 2 }).sum::<u64>();
dedup.insert(set, id);
id
};
multi_array.push(id);
}
};
}
let elapsed = t0.elapsed();
let rss_after = obisys::peak_rss_bytes();
let n_distinct_multi = dedup.len() as u64;
let n_singleton = singleton_array.len() as u64;
let n_multi = multi_array.len() as u64;
let is_multi_bytes = (n as u64).div_ceil(8);
let singleton_width_bits = (32 - (n_cols as u32).max(1).leading_zeros()).max(1);
let multi_width_bits = (32 - (n_distinct_multi as u32).max(1).leading_zeros()).max(1);
let singleton_array_bytes = (n_singleton * singleton_width_bits as u64).div_ceil(8);
let multi_array_bytes = (n_multi * multi_width_bits as u64).div_ceil(8);
let split_total = is_multi_bytes + singleton_array_bytes + multi_array_bytes + dict_values_bytes;
let dense_bytes = (n as u64 * n_cols as u64).div_ceil(8);
println!(
"elapsed={:?} rss_before={} rss_after={} rss_delta={}",
elapsed, fmt_mb(rss_before), fmt_mb(rss_after), fmt_mb(rss_after.saturating_sub(rss_before)),
);
println!(
"n_singleton={n_singleton} n_multi={n_multi} n_distinct_multi={n_distinct_multi} \
singleton_width_bits={singleton_width_bits} multi_width_bits={multi_width_bits}",
);
println!(
"is_multi_bytes={} singleton_array_bytes={} multi_array_bytes={} dict_values_bytes={} \
split_total_bytes={} ({:.1}x vs dense) dense_bytes={}",
fmt_mb(is_multi_bytes), fmt_mb(singleton_array_bytes), fmt_mb(multi_array_bytes),
fmt_mb(dict_values_bytes), fmt_mb(split_total),
dense_bytes as f64 / split_total as f64,
fmt_mb(dense_bytes),
);
}
fn fmt_mb(bytes: u64) -> String {
format!("{:.1}MB", bytes as f64 / (1024.0 * 1024.0))
}
#[test]
#[ignore]
fn bench_real_persistent_sparse_bit_matrix_on_disk_size() {
use std::time::Instant;
let idx = KmerIndex::open("/Users/coissac/travail/obiskim/data/phyloalps/phyloskims_sal_vac")
.expect("open real plant index");
let layer_dirs = super::family_scan::sibling_layer_dirs(&idx).expect("layer dirs");
let layer_dir = layer_dirs.iter()
.find(|p| p.to_string_lossy().contains("part_00018") && p.to_string_lossy().contains("layer_1"))
.expect("expected layer not found — layer_dirs order may have changed");
let dense = obicompactvec::PersistentBitMatrix::open(layer_dir).expect("open presence matrix");
let n = dense.n();
let n_cols = dense.n_cols();
println!("layer={} n_slots={n} n_cols={n_cols}", layer_dir.display());
let out_dir = std::env::temp_dir().join(format!("sparse_bench_{}", std::process::id()));
let _ = std::fs::remove_dir_all(&out_dir);
let rss_before = obisys::peak_rss_bytes();
let t0 = Instant::now();
let sparse = obicompactvec::PersistentSparseBitMatrixBuilder::build_from_dense(&dense, &out_dir)
.expect("build_from_dense")
.finish()
.expect("finish");
let elapsed = t0.elapsed();
let rss_after = obisys::peak_rss_bytes();
// Real on-disk size: sum of every file actually written under out_dir.
let mut sparse_bytes: u64 = 0;
for entry in std::fs::read_dir(&out_dir).unwrap() {
let entry = entry.unwrap();
let size = entry.metadata().unwrap().len();
println!(" {}: {}", entry.file_name().to_string_lossy(), fmt_mb(size));
sparse_bytes += size;
}
let dense_bytes = (n as u64 * n_cols as u64).div_ceil(8);
println!(
"elapsed={:?} rss_before={} rss_after={} rss_delta={}",
elapsed, fmt_mb(rss_before), fmt_mb(rss_after), fmt_mb(rss_after.saturating_sub(rss_before)),
);
println!(
"REAL on-disk: sparse_total={} dense={} ({:.1}x)",
fmt_mb(sparse_bytes), fmt_mb(dense_bytes), dense_bytes as f64 / sparse_bytes as f64,
);
// Sanity: reopen from disk (fresh mmap, not the just-built handle) and
// spot-check a handful of rows against the dense original.
drop(sparse);
let reopened = obicompactvec::PersistentSparseBitMatrix::open(&out_dir).expect("reopen");
let mut dense_buf = vec![0u32; n_cols];
for &slot in &[0usize, 1, 1000, n / 2, n - 1] {
dense.fill_row(slot, &mut dense_buf);
let sparse_row = reopened.row(slot);
for c in 0..n_cols {
assert_eq!(sparse_row[c], dense_buf[c] != 0, "slot {slot}, col {c}");
}
}
println!("spot-check rows: OK");
// ── Access-time comparison ──────────────────────────────────────────
// Both patterns matter in practice: sequential (a full-layer scan, the
// existing `scan_layer_families` shape) and random (a single-family
// cross-partition lookup, the entropy/`--shannon` shape). Same slot
// sequence used against both structures for a fair comparison.
const N_ACCESS: usize = 2_000_000;
// Deterministic xorshift, no extra dependency — fine for a benchmark's
// access pattern, not for anything security- or correctness-sensitive.
let mut rng_state: u64 = 0x9E3779B97F4A7C15;
let mut next_rand = move || {
rng_state ^= rng_state << 13;
rng_state ^= rng_state >> 7;
rng_state ^= rng_state << 17;
rng_state
};
let random_slots: Vec<usize> = (0..N_ACCESS).map(|_| (next_rand() as usize) % n).collect();
let sequential_slots: Vec<usize> = (0..N_ACCESS).map(|i| i % n).collect();
let mut buf = vec![0u32; n_cols];
let t = Instant::now();
for &slot in &sequential_slots {
dense.fill_row(slot, &mut buf);
}
let dense_seq = t.elapsed();
let t = Instant::now();
for &slot in &sequential_slots {
reopened.fill_row(slot, &mut buf);
}
let sparse_seq = t.elapsed();
let t = Instant::now();
for &slot in &random_slots {
dense.fill_row(slot, &mut buf);
}
let dense_rand = t.elapsed();
let t = Instant::now();
for &slot in &random_slots {
reopened.fill_row(slot, &mut buf);
}
let sparse_rand = t.elapsed();
println!(
"ACCESS ({N_ACCESS} rows) sequential: dense={:?} ({:.0}ns/row) sparse={:?} ({:.0}ns/row) — {:.2}x",
dense_seq, dense_seq.as_nanos() as f64 / N_ACCESS as f64,
sparse_seq, sparse_seq.as_nanos() as f64 / N_ACCESS as f64,
sparse_seq.as_secs_f64() / dense_seq.as_secs_f64(),
);
println!(
"ACCESS ({N_ACCESS} rows) random: dense={:?} ({:.0}ns/row) sparse={:?} ({:.0}ns/row) — {:.2}x",
dense_rand, dense_rand.as_nanos() as f64 / N_ACCESS as f64,
sparse_rand, sparse_rand.as_nanos() as f64 / N_ACCESS as f64,
sparse_rand.as_secs_f64() / dense_rand.as_secs_f64(),
);
// ── Column-major access, "just for fun" ─────────────────────────────
// The whole point of `DevDocMD/architecture/siblings.md`'s "Explicitly
// deferred" section: dense is genome-major (native, contiguous column
// access), sparse is k-mer-major (no column method at all — reading
// one column means decoding every row and keeping one bit each time).
// Extract one full column (all n rows) both ways.
let col = n_cols / 2;
let t = Instant::now();
let dense_col_view = dense.col_view(col);
let dense_col_ones: u64 = (0..n).filter(|&s| dense_col_view.get(s)).count() as u64;
let dense_col_time = t.elapsed();
let t = Instant::now();
let mut buf2 = vec![0u32; n_cols];
let sparse_col_ones: u64 = (0..n)
.filter(|&s| { reopened.fill_row(s, &mut buf2); buf2[col] != 0 })
.count() as u64;
let sparse_col_time = t.elapsed();
assert_eq!(dense_col_ones, sparse_col_ones, "column {col} popcount disagrees");
println!(
"COLUMN-MAJOR (col {col}, {n} rows) dense={:?} ({:.0}ns/row) sparse={:?} ({:.0}ns/row) — {:.1}x SLOWER on sparse",
dense_col_time, dense_col_time.as_nanos() as f64 / n as f64,
sparse_col_time, sparse_col_time.as_nanos() as f64 / n as f64,
sparse_col_time.as_secs_f64() / dense_col_time.as_secs_f64(),
);
let _ = std::fs::remove_dir_all(&out_dir);
}
#[test]
#[ignore]
fn diag_real_index_verify_kmer_resolution() {
use std::sync::Arc;
let idx = KmerIndex::open("/Users/coissac/Sync/travail/__MOI__/obikmer/benchmark/global_index_presence")
.expect("open real index");
let n_parts = idx.n_partitions();
let n_genomes = idx.meta().genomes.len();
let with_counts = idx.meta().config.with_counts;
let k = idx.kmer_size();
let n_bits = n_parts.trailing_zeros() as usize;
let partition = obikpartitionner::KmerPartition::open_with_config(
idx.root_path(), idx.kmer_size(), idx.minimizer_size(), n_bits,
).unwrap();
let cache = Arc::new(super::cache::PartitionCache::build(&partition, n_parts, with_counts).unwrap());
// The k-mer with A+other (mask 0b0011 = A+C)
let kmer_str = "AGCTAGCTATTGAGCCTGGTCCGTATGAAAC";
let kmer = canonical(kmer_str.as_bytes());
let rc_str = kmer_revcomp_to_string(&kmer, k);
println!("kmer_forward={} kmer_revcomp={}", kmer_str, rc_str);
println!("kmer partition={}", kmer.partition(n_parts));
// Check if this k-mer is found in its own partition
let dest_part = kmer.partition(n_parts);
if let Some(layer) = cache.find(dest_part, kmer) {
println!("FOUND in partition {} layer {}", dest_part, layer);
} else {
println!("NOT FOUND in partition {}", dest_part);
}
// Now check the variant with A at central position
// The mask is 0b0011 = A+C, so there should be a variant with A
// Let's generate the canonical neighbors and check which one has A
let mut found_variants = Vec::new();
for variant in kmer.central_canonical_neighbors() {
if variant == kmer {
continue;
}
let base = central_base(variant, k);
if base == 0 { // A
let v_part = variant.partition(n_parts);
println!("variant with A: partition={} found={}", v_part, cache.find(v_part, variant).is_some());
found_variants.push(variant);
}
}
// Also check the C variant
for variant in kmer.central_canonical_neighbors() {
if variant == kmer {
continue;
}
let base = central_base(variant, k);
if base == 1 { // C
let v_part = variant.partition(n_parts);
println!("variant with C: partition={} found={}", v_part, cache.find(v_part, variant).is_some());
}
}
}
fn kmer_to_string(kmer: &CanonicalKmer, k: usize) -> String {
let mut s = String::with_capacity(k);
for i in 0..k {
let n = kmer.nucleotide(i);
s.push(match n {
0 => 'A',
1 => 'C',
2 => 'G',
3 => 'T',
_ => unreachable!(),
});
}
s
}
fn kmer_revcomp_to_string(kmer: &CanonicalKmer, k: usize) -> String {
let rc = kmer.revcomp();
let mut s = String::with_capacity(k);
for i in 0..k {
let n = rc.nucleotide(i);
s.push(match n {
0 => 'A',
1 => 'C',
2 => 'G',
3 => 'T',
_ => unreachable!(),
});
}
s
}
-167
View File
@@ -1,167 +0,0 @@
ratio_ceiling: 0.5
cardinality_transitions:
- from: 0
to: 0
count: 122014779
probability: 0.870049812090934
- from: 0
to: 1
count: 17966272
probability: 0.12811195254940888
- from: 0
to: 2
count: 233683
probability: 0.001666321505518981
- from: 0
to: 3
count: 24109
probability: 0.00017191385413811494
- from: 0
to: 4
count: 0
probability: 0.0
- from: 1
to: 0
count: 17966272
probability: 0.967023782579572
- from: 1
to: 1
count: 602798
probability: 0.03244523972983381
- from: 1
to: 2
count: 9562
probability: 0.0005146688978673966
- from: 1
to: 3
count: 303
probability: 0.00001630879272681669
- from: 1
to: 4
count: 0
probability: 0.0
- from: 2
to: 0
count: 233683
probability: 0.9585500516842502
- from: 2
to: 1
count: 9562
probability: 0.039222603245442765
- from: 2
to: 2
count: 438
probability: 0.0017966429848885097
- from: 2
to: 3
count: 105
probability: 0.00043070208541847837
- from: 2
to: 4
count: 0
probability: 0.0
- from: 3
to: 0
count: 24109
probability: 0.9827171564831044
- from: 3
to: 1
count: 303
probability: 0.012350711286838137
- from: 3
to: 2
count: 105
probability: 0.004279949455834998
- from: 3
to: 3
count: 16
probability: 0.0006521827742224758
- from: 3
to: 4
count: 0
probability: 0.0
- from: 4
to: 0
count: 0
probability: 0.0
- from: 4
to: 1
count: 0
probability: 0.0
- from: 4
to: 2
count: 0
probability: 0.0
- from: 4
to: 3
count: 0
probability: 0.0
- from: 4
to: 4
count: 0
probability: 0.0
composition_transitions:
- from: 'A'
to: 'A'
count: 0
probability: 0.0
- from: 'A'
to: 'C'
count: 0
probability: 0.0
- from: 'A'
to: 'G'
count: 0
probability: 0.0
- from: 'A'
to: 'T'
count: 0
probability: 0.0
- from: 'C'
to: 'A'
count: 0
probability: 0.0
- from: 'C'
to: 'C'
count: 96389
probability: 0.9663832688335907
- from: 'C'
to: 'G'
count: 1290
probability: 0.012933368089671353
- from: 'C'
to: 'T'
count: 2063
probability: 0.020683363076737984
- from: 'G'
to: 'A'
count: 0
probability: 0.0
- from: 'G'
to: 'C'
count: 1290
probability: 0.005696948820201646
- from: 'G'
to: 'G'
count: 222720
probability: 0.9835848381669073
- from: 'G'
to: 'T'
count: 2427
probability: 0.010718213012891003
- from: 'T'
to: 'A'
count: 0
probability: 0.0
- from: 'T'
to: 'C'
count: 2063
probability: 0.007305266661709142
- from: 'T'
to: 'G'
count: 2427
probability: 0.008594223067362136
- from: 'T'
to: 'T'
count: 277909
probability: 0.9841005102709287