Files
obikmer/DevDocMD/implementation/benchmark_query_testing.md
T
Eric Coissac 89ea077456 Add benchmark pipeline for dense and sparse query testing
Introduces a complete query benchmark track to evaluate performance and verify consistency between dense and sparse index formats. Adds scripts to simulate fixed-size paired-end reads, pack a sparse presence index, execute queries in both modes, and capture wall time and RSS metrics. Includes a verification step that compares outputs by read ID to ensure content identity across parallel processing. Updates build configuration, documentation, and ignore patterns to support the new pipeline for two microbial specimens.
2026-08-20 13:59:12 +02:00

4.4 KiB

Benchmark: query-path testing

benchmark/Makefile exercises indexing, merge, and phylo distance reconstruction against simulated bacterial genomes. It now also covers obikmer query — the read-matching path — and the sparse packed presence-matrix format (obikmer pack --sparse), previously untested by this pipeline.

Motivation

  • query had no end-to-end coverage. A regression there would not be caught by verify_presence/verify_merge_presence, which only check index content against the .npz truth, never the query API.
  • pack --sparse produces a presence-matrix format documented (see siblings.md) as faster for single-row access (query) and slower for column-oriented access (phylo --metric). global_index_presence/ built by merge_presence.sh is always packed dense (packing is a stage inside merge, not a separate pack invocation) — there was no dense/sparse regression check.

Query read source

Query reads are independent of simulated_data/ (which is folded into the index being queried): reusing those reads would test against the exact error draw the index was built from. query_data/<species>/<strain>/ holds a second, independent iss generate run against the same reference genome, via simulate_query_one.sh — unseeded, so a second draw picks up different sequencing errors than simulate_one.sh's draw for the same genome. Fixed at 100,000 read pairs per genome (not coverage-proportional like the 15x used for simulated_data/), so wall/RSS numbers stay comparable across genomes of very different sizes.

Two query-source specimens, hardcoded as QUERY_SPECIMENS in make_deps.py: Escherichia_coli--K-12_MG1655 (common, well-represented bacterium) and Saccharolobus_islandicus--M.16.4 (the only archaeon in SPECIES — distant lineage, stresses the query path differently from a close-relative match). Two is enough to catch a dense/sparse regression without duplicating the exhaustive per-specimen coverage verify_merge_presence already provides across all SPECIMENS.

Sparse global index

global_index_presence_sparse/ is built by pack_sparse.sh: copy global_index_presence/ wholesale, then obikmer pack --sparse in place. This works directly because merge's pack stage (merge.rs:252, pack_matrices(false)) keeps the per-genome column files on disk after dense-packing — pack_sparse_bit_matrix (obicompactvec/src/bitmatrix/sparse.rs:447) reads those, is idempotent, and removes matrix.pbmx once the sparse form is written, so Persistent::open falls through to the sparse format afterward. No separate merge run needed.

Query runs

query_one.sh dense|sparse SPECIMEN runs obikmer query --count-missing against global_index_presence or global_index_presence_sparse, output gzipped to query_{dense,sparse}/SPECIMEN.fasta.gz, Reporter wall/RSS captured to stats/query_{dense,sparse}/SPECIMEN.stats (same stderr-parsing convention as merge_presence.sh).

Flags: --count-missing only. --mismatch is a no-op today (query/mod.rs:212-213, prints "not yet implemented, ignored") — left off rather than tested for a feature that doesn't exist yet.

Dense/sparse regression

verify_query.py compares the two query outputs per specimen, matched by read id (not stream position — the query pipeline chunks input across worker threads and doesn't guarantee output order). Compares kmer_count, kmer_missing, and the full kmer_strict_matches map per read. Any mismatch is a real regression: dense and sparse must be content-identical, only I/O access pattern differs. .statsstats/verify_query/, aggregated by aggregate_stats.sh query|verify_query-style cases (query_dense, query_sparse, verify_query).

Performance comparison

No dedicated script: the wall/RSS columns from the query_dense and query_sparse aggregated .stats CSVs are the dense-vs-sparse performance comparison — the expected win for query on sparse, per the pack --sparse help text.

Scope

count track excluded from the sparse branch: pack --sparse targets presence matrices only (per CLI help); pack_matrices leaves count matrices untouched regardless of the sparse flag (obikindex/src/index.rs:308).

New Makefile targets

simulate_query, pack_sparse, query_dense, query_sparse, aggregate_query_dense, aggregate_query_sparse, verify_query, aggregate_verify_query — the last three folded into all.