Files
obikmer/DevDocMD/implementation/benchmark_query_testing.md
T
Eric Coissac 89ea077456 Add benchmark pipeline for dense and sparse query testing
Introduces a complete query benchmark track to evaluate performance and verify consistency between dense and sparse index formats. Adds scripts to simulate fixed-size paired-end reads, pack a sparse presence index, execute queries in both modes, and capture wall time and RSS metrics. Includes a verification step that compares outputs by read ID to ensure content identity across parallel processing. Updates build configuration, documentation, and ignore patterns to support the new pipeline for two microbial specimens.
2026-08-20 13:59:12 +02:00

94 lines
4.4 KiB
Markdown

# Benchmark: query-path testing
`benchmark/Makefile` exercises indexing, merge, and phylo distance
reconstruction against simulated bacterial genomes. It now also covers
`obikmer query` — the read-matching path — and the sparse packed
presence-matrix format (`obikmer pack --sparse`), previously untested by
this pipeline.
## Motivation
- `query` had no end-to-end coverage. A regression there would not be caught
by `verify_presence`/`verify_merge_presence`, which only check index
*content* against the `.npz` truth, never the query API.
- `pack --sparse` produces a presence-matrix format documented (see
[siblings.md](../architecture/siblings.md)) as faster for single-row
access (query) and slower for column-oriented access (phylo `--metric`).
`global_index_presence/` built by `merge_presence.sh` is always packed
dense (packing is a stage inside `merge`, not a separate `pack`
invocation) — there was no dense/sparse regression check.
## Query read source
Query reads are independent of `simulated_data/` (which is folded into the
index being queried): reusing those reads would test against the exact
error draw the index was built from. `query_data/<species>/<strain>/` holds
a *second*, independent `iss generate` run against the same reference
genome, via `simulate_query_one.sh` — unseeded, so a second draw picks up
different sequencing errors than `simulate_one.sh`'s draw for the same
genome. Fixed at 100,000 read pairs per genome (not coverage-proportional
like the 15x used for `simulated_data/`), so wall/RSS numbers stay
comparable across genomes of very different sizes.
Two query-source specimens, hardcoded as `QUERY_SPECIMENS` in
`make_deps.py`: `Escherichia_coli--K-12_MG1655` (common, well-represented
bacterium) and `Saccharolobus_islandicus--M.16.4` (the only archaeon in
`SPECIES` — distant lineage, stresses the query path differently from a
close-relative match). Two is enough to catch a dense/sparse regression
without duplicating the exhaustive per-specimen coverage
`verify_merge_presence` already provides across all `SPECIMENS`.
## Sparse global index
`global_index_presence_sparse/` is built by `pack_sparse.sh`: copy
`global_index_presence/` wholesale, then `obikmer pack --sparse` in place.
This works directly because `merge`'s pack stage (`merge.rs:252`,
`pack_matrices(false)`) keeps the per-genome column files on disk after
dense-packing — `pack_sparse_bit_matrix` (`obicompactvec/src/bitmatrix/sparse.rs:447`)
reads those, is idempotent, and removes `matrix.pbmx` once the sparse form
is written, so `Persistent::open` falls through to the sparse format
afterward. No separate merge run needed.
## Query runs
`query_one.sh dense|sparse SPECIMEN` runs `obikmer query --count-missing`
against `global_index_presence` or `global_index_presence_sparse`, output
gzipped to `query_{dense,sparse}/SPECIMEN.fasta.gz`, Reporter wall/RSS
captured to `stats/query_{dense,sparse}/SPECIMEN.stats` (same
stderr-parsing convention as `merge_presence.sh`).
Flags: `--count-missing` only. `--mismatch` is a no-op today
(`query/mod.rs:212-213`, prints "not yet implemented, ignored") — left off
rather than tested for a feature that doesn't exist yet.
## Dense/sparse regression
`verify_query.py` compares the two query outputs per specimen, matched by
read id (not stream position — the query pipeline chunks input across
worker threads and doesn't guarantee output order). Compares `kmer_count`,
`kmer_missing`, and the full `kmer_strict_matches` map per read. Any
mismatch is a real regression: dense and sparse must be content-identical,
only I/O access pattern differs. `.stats``stats/verify_query/`,
aggregated by `aggregate_stats.sh query|verify_query`-style cases
(`query_dense`, `query_sparse`, `verify_query`).
## Performance comparison
No dedicated script: the wall/RSS columns from the `query_dense` and
`query_sparse` aggregated `.stats` CSVs are the dense-vs-sparse performance
comparison — the expected win for query on sparse, per the `pack --sparse`
help text.
## Scope
`count` track excluded from the sparse branch: `pack --sparse` targets
presence matrices only (per CLI help); `pack_matrices` leaves count
matrices untouched regardless of the `sparse` flag
(`obikindex/src/index.rs:308`).
## New Makefile targets
`simulate_query`, `pack_sparse`, `query_dense`, `query_sparse`,
`aggregate_query_dense`, `aggregate_query_sparse`, `verify_query`,
`aggregate_verify_query` — the last three folded into `all`.