Introduces a complete query benchmark track to evaluate performance and verify consistency between dense and sparse index formats. Adds scripts to simulate fixed-size paired-end reads, pack a sparse presence index, execute queries in both modes, and capture wall time and RSS metrics. Includes a verification step that compares outputs by read ID to ensure content identity across parallel processing. Updates build configuration, documentation, and ignore patterns to support the new pipeline for two microbial specimens.
4.4 KiB
Benchmark: query-path testing
benchmark/Makefile exercises indexing, merge, and phylo distance
reconstruction against simulated bacterial genomes. It now also covers
obikmer query — the read-matching path — and the sparse packed
presence-matrix format (obikmer pack --sparse), previously untested by
this pipeline.
Motivation
queryhad no end-to-end coverage. A regression there would not be caught byverify_presence/verify_merge_presence, which only check index content against the.npztruth, never the query API.pack --sparseproduces a presence-matrix format documented (see siblings.md) as faster for single-row access (query) and slower for column-oriented access (phylo--metric).global_index_presence/built bymerge_presence.shis always packed dense (packing is a stage insidemerge, not a separatepackinvocation) — there was no dense/sparse regression check.
Query read source
Query reads are independent of simulated_data/ (which is folded into the
index being queried): reusing those reads would test against the exact
error draw the index was built from. query_data/<species>/<strain>/ holds
a second, independent iss generate run against the same reference
genome, via simulate_query_one.sh — unseeded, so a second draw picks up
different sequencing errors than simulate_one.sh's draw for the same
genome. Fixed at 100,000 read pairs per genome (not coverage-proportional
like the 15x used for simulated_data/), so wall/RSS numbers stay
comparable across genomes of very different sizes.
Two query-source specimens, hardcoded as QUERY_SPECIMENS in
make_deps.py: Escherichia_coli--K-12_MG1655 (common, well-represented
bacterium) and Saccharolobus_islandicus--M.16.4 (the only archaeon in
SPECIES — distant lineage, stresses the query path differently from a
close-relative match). Two is enough to catch a dense/sparse regression
without duplicating the exhaustive per-specimen coverage
verify_merge_presence already provides across all SPECIMENS.
Sparse global index
global_index_presence_sparse/ is built by pack_sparse.sh: copy
global_index_presence/ wholesale, then obikmer pack --sparse in place.
This works directly because merge's pack stage (merge.rs:252,
pack_matrices(false)) keeps the per-genome column files on disk after
dense-packing — pack_sparse_bit_matrix (obicompactvec/src/bitmatrix/sparse.rs:447)
reads those, is idempotent, and removes matrix.pbmx once the sparse form
is written, so Persistent::open falls through to the sparse format
afterward. No separate merge run needed.
Query runs
query_one.sh dense|sparse SPECIMEN runs obikmer query --count-missing
against global_index_presence or global_index_presence_sparse, output
gzipped to query_{dense,sparse}/SPECIMEN.fasta.gz, Reporter wall/RSS
captured to stats/query_{dense,sparse}/SPECIMEN.stats (same
stderr-parsing convention as merge_presence.sh).
Flags: --count-missing only. --mismatch is a no-op today
(query/mod.rs:212-213, prints "not yet implemented, ignored") — left off
rather than tested for a feature that doesn't exist yet.
Dense/sparse regression
verify_query.py compares the two query outputs per specimen, matched by
read id (not stream position — the query pipeline chunks input across
worker threads and doesn't guarantee output order). Compares kmer_count,
kmer_missing, and the full kmer_strict_matches map per read. Any
mismatch is a real regression: dense and sparse must be content-identical,
only I/O access pattern differs. .stats → stats/verify_query/,
aggregated by aggregate_stats.sh query|verify_query-style cases
(query_dense, query_sparse, verify_query).
Performance comparison
No dedicated script: the wall/RSS columns from the query_dense and
query_sparse aggregated .stats CSVs are the dense-vs-sparse performance
comparison — the expected win for query on sparse, per the pack --sparse
help text.
Scope
count track excluded from the sparse branch: pack --sparse targets
presence matrices only (per CLI help); pack_matrices leaves count
matrices untouched regardless of the sparse flag
(obikindex/src/index.rs:308).
New Makefile targets
simulate_query, pack_sparse, query_dense, query_sparse,
aggregate_query_dense, aggregate_query_sparse, verify_query,
aggregate_verify_query — the last three folded into all.