Bumps core dependencies including ndarray, rand, hashbrown, niffler, ureq, sysinfo, indicatif, lru, and remove_dir_all. Adapts source code to accommodate breaking changes by migrating RNG initialization, adjusting HTTP response handling, and replacing the fs4 crate with standard library file locking. Adds a planning document for query benchmarking and sparse index regression tests.
4.3 KiB
Benchmark: query-path testing (discussion)
benchmark/Makefile exercises indexing, merge, and phylo distance
reconstruction against simulated bacterial genomes, but has no coverage of
obikmer query — the read-matching path — nor of the sparse packed
presence-matrix format (obikmer pack --sparse). This note captures the
planned extension.
Motivation
queryis untested end-to-end. A regression there would not be caught by the existingverify_presence/verify_merge_presencebranches, which only check index content against the.npztruth, never the query API.pack --sparseproduces a presence-matrix format documented (see siblings.md) as faster for single-row access (query) and slower for column-oriented access (phylo--metric).global_index_presence/built bymerge_presence.shis currently always packed dense (packing is a stage insidemerge, not a separatepackinvocation). There is no dense/sparse regression check.
Plan
New read source, independent of simulated_data/. Reusing
simulated_data/<species>/<strain>/reads_R1.fastq.gz for queries would bias
the test: those reads were already folded into the index being queried, with
the same sequencing-error draw. Query reads must come from a second,
independent iss generate run against the same reference genome(s) — new
random error draw, same underlying sequence — landing in a separate tree:
query_data/<species>/<strain>/reads_R1.fastq.gz, built by the existing
simulate_one.sh (unseeded, so a second invocation naturally draws different
reads).
Two specimens chosen as query sources (enough to catch a dense/sparse
regression without duplicating the exhaustive per-specimen coverage
verify_merge_presence already provides across all SPECIMENS):
Escherichia_coli--K-12_MG1655 (common, well-represented bacterium) and
Saccharolobus_islandicus--M.16.4 (the only archaeon in SPECIES — distant
lineage, different GC content, stresses the query path differently from a
close-relative match).
make_deps.py needs a QUERY_SPECIMENS list (explicit, short) and, for each,
an extra dependency line:
query_data/<species>/<strain>/reads_R1.fastq.gz: genomes/<genome>.fna.gz
distinct from the simulated_data/... rule for the same specimen.
Read count fixed at 100,000 read pairs per genome, independent of genome
size — unlike simulate_one.sh's simulated_data/ runs, which derive
n_reads from a fixed 15x coverage target. A query benchmark does not need
coverage-proportional depth; a fixed pair count keeps the two query runs
comparable to each other and keeps wall/RSS numbers meaningful across
genomes of very different sizes (bacterium vs archaeon). This likely needs a
dedicated simulate_query_one.sh (or a parameter to simulate_one.sh)
rather than reusing it unchanged, since n_reads is currently computed
in-script from genome size.
Phase 1 — sparse global index. New target
global_index_presence_sparse/index.done, built from global_index_presence/
via obikmer pack --sparse. Open question, to verify against the pack
implementation before writing the rule: does pack --sparse accept an
already dense-packed index in place (cp -r + repack), or does it require
the pre-pack column layout, forcing a dedicated merge run instead of reusing
global_index_presence/?
Phase 2 — query runs. For each of the two QUERY_SPECIMENS, run
obikmer query against both global_index_presence and
global_index_presence_sparse, capturing Reporter wall/RSS stats the same
way merge_presence.sh does (stderr capture + parse_reporter).
Phase 3 — dense/sparse regression. verify_query.py diffs the two query
JSON outputs per specimen (same matches, same per-genome presence
annotations) → .stats CSV (run,specimen,mismatches,pct), aggregated by
aggregate_stats.sh under a new query case. Any mismatch is a real
regression — dense and sparse must be content-identical, only I/O access
pattern differs.
Phase 4 — performance comparison. No dedicated script: the wall/RSS
columns from Phase 2's .stats files, aggregated, are the dense-vs-sparse
performance comparison (the expected win for query on sparse, per the pack --sparse help text).
count track is out of scope for the sparse branch: pack --sparse targets
presence matrices only (per CLI help), no count equivalent confirmed.