Add sparse storage variant to PersistentCompactIntMatrix
Introduce a new `Sparse` format alongside existing `Columnar` and `Packed` variants, enabling optimized row-major pairwise counting for distance and similarity metrics via the `CountPartials` trait. Update storage detection priorities, extend matrix dispatch logic to sparse backends, and correct diagonal/off-diagonal formulas in bit matrix partial computations. Expand layer APIs with format-agnostic `nonzero_iter`, update usage documentation for the `--sparse` flag, and add comprehensive tests verifying roundtrip integrity and metric equivalence against dense implementations.
This commit is contained in:
@@ -933,3 +933,92 @@ mismatches. The dense/sparse performance gap is gone — previously sparse
|
||||
systematic gap. `pack --sparse`'s claimed query win isn't confirmed
|
||||
outright by this (sparse should arguably now *beat* dense on truly sparse
|
||||
real data, not just tie), but the pathological regression is fixed.
|
||||
|
||||
## `PersistentCompactIntMatrix::Sparse` — implemented (2026-08-26)
|
||||
|
||||
Closes the gap flagged throughout this document ("no sparse count format
|
||||
exists yet", `traits.rs:9-12`'s "Explicitly deferred"): `obicompactvec`
|
||||
already had `PersistentSparseCompactIntMatrix` (row-major, built on top of
|
||||
`PersistentSparseBitMatrix` as its "which columns are non-zero" support,
|
||||
values *not* deduplicated — see that struct's own doc comment), but it was
|
||||
never wired into `PersistentCompactIntMatrix`, the dense-dispatching enum
|
||||
every real consumer (`TypedLayer<PersistentCompactIntMatrix>`,
|
||||
`KmerLayer::Count`) actually holds. Concretely: `kmer_index.rs::
|
||||
pack_matrices(sparse=true)` already called `pack_sparse_compact_int_matrix`
|
||||
on every layer's `counts/` — but `PersistentCompactIntMatrix::open` had no
|
||||
code path back to what that just wrote, so a `Count` layer became
|
||||
unreadable ("no count matrix found ... run 'obikmer upgrade'") the moment
|
||||
anyone ran `pack --sparse` on an index with count layers. Root cause, not a
|
||||
workaround: add the missing `Sparse` variant.
|
||||
|
||||
- **Enum + dispatch** (`intmatrix.rs`): `PersistentCompactIntMatrix::Sparse
|
||||
(PersistentSparseCompactIntMatrix)`, detected in `open`/`detect_storage`
|
||||
via a `singleton_values.pciv` marker (mirrors `PersistentBitMatrix`'s own
|
||||
`sparse_meta.json` check), reported via `storage_kind()`. `col`/
|
||||
`col_view`/`col_persist` panic/`Unsupported` on `Sparse`, same convention
|
||||
as the bit side. `sub_matrix`/`fill_sub_matrix` and `nonzero_iter`
|
||||
unified the same way `PersistentBitMatrix`'s already are (drain
|
||||
`nonzero_iter`, one traversal per format — see "Implemented
|
||||
(2026-08-20)" above); `nonzero_iter` had to become `Box<dyn Iterator<...>>`
|
||||
for the same reason (`Columnar`/`Packed`/`Sparse` are different concrete
|
||||
types). No change needed in `obikindex` at all — `KmerLayer::Count`
|
||||
already only ever holds `TypedLayer<PersistentCompactIntMatrix>`, so the
|
||||
enum absorbing `Sparse` fixes the unreadable-layer bug for free, same as
|
||||
`PersistentBitMatrix::Sparse` already did on the presence side.
|
||||
|
||||
- **`CountPartials`, non-naive** (`sparse_intmatrix.rs`): unlike
|
||||
`PersistentSparseBitMatrix`'s dict-driven `col_weights_and_pair_counts`,
|
||||
values here aren't deduplicated (two rows can share the same non-zero
|
||||
column set via the same `dict_id` while carrying different counts), so
|
||||
the "weight by how many rows share a dict entry" shortcut doesn't carry
|
||||
over. What does: a single row-major pass (`row_major_pairwise`, decodes
|
||||
each row once via `for_each_cell_in_row`, nests over that row's own
|
||||
co-present columns) — `O(Σ k̄²)` over populated rows instead of the naive
|
||||
`O(n_cols² × n)` column-pair rescan, same complexity class as the bit
|
||||
side minus the dict multiplicity discount. Kernels used: `min(a,b)`
|
||||
(bray, relfreq-bray — both vanish when either side is absent, so no
|
||||
correction needed), `a·b` and `√(a·b)` (euclidean/relfreq-euclidean and
|
||||
hellinger — these *do* need a correction, reconstructed from per-column
|
||||
marginals via `Σ(a-b)² = Σa²+Σb²-2Σab`, since `(a-0)² = a² ≠ 0` unlike
|
||||
the `min`-based formulas). `threshold_jaccard(1)` shortcuts straight to
|
||||
`support`'s own `BitPartials::partial_jaccard` (threshold 1 is exactly
|
||||
presence); `threshold_jaccard(0)` is closed-form (every `u32` is `≥ 0`).
|
||||
|
||||
- **Two pre-existing bugs found and fixed while wiring the `threshold==1`
|
||||
shortcut** (`bitmatrix/sparse.rs`, `BitPartials for
|
||||
PersistentSparseBitMatrix`, present since the 2026-08-15 implementation
|
||||
above, never caught because no test compared `Sparse`'s raw `partial_*`
|
||||
output against dense on real data — only the diagonal-blind
|
||||
`jaccard_dist_matrix`/`hamming_dist_matrix` finalisations were tested):
|
||||
1. `partial_jaccard`'s diagonal was `(0, 2×col_weights[i])` instead of a
|
||||
genuine self-comparison `(col_weights[i], col_weights[i])` —
|
||||
`col_weights_and_pair_counts`'s `inter` never pairs a column with
|
||||
itself by construction.
|
||||
2. `partial_hamming`'s off-diagonal formula itself was wrong: `total -
|
||||
union` (count of rows where *neither* column is present) instead of
|
||||
the actual Hamming distance `col_weights[i] + col_weights[j] -
|
||||
2×inter[i,j]` (symmetric-difference size). Only coincides with the
|
||||
correct value when `col_weights[i] + col_weights[j] == total`, so
|
||||
small/synthetic test data could easily have hidden it.
|
||||
|
||||
Neither surfaced through `jaccard_dist_matrix`/`hamming_dist_matrix`
|
||||
(both explicitly zero their own diagonal at finalisation, and the
|
||||
off-diagonal `partial_hamming` bug had gone untested against dense
|
||||
entirely) — only visible to a caller of the raw `partial_*` methods
|
||||
directly, which is exactly what `partial_threshold_jaccard(1)`'s new
|
||||
shortcut became. Fixed at the source, not patched around at the call
|
||||
site; regression test added:
|
||||
`tests::sparse::partial_jaccard_and_hamming_match_dense_including_diagonal`.
|
||||
|
||||
- **Tests**: `tests::intmatrix::sparse_roundtrip_matches_columnar`/
|
||||
`sparse_roundtrip_from_packed` (the `open`-dispatch fix, both build
|
||||
paths); `tests::intmatrix::sparse_count_partials_match_dense` (all six
|
||||
`CountPartials` formulas, thresholds 0/1/2/3, against `Columnar` on
|
||||
asymmetric-presence data — this is what caught the diagonal gap in the
|
||||
int side's own new code before it shipped, the same way it exposed the
|
||||
two pre-existing bit-side bugs above); `obikindex`'s
|
||||
`count_layer_transparently_reads_sparse_after_pack` — the actual
|
||||
end-to-end regression test for the original "layer unreadable after
|
||||
`pack --sparse`" bug, built → packed sparse → reopened, compared against
|
||||
the pre-pack dense read. `cargo test -p obicompactvec -p obikindex`:
|
||||
green, no regressions (180 + 12 tests).
|
||||
|
||||
@@ -292,16 +292,18 @@ Pass 1 — byte max, SIMD-vectorizable, O(n)
|
||||
|
||||
## Matrix types
|
||||
|
||||
Both matrix types are enums behind a transparent API — the caller never matches on the variant. `PersistentCompactIntMatrix` has two variants (`Columnar`, `Packed`). `PersistentBitMatrix` has four:
|
||||
Both matrix types are enums behind a transparent API — the caller never matches on the variant. `PersistentCompactIntMatrix` has three variants (`Columnar`, `Packed`, `Sparse`). `PersistentBitMatrix` has four:
|
||||
|
||||
| Variant | Storage | When |
|
||||
|---|---|---|
|
||||
| `Columnar` | one `.pbiv`/`.pciv` file per column + `meta.json` | build-time default (`*Builder::new`) |
|
||||
| `Packed` | single `matrix.pbmx` mmap file | query-optimised, produced by `pack_bit_matrix`/`pack_compact_int_matrix` |
|
||||
| `Sparse` (bit only) | `sparse_meta.json` + PFIV/Elias-Fano component files, row-major | `pack --sparse`; see [siblings.md](../architecture/siblings.md) for the sparse-vs-dense access-pattern trade-off |
|
||||
| `Packed` | single `matrix.pbmx`/`matrix.pcmx` mmap file | query-optimised, produced by `pack_bit_matrix`/`pack_compact_int_matrix` |
|
||||
| `Sparse` | bit: `sparse_meta.json` + PFIV/Elias-Fano component files, row-major. Int: same support files (built on `PersistentSparseBitMatrix` internally) plus `singleton_values.pciv`/`multi_values.pciv`/`multi_offsets` for the per-row, non-deduplicated values | `pack --sparse`; see [siblings.md](../architecture/siblings.md) for the sparse-vs-dense access-pattern trade-off |
|
||||
| `Implicit` (bit only) | no file at all | mono-genome presence layers — `n_cols` is always reported as `1`, every value is `true` |
|
||||
|
||||
`PersistentBitMatrix::open(layer_dir)` auto-detects the variant, in order: `matrix.pbmx` → Packed, `presence/meta.json` → Columnar, `presence/sparse_meta.json` → Sparse, `layer_meta.json` (no presence dir at all) → Implicit. `col_view`/`col`/`sub_matrix` panic on `Sparse`/`Implicit` where the operation has no direct-slice equivalent (Sparse is k-mer-major, not column-major; Implicit has no backing storage) — callers needing per-column data on those variants go through `row`/`fill_row`.
|
||||
`PersistentBitMatrix::open(layer_dir)` auto-detects the variant, in order: `matrix.pbmx` → Packed, `presence/meta.json` → Columnar, `presence/sparse_meta.json` → Sparse, `layer_meta.json` (no presence dir at all) → Implicit. `PersistentCompactIntMatrix::open(layer_dir)` mirrors the same priority order minus `Implicit` (there's no implicit count matrix — counts always have at least one on-disk column): `matrix.pcmx` → Packed, `counts/meta.json` → Columnar, `counts/singleton_values.pciv` → Sparse. `col_view`/`col`/`sub_matrix` panic on `Sparse`/`Implicit` where the operation has no direct-slice equivalent (Sparse is k-mer-major, not column-major; Implicit has no backing storage) — callers needing per-column data on those variants go through `row`/`fill_row`.
|
||||
|
||||
Unlike the bit side, `PersistentSparseCompactIntMatrix`'s values are *not* deduplicated across rows — two rows can share the same non-zero column set (same `dict_id` in the shared support) while carrying different counts — so its `CountPartials` impl can't reuse the support's dict-multiplicity shortcut the way `BitPartials for PersistentSparseBitMatrix` does. It still avoids the naive `O(n_cols² × n)` column-pair scan via a single row-major pass (`row_major_pairwise` in `sparse_intmatrix.rs`), reconstructing the squared-difference formulas (`euclidean`/`relfreq_euclidean`/`hellinger`) from per-column marginals via `Σ(a-b)² = Σa²+Σb²-2Σab` — see [siblings.md](../architecture/siblings.md)'s "`PersistentCompactIntMatrix::Sparse` — implemented" entry for the full derivation.
|
||||
|
||||
`col_view(c)` returns the appropriate view directly:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user