Files
obikmer/DevDocMD/implementation/select.md
T
Eric Coissac 579cfd8752 docs: document partition runner deadlock fix and select CLI updates
Update documentation to reflect architectural fixes in the NUMA partition runner and implementation changes in the merge and select modules. The partition runner now wraps task execution in `catch_unwind`, emitting panic events that count toward completion thresholds to prevent controller deadlocks, while re-raising the first payload on return. CLI documentation for select has been updated to mandate output directories, introduce default hard-linking with automatic fallback, and document presence aggregation optimizations. Known issues such as self-merge crashes and sparse-packed index gaps are now documented as failing fast due to these improvements.
2026-08-28 23:25:46 +02:00

255 lines
9.1 KiB
Markdown

# `select` — column projection and aggregation
`select` transforms an index by operating on its **genome columns**: projecting a
subset of columns, aggregating groups of genomes into synthetic columns, or both.
It is the column-axis counterpart of `filter` (row-axis operations).
Following relational algebra conventions:
| Command | Relational operation | Axis |
|----------|---------------------|----------|
| `filter` | σ — selection | rows (k-mers) |
| `select` | π — projection | columns (genomes) |
The two commands compose naturally: run `filter` first to restrict the kmer set,
then `select` to reshape the genome columns.
`select` never changes the kmer set. The MPHF and `unitigs.bin` of each layer
are preserved unchanged; only the data matrices are rewritten.
---
## Synopsis
```sh
obikmer select <input-index>
--output <dir>
[--group <name>:<pred> ...]
[--group-op <name>:<op> ...]
[--aggregate-by <key> ]
[--aggregate-op <op> ]
[--select <col1,col2,...> ]
[--presence-threshold <N> ]
[--force-copy ]
```
---
## Output destination
`--output <dir>` is required — `select` always writes a new index; there is no
`--in-place` mode (2026-08-28: never implemented, removed from the design). The
source index is unchanged.
Each layer's kmer-identity files (`mphf.bin`/`unitigs.bin`/`evidence.bin`/
`unitigs.bin.idx`/`fingerprint.bin`/`layer_meta.json`) are never rewritten by a
column projection/aggregation, so they are hard-linked into the output rather
than copied — no extra disk for them even on a large index. Falls back to a
real copy automatically if linking fails (different filesystems); `--force-copy`
forces a real copy always, for an output that must survive independently of the
source on disk (a hard link shares the same inode — rewriting one path outside
`select` itself would affect the other). Only the `presence`/`counts`
subdirectory is ever a genuinely new, independent file.
To replace an index with a selected version of itself, select to a temporary
directory and swap it in (`rm -rf INDEX && mv INDEX.tmp INDEX`) — the case
`--in-place` used to cover.
---
## Defining output columns
### Named groups — `--group`
```
--group <name>:<pred>
```
Defines a named group of genomes using the same predicate syntax as `filter`.
Repeatable; a genome can belong to several groups.
```sh
--group "pub:species=Betula_pubescens"
--group "nan:species=Betula_nana"
```
### Per-group operator — `--group-op`
```
--group-op <name>:<op>
```
Assigns an aggregation operator to a named group. Optional; if absent, the
default operator applies (see below).
```sh
--group-op "pub:any"
--group-op "nan:all"
```
### Shorthand — `--aggregate-by` / `--aggregate-op`
`--aggregate-by <key>` automatically creates one group per unique value of the
metadata key `<key>`. Equivalent to one `--group <val>:<key>=<val>` per distinct
value. `--aggregate-op <op>` sets the operator for all auto-generated groups.
`--aggregate-by` and `--group` are mutually exclusive.
### Column selection and ordering — `--select`
```
--select col1,col2,...
```
Lists the output columns in order. Each element is either a group name (defined
by `--group` or generated by `--aggregate-by`) or a genome label from the source
index (pass-through, no aggregation).
**Default when `--select` is absent:**
all defined groups in declaration order (for `--group`), or all generated groups
in metadata-value order (for `--aggregate-by`). Individual genomes not in any
group are excluded unless named explicitly.
**When neither `--group` nor `--aggregate-by` is specified:**
`--select` can still reference genome labels for pure column projection (no
aggregation). If `--select` is also absent, all genomes are output unchanged
(identity transform — useful combined with row filtering via a prior `filter`
run).
---
## Aggregation operators
| Operator | Input | Output | Semantics |
|----------|-------------|----------|-----------|
| `any` | pres / count | presence | 1 if ≥ 1 genome in group carries the k-mer |
| `all` | pres / count | presence | 1 if every genome in group carries the k-mer |
| `none` | pres / count | presence | 1 if no genome in group carries the k-mer |
| `sum` | count | count | sum of counts across the group |
| `min` | count | count | minimum count |
| `max` | count | count | maximum count |
**Default operator:**
- Presence index: `any`
- Count index: `sum`
Logical operators (`any`/`all`/`none`) on a count index use
`--presence-threshold N` (default 0): a genome "carries" the k-mer if its count
is > N.
**Output index type:**
- If the source is a presence index, the output is always a presence index.
- If the source is a count index and every output column uses a logical operator
or is a pass-through from a presence source, the output is a presence index.
- Otherwise (at least one arithmetic operator on a count source), the output is
a count index.
---
## Behaviour for edge cases
| Situation | Behaviour |
|-----------|-----------|
| Genome missing the metadata key in `--aggregate-by` | genome ignored (no `NA` group) |
| Genome in multiple groups | contributes independently to each |
| `--group-op` references undefined group | error |
| `--select` element is neither group name nor genome label | error |
| `--output` and `--in-place` both specified | error |
| Neither `--output` nor `--in-place` | error |
| Group with zero matching genomes | column is all-zeros (or all-ones for `none`) |
---
## Examples
### Aggregate by metadata group, default operators
```sh
obikmer select myindex --output out --aggregate-by group
# one column per unique value of "group"; presence→any, count→sum
```
### Named groups with different operators
```sh
obikmer select myindex --output out \
--group "pub:species=Betula_pubescens" \
--group "nan:species=Betula_nana" \
--group-op "pub:any" \
--group-op "nan:all" \
--select "pub,nan"
```
### Mix aggregated group and individual genome
```sh
obikmer select myindex --output out \
--group "A:group=A" \
--select "A,Betula_nana--IGA-24-39"
```
### Pure column projection (no aggregation)
```sh
obikmer select myindex --output out \
--select "Betula_nana--TROM-V-149986,Betula_nana--AG-P04-25-01"
```
### Compose with filter
```sh
# Step 1: keep only B. nana-specific k-mers
obikmer filter myindex --output filtered \
--ingroup "species=Betula_nana" --outgroup "*"
# Step 2: aggregate genome columns by collection site
obikmer select filtered --output final --aggregate-by site
```
---
## Implementation notes
`select` does not rebuild the MPHF. Every partition is processed independently
(`PartitionRunner`), each writing its own output layers; no cross-partition
synchronisation is needed.
For each layer in each partition (`obikselect::select_layer::select_partition`):
1. `copy_layer_files` hard-links the source layer's kmer-identity files
(`mphf.bin`/`unitigs.bin`/`evidence.bin`/`unitigs.bin.idx`/
`fingerprint.bin`/`layer_meta.json`) into the destination — never a real
copy unless linking fails or `--force-copy` is given.
2. A new data matrix is built with M columns (M = number of output columns),
under a fresh `presence/`/`counts/` subdirectory (never touching the
source's own).
3. **Presence source (2026-08-28: `batch_presence_counts`)**: one shared pass
over the source bit matrix computes every output group's presence count at
once — row-major native for a `Sparse` source (`for_each_genome_in_row`,
which has no column representation to read a `col_view` from at all — the
reason this replaced the old per-group loop, not just an optimisation of
it), deduplicated column-major (one `col_view` per *distinct referenced
column*, not per group) for `Columnar`/`Packed`. Every `AggOp` for a bit
matrix is then a cheap derivation of that one count vector (`sum` = the
count itself, `any`/`max` = `count ≥ 1`, `all`/`min` = `count == group
size`, `none` = `count == 0`) — see
`obikselect::select_layer::agg_result_from_count`.
4. **Count source**: unchanged, one `col_view`-driven pass per output column
via `MatrixGroupOps``sum`/`min`/`max` are genuine per-value reductions
for a count matrix, not derivable from a single presence count the way
they are for a bit matrix.
5. `index.meta` is rewritten with the new genome list and updated `with_counts`.
### Known gap (not yet fixed, 2026-08-28)
Step 4 above still panics (`col_view() not available on Sparse
PersistentCompactIntMatrix`) if the source is a **count** index packed
sparse — `batch_presence_counts`' row-major treatment was only ported to the
bit-matrix (`Presence`) case, since that was the one actually blocking a real
benchmark run. `select`/`filter` on a sparse-packed count index still hits
this; the fix would follow the same shape (a `PersistentSparseCompactIntMatrix`
row-major decode, analogous to `for_each_genome_in_row`), just not done. Since
`obisys::numa::runner::PartitionRunner`'s panic-propagation fix (see
`architecture/numa_partition_runner.md`), this at least fails fast (process
panic, exit 101) instead of hanging.