Skip to content

Re-run kb e2e-cost on arc: the published hg38 curve was priced against a whitelist STARsolo could not match #495

Description

@lhqing

This was generated by AI during triage.

Follow-up to #493 / #494. The code defect is fixed; this is the measurement half, which needs a human with arc.

What needs re-measuring

Every kb e2e-cost reading between 2026-07-16 and 2026-08-22 — the whole hg38 table in
docs/research/e2e-gate-runs.md — was taken with
--soloCBwhitelist pointed at a 17-byte file whose one line was 3M-february-2018. Every
synthesized read's barcode was off-list, every point exited 0, and the sweep records only a wall
time and a peak, so nothing showed.

The doc already carries a caveat block saying which numbers that touches. This issue is to replace
the caveat with a re-measured table.

What is actually in doubt

Wall clock — the number most affected. Per-read CB lookup ran against one entry instead of
6 794 880, with no 1MM correction, no UMI collapse and no Solo.out write. All of that scales with
depth, so the published ~21 s + 3 s per million reads is a floor, not the slope.

Peak RSS intercept. 31.1 GB omits whatever a 6.8 M-entry whitelist costs STARsolo, and omits the
five Solo count matrices that were allocated empty.

Flat-with-depth is the finding most likely to survive — the largest omitted term (the whitelist
itself) is constant in depth — but the per-read CB/UMI records and the Solo matrices are not, so
reading the fix does not certify the slope. This re-run is what settles it, and it is the one thing
worth checking first, because the recipe's memory defaults are sized against that intercept.

What is NOT in doubt, and does not need re-running

  • Every matrix assertion on that page. kb e2e and kb e2e-introns drive the composed
    Snakefile through run_composed, whose onlist rule materializes correctly, and
    test_the_correctness_arms_run_the_composed_snakefile forbids them from reaching the instrument.
    The recovery figures, 0-spurious/0-inflated, the strand inversion and the 40.7 % Gene-only loss
    all stand.
  • Log.final.out counts and STAR's reported sort requirement. CB matching does not gate
    alignment, so the record count is identical either way. The ~4.9 B/read sort figure is unaffected.
  • No published cell count or matrix density came through the defect — the sweep reads no Solo.out.

How to run it

The instrument is fixed and now refuses to run at all unless the sampled cells are members of the
list STAR receives, so a green run is self-certifying in a way the old one was not.

INSTRUMENT_VERSION moved to 2026.8.22, which means existing arc workdirs will not resume
their old points into the new curve — the stamp is in _cost_fingerprint. No need to delete
cost_sweep.partial.json; it will simply be ignored.

Same conditions as the 2026-08-11 run: hg38 + gencode_v50, 8 threads, the shipped
--outSAMtype BAM SortedByCoordinate, all-five counting, sweep to 250 M. Check
star_peak_rss_process is not bash before believing anything.

Why this is not ready-for-agent

It needs arc, a ~31 GB hg38 index and hours of STAR, and then a judgement about which published
figures to withdraw versus amend — the same judgement call that kept #493 from being fully
delegable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingready-for-humanRequires human implementation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions