Skip to content

chore(metrics): evaluate realized benefits of epics #2164 and #2172 #2200

Description

@bdfinst

Context

Both nWave-comparison epics have shipped:

What was measured before building: #2165's cross-checkpoint duplication ceiling (20.69% byte-estimate at the Step 6 backstop; one epic, theoretical, not observed) and #2170's report-size cut (59.0%, one round). Nothing has been measured after. The superseded #2174 required a completion-divergence rate over ≥5 sessions before deciding on blocking; #2188 dropped that, so it was never run.

North Star: a change that can't name the friction it removes doesn't ship. This epic checks whether each mechanism removes friction in practice, and retires or tunes what doesn't.

Goal

For every shipped mechanism: one measured number from real sessions, compared against a threshold written down before the data is read, ending in a keep / tune / retire decision.

Rules

  • Data source: one local machine. All session data is gathered from the maintainer's local machine, where .claude/metrics/ and ~/.claude/metrics/ persist. Cloud sessions are out of scope because their containers, and the metrics in them, are recycled.
  • Real session: a /build or /code-review run on actual repo work on that machine. Eval fixtures, dry runs and replays are not real sessions; a slice that uses one must label it as such.
  • Pre-register metrics, thresholds and the data window in slice 0 before reading any post-merge data.
  • Measure, don't build. If an instrument is missing, the only in-scope code is the smallest logging addition that makes the number exist.
  • Name each instrument (script, JSONL stream, or eval) next to every number.
  • A null result is a valid outcome. Record it so the question doesn't get reopened from scratch.

Slices

  1. Pre-register metrics and thresholds; close instrument gaps (gates 1–4)
  2. feat(review): hybrid review flow — verdict ledger spine, cheap-gate phase, tiered findings #2164: realized delta-scoping savings and findings parity
  3. feat(review): hybrid review flow — verdict ledger spine, cheap-gate phase, tiered findings #2164: abort, countable pre-phase and tiering, measured in cost and quality
  4. feat(plugin): nWave comparison follow-ups — subagent lifecycle hooks and two skill gaps #2172: skill-context injection uptake and completion-guard divergence
  5. feat(plugin): nWave comparison follow-ups — subagent lifecycle hooks and two skill gaps #2172: adoption of the source-verification and PBT skills
  6. Decision rollup: keep / tune / retire per mechanism, posted to feat(review): hybrid review flow — verdict ledger spine, cheap-gate phase, tiered findings #2164 and feat(plugin): nWave comparison follow-ups — subagent lifecycle hooks and two skill gaps #2172

Slices 1–4 are independent of each other once slice 0 lands, with one exception: #2202's ledger-off replays run on copies of past diffs, never inside the live sessions that #2203–#2205 measure, so that disabling the ledger can't distort their numbers.

Out of scope


Generated by Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions