You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
What was measured before building: #2165's cross-checkpoint duplication ceiling (20.69% byte-estimate at the Step 6 backstop; one epic, theoretical, not observed) and #2170's report-size cut (59.0%, one round). Nothing has been measured after. The superseded #2174 required a completion-divergence rate over ≥5 sessions before deciding on blocking; #2188 dropped that, so it was never run.
North Star: a change that can't name the friction it removes doesn't ship. This epic checks whether each mechanism removes friction in practice, and retires or tunes what doesn't.
Goal
For every shipped mechanism: one measured number from real sessions, compared against a threshold written down before the data is read, ending in a keep / tune / retire decision.
Rules
Data source: one local machine. All session data is gathered from the maintainer's local machine, where .claude/metrics/ and ~/.claude/metrics/ persist. Cloud sessions are out of scope because their containers, and the metrics in them, are recycled.
Real session: a /build or /code-review run on actual repo work on that machine. Eval fixtures, dry runs and replays are not real sessions; a slice that uses one must label it as such.
Pre-register metrics, thresholds and the data window in slice 0 before reading any post-merge data.
Measure, don't build. If an instrument is missing, the only in-scope code is the smallest logging addition that makes the number exist.
Name each instrument (script, JSONL stream, or eval) next to every number.
A null result is a valid outcome. Record it so the question doesn't get reopened from scratch.
Slices
Pre-register metrics and thresholds; close instrument gaps (gates 1–4)
Slices 1–4 are independent of each other once slice 0 lands, with one exception: #2202's ledger-off replays run on copies of past diffs, never inside the live sessions that #2203–#2205 measure, so that disabling the ledger can't distort their numbers.
Out of scope
New features or behavior changes other than the minimum instrumentation from slice 0.
Context
Both nWave-comparison epics have shipped:
test-reviewpre-phase feat(agents): countable pre-phase in test-review, as a pilot #2169, tiered findings feat(code-review): tiered findings output with on-demand expansion #2170 (PR feat(review): abort-on-cheap-blocker, countable test-review, tiered findings #2197).hooks/subagent_skill_context.py(feat(hooks): inject skill-loading context via PreToolUse on Agent/Task dispatch #2187),hooks/subagent_completion_guard.py(feat(hooks): make SubagentStop validate completion, not only record metrics #2188),source-verification(feat(skills): add a source-verification skill #2189),property-based-testing(feat(skills): add a property-based-testing skill #2190), all via PR feat: agent lifecycle improvements — skill hints, completion guard, verification, PBT #2191.What was measured before building: #2165's cross-checkpoint duplication ceiling (20.69% byte-estimate at the Step 6 backstop; one epic, theoretical, not observed) and #2170's report-size cut (59.0%, one round). Nothing has been measured after. The superseded #2174 required a completion-divergence rate over ≥5 sessions before deciding on blocking; #2188 dropped that, so it was never run.
North Star: a change that can't name the friction it removes doesn't ship. This epic checks whether each mechanism removes friction in practice, and retires or tunes what doesn't.
Goal
For every shipped mechanism: one measured number from real sessions, compared against a threshold written down before the data is read, ending in a keep / tune / retire decision.
Rules
.claude/metrics/and~/.claude/metrics/persist. Cloud sessions are out of scope because their containers, and the metrics in them, are recycled./buildor/code-reviewrun on actual repo work on that machine. Eval fixtures, dry runs and replays are not real sessions; a slice that uses one must label it as such.Slices
Slices 1–4 are independent of each other once slice 0 lands, with one exception: #2202's ledger-off replays run on copies of past diffs, never inside the live sessions that #2203–#2205 measure, so that disabling the ledger can't distort their numbers.
Out of scope
Generated by Claude Code