Skip to content

chore(metrics): pre-register benefit metrics and close instrument gaps #2201

Description

@bdfinst

Part of #2200. Gates slices 1–4.

Goal

Fix every metric, threshold and data window in writing before any post-merge data is read, and make sure each metric has an instrument that actually records it.

Tasks

  • The table must fix a numeric keep/tune/retire line for at least: the delta-scoping skip rate (chore(metrics): measure realized delta-scoping savings and findings parity (#2164) #2202); the findings-parity rule (any miss is a regression); the abort's deferred-lens yield; the test-review true-positive change; the --expand rate; injection overhead in tokens per dispatch; injected-skill uptake; the completion-guard divergence rate and the false-positive rate that allows blocking; and the contradicted-verdict rate per invocation for source-verification.
  • For every before/after comparison (chore(metrics): measure cheap-blocker abort, countable pre-phase and tiered findings (#2164) #2203, chore(metrics): measure skill-context injection uptake and completion-guard divergence (#2172) #2204), record the other merges, model changes and prompt changes that land in the window as known confounds. Those deltas are reported as directional, not causal.
  • Post a table to chore(metrics): evaluate realized benefits of epics #2164 and #2172 #2200 with one row per mechanism: metric, instrument, baseline, keep / tune / retire thresholds, and the data window. Suggested window: ≥10 real /build or /code-review sessions, or 14 days after the latest merge (feat(code-review): scope backstop and repeat runs to the unreviewed delta #2199), whichever comes later. If fewer than 10 sessions exist after 28 days, close the window anyway and record a null result with the actual N.
  • Audit each instrument against a real session:
    • review-verdicts.jsonl: are rows written for /build sub-step 4/6 dispatches, or only for /code-review? Count the missing-scope-marker boundary events.
    • ledgerSkipped / fullySkippedLenses: is the count persisted anywhere, or only rendered in a report?
    • checkpoint_abort.py: are abort outcomes (aborted, deferredLenses, and whether the deferred lenses found anything) recorded to a stream, or only printed?
    • subagent_skill_context.py: is there any signal that an injected skill was actually loaded? (Skill tool calls in the subagent transcript, read via scripts/lib/session_log.)
    • subagent_completion_guard.py: which boundary-event classifications does it emit?
  • For each gap, add the smallest possible log line (a JSONL row, no behavior change) and document it in knowledge/telemetry-schema.md.

Acceptance criteria


Generated by Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions