Skip to content

feat(hooks): make SubagentStop validate completion, not only record metrics #2174

Description

@bdfinst

Part of #2172.

Context

Our SubagentStop registration runs cost_meter.py and task_completion_metrics.py — both purely observational. Nothing inspects what the subagent actually did before its result is accepted.

nWave's subagent_stop_handler.py (540 lines) returns allow/block: it extracts context from the agent transcript, manages a signal-file lifecycle, emits audit events, and validates step completion before the orchestrator moves on. A second SubagentStop handler writes delivery progress from the transcript rather than asking the model to update a progress file.

The second half is the more interesting idea for us. /build tracks step and slice completion by having the model check boxes in the plan; a hook that derives completion from what the subagent actually returned is the deterministic version of the same fact — and this repo's own rule is to prefer a program over a model for a mechanical question.

Goal

Decide, with evidence, whether SubagentStop should gain a validating handler — and if so, build the narrowest useful one.

Approach

Deliberately staged, because a blocking hook on every subagent return is a large blast radius and the failure mode (a wrongly-blocked legitimate result) is worse than the problem it solves:

  1. Observe first. Add a non-blocking handler that records, per subagent return, whether the claimed outcome is corroborated by the transcript (files actually modified, tests actually run). Emit as a boundary event; block nothing.
  2. Measure. Over real sessions, how often does a subagent's claimed completion diverge from what its transcript shows? If the answer is near zero, close this slice — the prose instructions are working and a hook would be pure risk.
  3. Only then decide whether to block, and on what narrow, mechanically-decidable condition.

Acceptance criteria

  • Stage 1 handler registered, non-blocking, stdlib-only Python (ADR 0014/0015), fail-open on any exception.
  • knowledge/telemetry-schema.md gains a section for the new event — fields, types, emitter, consent gating, consumers — keeping test_schema_doc_covers_all_metrics_paths green.
  • Divergence rate measured across at least 5 real sessions and posted to feat(plugin): nWave comparison follow-ups — subagent lifecycle hooks and two skill gaps #2172, with the go/no-go threshold stated before the numbers are read.
  • Stage 3 does not start without that number. If this slice closes at stage 2 with "no divergence found", that is a successful outcome, not a failed one — record it so the question is not reopened from scratch.
  • Full pre-push pytest directory list green.

Out of scope

  • Blocking behavior, until stage 2 justifies it.
  • Replacing /build's plan checkbox tracking. Corroborating it is in scope; replacing it is not.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions