Skip to content

feat: agent lifecycle improvements — skill hints, completion guard, verification, PBT - #2191

Merged
bdfinst merged 25 commits into
mainfrom
feat/2172-agent-lifecycle-improvements
Sep 22, 2026
Merged

bdfinst merged 25 commits into
mainfrom
feat/2172-agent-lifecycle-improvements

Conversation

@bdfinst

@bdfinst bdfinst commented Sep 21, 2026

Copy link
Copy Markdown
Owner

Summary

Four vertical slices of epic #2172, built as one batch per plan approval:

  • feat(hooks): inject skill-loading context via PreToolUse on Agent/Task dispatch #2187 — subagent_skill_context.py: a PreToolUse hook on Agent|Task dispatch that resolves the dispatched agent's declared skills: frontmatter (hooks/lib/agent_skill_hints.py) and injects a short reminder note into the dispatch's additionalContext, so a subagent is nudged toward the skills its own frontmatter already declares as relevant.
  • feat(hooks): make SubagentStop validate completion, not only record metrics #2188 — subagent_completion_guard.py: a SubagentStop hook that classifies a subagent's own transcript tail (clean / empty-final-turn / truncated-final-turn / unreadable), derived from inspecting 70 real subagent transcripts from this session, and records a boundary-events.jsonl observation for the two explainable non-clean outcomes.
  • feat(skills): add a source-verification skill #2189 — /source-verification: a new user-invocable skill (plus claim_extractor.py) that extracts and verifies factual claims in generated content against this repo's own code and external sources, flagging each claim verified/contradicted/unverifiable — wired into doc-review's External claim verification checklist.
  • feat(skills): add a property-based-testing skill #2190 — /property-based-testing: a new user-invocable skill that scaffolds a runnable property-based test for a target function — a round-trip test for an encode/decode pair, or an invariant test for a documented postcondition — using Hypothesis for Python and fast-check for JS/TS.

Built via /build's Code-First Small Batches cadence with inline review checkpoints, then closed out through a full /code-review --internal backstop (correctness, test, architecture, domain, and security lenses) whose findings are folded into the commit history — including two real security hardening fixes (a path-traversal guard on the skill-hint lookup, and an identifier-validation guard on the property-test scaffold's generated import) and a correctness fix for the JS/TS stack-detection token mismatch.

Test Plan

  • Full pre-PR pytest gate green: plugins/dev-team/tests tests/repo tests/agents tests/commands tests/docs tests/knowledge tests/stack_aware tests/skills tests/scripts tests/hooks — 10831 passed, 48 skipped.
  • scripts/check_md_references.py and hooks/lib/build_knowledge_index.py clean.
  • repo_invariants.py and internal_double_detector.py static pre-passes clean over every changed file.
  • /code-review --internal backstop: correctness-review, test-review, arch-review, domain-review, security-review all dispatched; every error/warning-severity finding fixed and re-verified; suggestion-tier findings triaged (fixed where trivial, otherwise left as non-blocking).
  • git push completed with the repo's own pre-push CI gate green (10754 passed, 46 skipped).

Closes #2187
Closes #2188
Closes #2189
Closes #2190
Part of #2172

🤖 Generated with Claude Code

https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n


Generated by Claude Code

…batch

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…roperties (#2190)

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…lls index (#2189)

Registry-completeness gate (test_registry_sync) and the skills-index
freshness gate (test_skills_index_current) both went red once
source-verification/SKILL.md landed — regenerate docs/skills.md and add
the missing Skills Registry row.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…skills index (#2190)

Same drift class as the earlier source-verification fix — the
registry-completeness and skills-index freshness gates went red once
property-based-testing/SKILL.md landed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…ast-check reference (#2190)

Step 4.2's SKILL.md still reported "not yet implemented" for JS/TS after
Step 4.3 landed references/languages/javascript.md — point Step 4's JS/TS
branch at it instead of the stale stub, and add the allowed-tools Bash()
grant covering detect_and_dispatch.py/hypothesis_scaffold.py/npm/npx that
tests/repo/test_skill_bash_grants_cover_invocations.py requires for every
fenced shell invocation a SKILL.md instructs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Full pre-PR gate run surfaced three drift classes not caught by any
individual step's narrower test run:

- doc-review.md's new source-verification reference lacked the anchor or
  "Whole-file load:" token test_agent_knowledge_anchor.py requires.
- subagent_completion_guard.py's module docstring mentions "isSidechain"
  while recording Step 2.1a's research finding, tripping
  repo_invariants.py's transcript-parsing-confined-to-session-log check
  (ADR 0036/#2048) even though the hook itself never reads that field —
  allowlisted with a stated reason, same shape as cost_meter.py's/
  context_ceiling_guard.py's existing entries.
- hypothesis (added to requirements-dev.txt for property-based-testing)
  was missing from dev-setup.sh's dev_deps_satisfied() probe and its
  test's DIST_TO_MODULE mapping.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
- subagent_skill_context.py: actually call strip_plugin_prefix on the
  dispatched subagent_type, and wrap main() in a fail-open try/except
  matching its own docstring contract. Without the prefix strip this
  hook was a silent no-op on every real "dev-team:<agent>" dispatch.
- agent_skill_hints.py: broaden the read-failure guard to also catch
  UnicodeDecodeError, matching its "every failure degrades to no hint"
  docstring claim.
- hypothesis_scaffold.py: restrict _find_function to module-level
  functions (tree.body) instead of ast.walk, so a same-named class
  method's docstring never gets scaffolded as an unimportable
  module-level property test.
- test_subagent_completion_guard.py: replace the monkeypatched
  read_stdin_json double with a real subprocess invocation of the
  hook, closing an internal-collaborator-doubling finding and adding
  coverage for main()'s real stdin-parsing path (malformed JSON,
  missing transcript_path).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
- subagent_completion_guard.py: document the deliberate reason its
  local _tail_lines/_last_row stays independent of
  session_log.records.iter_file_records (streaming skip-and-continue
  vs this hook's need to distinguish a malformed trailing line as its
  own "unreadable" outcome) instead of the prior "a second call site
  doesn't justify an abstraction" framing, which answered the wrong
  question. Replace the two plans/2172-*.md citations (a gitignored,
  transient file) with the durable #2188/#2172 reference, and
  generalize the sample transcript path off one maintainer's own
  project directory name.
- repo_invariants.py: correct the subagent_completion_guard.py
  allowlist entry's stated reason — it does carry its own small
  transcript-row reader (just not one reading the four flagged
  identifiers), and there are two 'isSidechain' occurrences in its
  docstring, not one.
- agent_skill_hints.py: adopt this repo's guarded sys.path.insert
  idiom instead of a bare unconditional insert.
- skills-registry.md: add the two new user-invocable commands
  (/property-based-testing, /source-verification) to the Command
  Table — they were registered in agent-registry.md but missing from
  the reference the plugin CLAUDE.md points at.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Security (both untrusted-input handling gaps in new code):
- agent_skill_hints.py: reject any subagent_type that isn't a bare
  agent-name identifier before joining it into a filesystem path — a
  crafted "../"-shaped dispatch value could otherwise read an
  arbitrary .md file outside agents/ and echo its skills: list into
  the dispatch prompt.
- hypothesis_scaffold.py: reject a module filename that isn't a valid
  Python identifier before rendering a generated test — its basename
  is interpolated unescaped into the generated file's import
  statement, so a crafted filename could otherwise inject arbitrary
  statements into a file pytest later collects and executes.

Domain:
- detect_and_dispatch.py / SKILL.md / tests: match project-init's
  actual "JS/TS" stack-detection token instead of an invented
  "JavaScript/TypeScript" spelling — the mismatch broke the JS/TS
  lane end-to-end (every real JS/TS project would report
  "unsupported language").
- doc-review.md: add skills: [source-verification] frontmatter and a
  ## Skills section so the new subagent_skill_context.py hint hook
  actually fires for the one agent this batch wired the skill into;
  clarify it applies the procedure by direct inspection since it has
  no Bash/Skill/WebFetch grant.
- agent-registry.md: drop property-based-testing's unverified "Used
  by QA Engineer, Software Engineer" claim — neither agent's body
  references the skill.
- subagent_completion_guard.py: emit boundary-events.jsonl decision
  "record" instead of "warn" — this hook's own docstring already
  calls its posture record-only/non-verdict, and "warn" collides with
  the stream's reserved verdict-counted term.
- telemetry-schema.md: register SubagentStop as a tool value, the two
  new matched_rule constants, and subagent_completion_guard.py as an
  emitter — this vocabulary shipped without updating the stream's
  schema doc.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
@bdfinst bdfinst changed the title feat: agent lifecycle improvements — skill hints, completion guard, source verification, property testing feat: agent lifecycle improvements — skill hints, completion guard, verification, PBT Sep 21, 2026
…racefully

test_hypothesis_scaffold.py (#2190) runs its scaffolded test files via a
real pytest subprocess to prove the generated code actually executes —
those files `import hypothesis`.

For the "Python ceiling" job: it provisions an isolated uv venv via an
explicit --with list in scripts/ci-local.sh's chk_python_ceiling, which
did not include hypothesis despite requirements-dev.txt already
declaring it. Added, with test_python_ceiling.py's
CEILING_INVOCATION_PREFIX pin updated to match.

For "Plugin content & hooks": that job installs an explicit package list
in .github/workflows/plugin-tests.yml, which this session's credentials
cannot push a change to (no `workflow` OAuth scope). Instead, the two
execution tests now `pytest.importorskip("hypothesis")` first, so they
skip cleanly rather than error when the dependency isn't provisioned in
a given CI job's environment — the property is still fully exercised by
"Python ceiling", which now reliably has it.

Also shortens the PR title, which exceeded commitlint's 100-char header
limit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
@bdfinst

bdfinst commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Fixed the two CI failures from the initial push:

  • PR title lint (header > 100 chars) — shortened the title.
  • Python ceiling / Plugin content & hooks — both failed on test_hypothesis_scaffold.py's two execution tests (ModuleNotFoundError: No module named 'hypothesis'). requirements-dev.txt already declares hypothesis, but neither CI job installs it: "Python ceiling" provisions an isolated uv venv via an explicit --with list (scripts/ci-local.sh), and "Plugin content & hooks" installs an explicit package list in .github/workflows/plugin-tests.yml.

I fixed the former (added hypothesis to the --with list, updated the matching test pin in test_python_ceiling.py). For the latter, this session's credentials don't have the GitHub workflow OAuth scope, so pushing a change to .github/workflows/plugin-tests.yml was rejected server-side. Rather than leave the job red, I made the two subprocess-execution tests pytest.importorskip("hypothesis") first, so they skip cleanly instead of erroring when that job's environment doesn't have it — the property they check is still fully exercised by "Python ceiling", which now reliably installs it.

Follow-up for a human (or a session with workflow scope): add hypothesis to the pip install step in .github/workflows/plugin-tests.yml's "Plugin content & hooks" job so those two tests run for real there too, then the importorskip guards can stay as defense-in-depth or be removed.


Generated by Claude Code

@bdfinst
bdfinst enabled auto-merge (squash) September 22, 2026 12:04
@bdfinst
bdfinst merged commit 4e81194 into main Sep 22, 2026
15 checks passed
@bdfinst
bdfinst deleted the feat/2172-agent-lifecycle-improvements branch September 22, 2026 12:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants