feat: agent lifecycle improvements — skill hints, completion guard, verification, PBT - #2191
Conversation
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…batch Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…spatch (#2187) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…#2187) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…roperties (#2190) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…lls index (#2189) Registry-completeness gate (test_registry_sync) and the skills-index freshness gate (test_skills_index_current) both went red once source-verification/SKILL.md landed — regenerate docs/skills.md and add the missing Skills Registry row. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…skills index (#2190) Same drift class as the earlier source-verification fix — the registry-completeness and skills-index freshness gates went red once property-based-testing/SKILL.md landed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
… tail (#2188) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…ast-check reference (#2190) Step 4.2's SKILL.md still reported "not yet implemented" for JS/TS after Step 4.3 landed references/languages/javascript.md — point Step 4's JS/TS branch at it instead of the stale stub, and add the allowed-tools Bash() grant covering detect_and_dispatch.py/hypothesis_scaffold.py/npm/npx that tests/repo/test_skill_bash_grants_cover_invocations.py requires for every fenced shell invocation a SKILL.md instructs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Full pre-PR gate run surfaced three drift classes not caught by any individual step's narrower test run: - doc-review.md's new source-verification reference lacked the anchor or "Whole-file load:" token test_agent_knowledge_anchor.py requires. - subagent_completion_guard.py's module docstring mentions "isSidechain" while recording Step 2.1a's research finding, tripping repo_invariants.py's transcript-parsing-confined-to-session-log check (ADR 0036/#2048) even though the hook itself never reads that field — allowlisted with a stated reason, same shape as cost_meter.py's/ context_ceiling_guard.py's existing entries. - hypothesis (added to requirements-dev.txt for property-based-testing) was missing from dev-setup.sh's dev_deps_satisfied() probe and its test's DIST_TO_MODULE mapping. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
- subagent_skill_context.py: actually call strip_plugin_prefix on the dispatched subagent_type, and wrap main() in a fail-open try/except matching its own docstring contract. Without the prefix strip this hook was a silent no-op on every real "dev-team:<agent>" dispatch. - agent_skill_hints.py: broaden the read-failure guard to also catch UnicodeDecodeError, matching its "every failure degrades to no hint" docstring claim. - hypothesis_scaffold.py: restrict _find_function to module-level functions (tree.body) instead of ast.walk, so a same-named class method's docstring never gets scaffolded as an unimportable module-level property test. - test_subagent_completion_guard.py: replace the monkeypatched read_stdin_json double with a real subprocess invocation of the hook, closing an internal-collaborator-doubling finding and adding coverage for main()'s real stdin-parsing path (malformed JSON, missing transcript_path). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
- subagent_completion_guard.py: document the deliberate reason its local _tail_lines/_last_row stays independent of session_log.records.iter_file_records (streaming skip-and-continue vs this hook's need to distinguish a malformed trailing line as its own "unreadable" outcome) instead of the prior "a second call site doesn't justify an abstraction" framing, which answered the wrong question. Replace the two plans/2172-*.md citations (a gitignored, transient file) with the durable #2188/#2172 reference, and generalize the sample transcript path off one maintainer's own project directory name. - repo_invariants.py: correct the subagent_completion_guard.py allowlist entry's stated reason — it does carry its own small transcript-row reader (just not one reading the four flagged identifiers), and there are two 'isSidechain' occurrences in its docstring, not one. - agent_skill_hints.py: adopt this repo's guarded sys.path.insert idiom instead of a bare unconditional insert. - skills-registry.md: add the two new user-invocable commands (/property-based-testing, /source-verification) to the Command Table — they were registered in agent-registry.md but missing from the reference the plugin CLAUDE.md points at. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Security (both untrusted-input handling gaps in new code): - agent_skill_hints.py: reject any subagent_type that isn't a bare agent-name identifier before joining it into a filesystem path — a crafted "../"-shaped dispatch value could otherwise read an arbitrary .md file outside agents/ and echo its skills: list into the dispatch prompt. - hypothesis_scaffold.py: reject a module filename that isn't a valid Python identifier before rendering a generated test — its basename is interpolated unescaped into the generated file's import statement, so a crafted filename could otherwise inject arbitrary statements into a file pytest later collects and executes. Domain: - detect_and_dispatch.py / SKILL.md / tests: match project-init's actual "JS/TS" stack-detection token instead of an invented "JavaScript/TypeScript" spelling — the mismatch broke the JS/TS lane end-to-end (every real JS/TS project would report "unsupported language"). - doc-review.md: add skills: [source-verification] frontmatter and a ## Skills section so the new subagent_skill_context.py hint hook actually fires for the one agent this batch wired the skill into; clarify it applies the procedure by direct inspection since it has no Bash/Skill/WebFetch grant. - agent-registry.md: drop property-based-testing's unverified "Used by QA Engineer, Software Engineer" claim — neither agent's body references the skill. - subagent_completion_guard.py: emit boundary-events.jsonl decision "record" instead of "warn" — this hook's own docstring already calls its posture record-only/non-verdict, and "warn" collides with the stream's reserved verdict-counted term. - telemetry-schema.md: register SubagentStop as a tool value, the two new matched_rule constants, and subagent_completion_guard.py as an emitter — this vocabulary shipped without updating the stream's schema doc. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
…racefully test_hypothesis_scaffold.py (#2190) runs its scaffolded test files via a real pytest subprocess to prove the generated code actually executes — those files `import hypothesis`. For the "Python ceiling" job: it provisions an isolated uv venv via an explicit --with list in scripts/ci-local.sh's chk_python_ceiling, which did not include hypothesis despite requirements-dev.txt already declaring it. Added, with test_python_ceiling.py's CEILING_INVOCATION_PREFIX pin updated to match. For "Plugin content & hooks": that job installs an explicit package list in .github/workflows/plugin-tests.yml, which this session's credentials cannot push a change to (no `workflow` OAuth scope). Instead, the two execution tests now `pytest.importorskip("hypothesis")` first, so they skip cleanly rather than error when the dependency isn't provisioned in a given CI job's environment — the property is still fully exercised by "Python ceiling", which now reliably has it. Also shortens the PR title, which exceeded commitlint's 100-char header limit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
|
Fixed the two CI failures from the initial push:
I fixed the former (added Follow-up for a human (or a session with Generated by Claude Code |
Summary
Four vertical slices of epic #2172, built as one batch per plan approval:
subagent_skill_context.py: aPreToolUsehook onAgent|Taskdispatch that resolves the dispatched agent's declaredskills:frontmatter (hooks/lib/agent_skill_hints.py) and injects a short reminder note into the dispatch'sadditionalContext, so a subagent is nudged toward the skills its own frontmatter already declares as relevant.subagent_completion_guard.py: aSubagentStophook that classifies a subagent's own transcript tail (clean/empty-final-turn/truncated-final-turn/unreadable), derived from inspecting 70 real subagent transcripts from this session, and records aboundary-events.jsonlobservation for the two explainable non-clean outcomes./source-verification: a new user-invocable skill (plusclaim_extractor.py) that extracts and verifies factual claims in generated content against this repo's own code and external sources, flagging each claim verified/contradicted/unverifiable — wired intodoc-review's External claim verification checklist./property-based-testing: a new user-invocable skill that scaffolds a runnable property-based test for a target function — a round-trip test for an encode/decode pair, or an invariant test for a documented postcondition — using Hypothesis for Python and fast-check for JS/TS.Built via
/build's Code-First Small Batches cadence with inline review checkpoints, then closed out through a full/code-review --internalbackstop (correctness, test, architecture, domain, and security lenses) whose findings are folded into the commit history — including two real security hardening fixes (a path-traversal guard on the skill-hint lookup, and an identifier-validation guard on the property-test scaffold's generated import) and a correctness fix for the JS/TS stack-detection token mismatch.Test Plan
plugins/dev-team/tests tests/repo tests/agents tests/commands tests/docs tests/knowledge tests/stack_aware tests/skills tests/scripts tests/hooks— 10831 passed, 48 skipped.scripts/check_md_references.pyandhooks/lib/build_knowledge_index.pyclean.repo_invariants.pyandinternal_double_detector.pystatic pre-passes clean over every changed file./code-review --internalbackstop: correctness-review, test-review, arch-review, domain-review, security-review all dispatched; every error/warning-severity finding fixed and re-verified; suggestion-tier findings triaged (fixed where trivial, otherwise left as non-blocking).git pushcompleted with the repo's own pre-push CI gate green (10754 passed, 46 skipped).Closes #2187
Closes #2188
Closes #2189
Closes #2190
Part of #2172
🤖 Generated with Claude Code
https://claude.ai/code/session_016uBnw1i52qEik2k9LSHa4n
Generated by Claude Code