Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .claude/metrics/config-changelog.jsonl
Original file line number Diff line number Diff line change
Expand Up @@ -86,3 +86,4 @@
{"timestamp": "2026-09-21T18:51:16Z", "type": "approval", "proposed": "Acceptance-criteria gate for plans/2172-agent-lifecycle-improvements.md", "evidence_shown": "plans/2172-agent-lifecycle-improvements.md", "risks_surfaced": ["Step 1.4 latency check has no quantitative SLA threshold", "Step 1.2 missing idempotent-double-fire test case", "Step 2.1a conditional AC pass/fail ambiguity if no hand-back signal found", "Step 2.1b missing invalid-JSON (not just unreadable) transcript test", "Step 3.1 claim-extraction heuristics under-specified with only 2 of 5 pinned", "Step 3.2 missing malformed-but-parseable WebFetch response case", "Step 3.3 doc-review integration test is structural-only, not behavioral", "Step 4.1 property-derivation edge cases (cross-module pair, ambiguous invariant) untested", "Step 4.2 missing language-specific runtime-failure cases (Hypothesis/fast-check install failure)", "Step 4.3 vendoring-impractical threshold undefined"], "description": "Acceptance-criteria gate auto-passed with 10 flagged (0 blocker, 2 warning, 8 suggestion/minor) criterion findings (non-interactive) - no human gate. Trigger: --yes flag. Findings will be resolved as implementation-time decisions during Step 4, per the plan own assumptions-recording convention."}
{"ts": "2026-09-22T14:33:36Z", "type": "approval", "proposed": "Approve plan for #2166+#2171 verdict-ledger writer batch (part of #2164)", "evidence_shown": "plans/2164-verdict-ledger-writer.md", "risks_surfaced": ["Bash write-shape detection is a heuristic, not a sandbox", "override-audit.jsonl/config-changelog.jsonl remain unguarded (Decision 5 scope-narrowing)", "Step 2.1 marker/Step 2.3 parser format drift risk", "No consumer exists yet for review-verdicts.jsonl until #2167", "plan-review-strategic suggested splitting into two PRs (not adopted)", "Decision 4a: verdict trust boundary relies on the orchestrating session's own declared scope marker, not independently re-verified"], "description": "Auto-approved (non-interactive, --yes via /ship)"}
{"ts": "2026-09-22T14:37:03Z", "type": "approval", "proposed": "Acceptance criteria set for #2166+#2171 verdict-ledger batch (part of #2164)", "evidence_shown": "plans/2164-verdict-ledger-writer.md", "risks_surfaced": [], "description": "spec-compliance-review criteria-verification pass: all 10 criteria PASS, no flags"}
{"timestamp": "2026-09-22T17:18:55Z", "type": "approval", "proposed": "Approve plan for #2168+#2169+#2170 batch (abort-on-cheap-blocker, countable test-review pilot, tiered findings output)", "evidence_shown": "plans/2164-abort-countable-tiered.md", "risks_surfaced": ["Strategic critic questioned bundling 3 independent slices into one PR (revert blast-radius, unsubstantiated Slice-3 cost-cutting framing) - acknowledged in Goal section, not split (mirrors prior epic batching pattern)", "Slice 2's mechanical/judgment classification (Step 2.1) is this plan's own derivation, not a separate authority - may need revision if /agent-eval shows a detection regression", "Slice 3's report-size measurement acceptance criterion needs a real multi-finding review round and is deferred to a post-merge follow-up comment on issue #2164, not a build step"], "description": "Auto-approved (non-interactive) - no usable TTY in this remote session. Trigger: --yes. Plan review ran to convergence first: 4 reviewers dispatched, Acceptance and Design each required 3 rounds to resolve real blockers (finding-id collision, severity-vocabulary mismatch with internal_double_detector.py, an unfalsifiable fixture test); all now approve."}
95 changes: 50 additions & 45 deletions plugins/dev-team/agents/test-review.md

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions plugins/dev-team/docs/eval-maintenance.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ for the operational run procedure see [`eval-running-guide.md`](eval-running-gui
| Fixtures | `evals/fixtures/` | Input code (deliberately good or bad) the agents review. |
| Expectations | `evals/expected/*.json` | The **contract**: what a correct verdict looks like per fixture/agent. |
| Grader | `scripts/eval_grade.py` | Deterministic, model-free: compares recorded actuals to expectations. |
| Regression diff | `scripts/compare_eval_results.py` | Diffs two `--actuals` result files against `evals/expected/*.json`, gating on true/false-positive-proxy count regression. Repo-root placement (ADR 0032 category 2, monorepo-dev-only) next to `eval_grade.py` — not shipped, unlike `eval_ablation.py`, which ships only for its unrelated generic `--find-latest` reader mode. |
| Variance | `scripts/eval_variance.py` | Aggregates K trials → pass@k, flap rate, quarantine. |
| Trend | `.claude/metrics/eval-variance.jsonl` | Append-only stability history (metrics only). |
| Semver contract | `scripts/eval_semver_classify.sh` | The eval corpus IS the version contract (#101). |
Expand Down
2 changes: 1 addition & 1 deletion plugins/dev-team/docs/skills.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ Most skills are **user-invocable** as slash commands — shown as `/name`; run t
| Skill | Options | File | Description |
| --- | --- | --- | --- |
| `/apply-fixes` | <corrections-dir> [--dry] [--skip-tests] [--skip-build] [--skip-lint] | [`apply-fixes/SKILL.md`](https://github.com/bdfinst/agentic-dev-team/blob/main/plugins/dev-team/skills/apply-fixes/SKILL.md) | Apply correction prompts generated by /code-review. Use this whenever the user wants to apply, fix, or action the results of a code review — phrases like "apply the fixes", "fix the issues", "apply corrections", or after /code-review has run and produced a corrections/ directory. |
| `/code-review` | [--agent <name>] [--since <ref>] [--path <dir>] [--all] [--json] [--internal] [--force --reason "<text>"] [--static-analysis\|--no-static-analysis] [--init-risks] [--background] [--pdf] | [`code-review/SKILL.md`](https://github.com/bdfinst/agentic-dev-team/blob/main/plugins/dev-team/skills/code-review/SKILL.md) | Run all enabled review agents against target files. Use this whenever the user asks for a code review, wants feedback on their code, says "review my code", "check this before I PR", "what's wrong with this", "run the agents", or has just finished implementing a feature. Use proactively before commits and pull requests. |
| `/code-review` | [--agent <name>] [--since <ref>] [--path <dir>] [--all] [--json] [--expand <finding-id>\|all] [--internal] [--force --reason "<text>"] [--static-analysis\|--no-static-analysis] [--init-risks] [--background] [--pdf] | [`code-review/SKILL.md`](https://github.com/bdfinst/agentic-dev-team/blob/main/plugins/dev-team/skills/code-review/SKILL.md) | Run all enabled review agents against target files. Use this whenever the user asks for a code review, wants feedback on their code, says "review my code", "check this before I PR", "what's wrong with this", "run the agents", or has just finished implementing a feature. Use proactively before commits and pull requests. |
| `/frontend-architecture` | [--path <dir>] [--since <ref>] [--all] [--json] | [`frontend-architecture/SKILL.md`](https://github.com/bdfinst/agentic-dev-team/blob/main/plugins/dev-team/skills/frontend-architecture/SKILL.md) | Frontend component architecture review — dispatch the component-architecture-review agent over the frontend component files to catch reusable components that should be extracted, duplicated UI patterns, prop drilling, component-granularity problems, and inconsistent component APIs as a frontend evolves. Use when the user says "review the frontend architecture", "are my components reusable", "is this UI duplicated", "should this be a shared component", "check for prop drilling", or before extracting a component library. Advisory — it recommends, it does not edit. |
| `/review` | [--agent <name>] [--since <ref>] [--path <dir>] [--all] [--json] [--internal] [--force --reason "<text>"] [--static-analysis\|--no-static-analysis] [--init-risks] [--background] | [`review/SKILL.md`](https://github.com/bdfinst/agentic-dev-team/blob/main/plugins/dev-team/skills/review/SKILL.md) | Alias for /code-review. Run all enabled review agents against target files. Use this whenever the user asks for a code review, wants feedback on their code, says "review my code", "check this before I PR", "what's wrong with this", "run the agents", or has just finished implementing a feature. |
| `/review-agent` | <agent-name> [--since <ref>] [--path <dir>] [--internal] [--json] | [`review-agent/SKILL.md`](https://github.com/bdfinst/agentic-dev-team/blob/main/plugins/dev-team/skills/review-agent/SKILL.md) | Run a single named review agent against target files. Use this when the user names a specific agent (e.g. "run security-review", "check for test issues", "run js-fp-review on this file") rather than wanting the full suite. Prefer this over /code-review when only one concern is relevant or speed matters. Also used by the orchestrator for inline review checkpoints during Phase 3 implementation. |
Expand Down
Loading
Loading