Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
8b79f9d
chore(metrics): record plan-approval audit entry for #2168+#2169+#217…
claude Sep 22, 2026
27186ca
feat(scripts): add checkpoint_abort.py — cheap-lens blocker decision
claude Sep 22, 2026
3dfe8a5
feat(build): abort remaining checkpoint lenses on a cheap blocker
claude Sep 22, 2026
908666f
fix(build): make checkpoint_abort's merge/outcome calls invocable, fi…
claude Sep 22, 2026
ce29b43
docs(test-review): annotate mechanical vs judgment checks
claude Sep 22, 2026
8c99066
feat(scripts): add test_review_mechanics.py — mechanical pre-phase fo…
claude Sep 22, 2026
f40878f
feat(test-review): wire mechanical pre-phase into protocol, verified …
claude Sep 22, 2026
d0d2190
fix(test-review): wire mechanical pre-phase into /code-review dispatc…
claude Sep 22, 2026
7867121
fix(test-review): fix stale wording, parametrize and harden mechanica…
claude Sep 22, 2026
11eb121
feat(code-review): add render_tiered_findings.py — Tier-1/Tier-2 repo…
claude Sep 22, 2026
b26203e
feat(code-review): render Tier-1 findings by default, Tier-2 via --ex…
claude Sep 22, 2026
236d40d
fix(code-review): print Tier-1 alongside --expand, mirror category fa…
claude Sep 22, 2026
27d0c71
fix(code-review): fix backstop review findings across #2168/#2169/#2170
claude Sep 22, 2026
04d0953
Merge remote-tracking branch 'origin/main' into feat/2164-abort-count…
claude Sep 22, 2026
1878555
fix(agents): avoid bare skills/code-review/SKILL.md path reference
claude Sep 22, 2026
86b5904
feat(agent-eval): wire test-review Phase 0 into unit-tier dispatch
claude Sep 22, 2026
a2605ac
fix(agents): unbreak test-review.md's anchor-citation gate
claude Sep 22, 2026
86996b6
Merge remote-tracking branch 'origin/feat/2164-abort-countable-tiered…
claude Sep 22, 2026
afdea33
Merge branch 'main' into claude/nifty-albattani-82ujmh
bdfinst Sep 23, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
62 changes: 52 additions & 10 deletions plugins/dev-team/skills/agent-eval/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,8 @@ allowed-tools: >-
Bash(readlink *, ls *, date *, mkdir *, command -v claude, claude -p *,
python3 scripts/eval_cache.py *, python3 scripts/run_integration_eval.py *,
python3 scripts/eval_variance.py *, python3 "$CLAUDE_PLUGIN_ROOT/scripts/eval_ablation.py" *,
python3 scripts/citation_lint.py *),
python3 scripts/citation_lint.py *,
python3 "$CLAUDE_PLUGIN_ROOT/scripts/test_review_mechanics.py" *),
Skill(review-agent *), Skill(test-design-advisor *)
---

Expand All @@ -36,7 +37,11 @@ against eval fixtures and grade the results.
keyword checks). Do not apply judgment.
3. **Minimize context per agent.** Pass only the fixture file to the
agent — not the expected results, not other fixtures, not prior
transcripts.
transcripts. One narrow exception: a `test-review` dispatch also
receives its fixture's own `test_review_mechanics.py` mechanical
pre-phase result (Step 3 below, #2169 follow-up) — computed from the
fixture file itself, not from grading data, so it does not weaken this
constraint's intent.
4. **Track results.** Save transcripts for saturation detection. Do
not modify fixtures or expected files.
5. **Be concise.** Output the report table and failure details. No
Expand Down Expand Up @@ -216,9 +221,13 @@ deliberately not relocated under `.claude/`) the record was appended to.

`scripts/eval_cache.py` SHAs each `fixture::target` over the target's
definition + the **transitive closure** of files it reaches (knowledge/,
skills/) + the fixture + expected JSON + grader version. An unchanged SHA
with a stored PASS replays at zero token cost; any changed input busts the
cache.
skills/, and any bare `scripts/<n>.py` it names — e.g. `test-review.md`'s
pointer to `test_review_mechanics.py`, #2169) + the fixture + expected JSON
+ grader version. An unchanged SHA with a stored PASS replays at zero token
cost; any changed input busts the cache — including an edit to
`test_review_mechanics.py` itself, so a Phase 0 detection-logic change
correctly busts cached `test-review` pairs instead of silently replaying a
result computed under the old script.

Default behaviour (cache-on):

Expand Down Expand Up @@ -346,16 +355,49 @@ above) and consult the cache (see *Cache* above):

For each fixture/agent pair (agent fixtures):

1. Dispatch the named review agent against the fixture file/directory:
1. **Test-review Phase 0 pre-phase (#2169 follow-up).** When `<agent-name>`
is `test-review`, first compute the mechanical pre-phase result for the
fixture file — the same invocation `/code-review` step 2b uses:

```bash
python3 "$CLAUDE_PLUGIN_ROOT/scripts/test_review_mechanics.py" . <fixture-path>
```

Without this, a fresh `test-review` dispatch has no Phase 0 result
supplied and falls straight through `agents/test-review.md`'s own
Protocol rule "No result supplied for a file — run Phase 1/2 for it as
usual; say nothing about Phase 0" — so `/agent-eval` would only ever
grade Phase 1/2's judgment, never Phase 0's own detection accuracy
(`mechanicalFail` gating). Skip this sub-step entirely for every other
agent.
2. Dispatch the named review agent against the fixture file/directory:
- **Default:** a fresh subprocess that reads the agent from disk —
`claude -p "/review-agent <agent-name> <fixture-path>" --output-format json`
(add `--model` per the agent's tier when known). Pass **only** the
fixture path, never the expected JSON.
fixture path, never the expected JSON — **except** for `test-review`,
where sub-step 1's result is appended to the prompt text, framed
exactly as `skills/code-review/SKILL.md` step 2b's "Test-review
mechanical pre-phase" block frames it for a production dispatch
("detected by static analysis, do not re-derive" — the agent still
reports it as this file's own finding when `mechanicalFail` is true):

```text
claude -p "/review-agent test-review <fixture-path>

Test-review mechanical pre-phase result for this file (computed by
scripts/test_review_mechanics.py — detected by static analysis, do
not re-derive, cite verbatim):
<sub-step-1 JSON>" --output-format json
```

- **`--in-session`:** invoke `/review-agent <agent-name>` with the
fixture file/directory as the target.
2. Parse the agent's JSON output to extract: `status`, `issues[]`,
fixture file/directory as the target — for `test-review`, pass
sub-step 1's result as additional context in the same invocation
(`Skill(review-agent test-review, <fixture-path> plus the same
Phase 0 framing and JSON above)`).
3. Parse the agent's JSON output to extract: `status`, `issues[]`,
`summary`
3. If running multiple trials (`--trials`), repeat and collect all
4. If running multiple trials (`--trials`), repeat and collect all
results

For each fixture/skill pair (skill fixtures, e.g. `tlg-*`):
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
"""Content guard for agent-eval/SKILL.md's test-review Phase 0 pre-phase
wiring (#2169 follow-up, noted in plans/2164-abort-countable-tiered.md's
Risks section as "Correction to the above").

`/agent-eval`'s unit-tier dispatch passes only the fixture file to a review
agent (constraint 3) -- but `test-review`'s Phase 0 mechanical pre-phase
(`agents/test-review.md` -> Protocol) needs its `test_review_mechanics.py`
result supplied as caller context, since the agent has no Bash tool and
never runs the script itself. Without a dispatch site inside `/agent-eval`
itself, an eval run can never exercise Phase 0's `mechanicalFail` gating --
only Phase 1/2's judgment, unaffected by Phase 0. This test pins that the
dispatch site stays present, so a future SKILL.md edit can't silently drop
it the same way the first `/code-review` step 2b wiring attempt did (see
`test_code_review_test_review_mechanics_pre_pass_marker.py`).
"""

from __future__ import annotations

from _repo_root import REPO_ROOT

SKILL = REPO_ROOT / "plugins" / "dev-team" / "skills" / "agent-eval" / "SKILL.md"

_STEP3_MARKER = "### 3. Run agents against fixtures"
_NEXT_STEP_MARKER = "### 4. Grade each result"


def _step3_section() -> str:
text = SKILL.read_text(encoding="utf-8")
assert _STEP3_MARKER in text, "Step 3 heading not found"
return text.split(_STEP3_MARKER, 1)[1].split(_NEXT_STEP_MARKER, 1)[0]


def test_step3_invokes_test_review_mechanics_script() -> None:
section = _step3_section()
assert "test_review_mechanics.py" in section, (
"agent-eval SKILL.md no longer names test_review_mechanics.py -- "
"test-review's Phase 0 would have no dispatch site during /agent-eval"
)


def test_step3_gates_the_pre_phase_on_test_review_only() -> None:
section = _step3_section()
pre_phase = section.split("test_review_mechanics.py", 1)[0][-400:]
assert "test-review" in pre_phase, (
"the pre-phase block should be scoped to the test-review agent, "
"not run unconditionally for every fixture/agent pair"
)


def test_step3_default_dispatch_appends_phase0_result_for_test_review() -> None:
section = _step3_section()
assert "detected by static analysis, do not re-derive" in section
assert "cite verbatim" in section


def test_constraint_3_documents_the_test_review_exception() -> None:
text = SKILL.read_text(encoding="utf-8")
constraints = text.split("## Orchestrator constraints", 1)[1].split("## Parse Arguments", 1)[0]
assert "test_review_mechanics.py" in constraints
assert "test-review" in constraints


def test_allowed_tools_permits_test_review_mechanics_invocation() -> None:
text = SKILL.read_text(encoding="utf-8")
frontmatter = text.split("---", 2)[1]
assert "test_review_mechanics.py" in frontmatter, (
"allowed-tools frontmatter must permit invoking test_review_mechanics.py "
"or the Step 3 pre-phase call is blocked at the tool-permission layer"
)
28 changes: 24 additions & 4 deletions scripts/eval_cache.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,10 @@
2. the target's root definition (agent ``agents/<t>.md`` or skill
``skills/<t>/SKILL.md``),
3. the **transitive closure** of dependency files reachable from it
(``knowledge/*.md`` it reads, ``skills/*`` it invokes — recursively),
(``knowledge/*.md`` it reads, ``skills/*`` it invokes — recursively,
plus any ``scripts/<name>.py`` it names directly, e.g. a mechanical
pre-phase script an agent cites but never runs itself — issue #2169's
`/agent-eval` follow-up),
4. the fixture file(s) for the stem,
5. the expected JSON,
6. the grader version (hash of the eval_graders package sources).
Expand Down Expand Up @@ -75,6 +78,12 @@
_KNOWLEDGE_RE = re.compile(r"knowledge/([A-Za-z0-9_-]+)\.md")
_SKILL_PATH_RE = re.compile(r"skills/([A-Za-z0-9_-]+)/")
_SKILL_TOOL_RE = re.compile(r"Skill\(([A-Za-z0-9_-]+)")
# A bare `scripts/<name>.py` reference, e.g. test-review.md's Phase 0 pointer
# to `scripts/test_review_mechanics.py` (#2169) — the mechanical pre-phase
# result an agent cites is itself computed by that script, so a change to the
# script's detection logic must bust the fingerprint exactly like a change to
# the agent's own prose would.
_SCRIPT_RE = re.compile(r"scripts/([A-Za-z0-9_]+)\.py")


class Dirs:
Expand All @@ -88,6 +97,7 @@ def __init__(self, repo_root: Path, expected_dir: Path | None = None,
self.fixtures = fixtures_dir or repo_root / "evals" / "fixtures"
self.plugin = plugin_root or repo_root / "plugins" / "dev-team"
self.agents = self.plugin / "agents"
self.scripts = self.plugin / "scripts"
self.skills = self.plugin / "skills"
self.knowledge = self.plugin / "knowledge"
self.graders = Path(__file__).resolve().parent / "eval_graders"
Expand Down Expand Up @@ -139,9 +149,15 @@ def resolve_root_file(pair: str, dirs: Dirs):
def transitive_deps(root_file: Path, dirs: Dirs) -> list[Path]:
"""Sorted, de-duplicated closure of dep files reachable from root_file.

Follows ``knowledge/<n>.md`` reads, ``skills/<n>/`` paths and ``Skill(<n>``
invocations recursively. Only existing files are returned; the root file
itself is excluded (the caller hashes it separately).
Follows ``knowledge/<n>.md`` reads, ``skills/<n>/`` paths, ``Skill(<n>``
invocations, and bare ``scripts/<n>.py`` references recursively. Only
existing files are returned; the root file itself is excluded (the
caller hashes it separately).

Script matches are added as leaves only (never pushed back onto the
walk stack) — they're Python source, not prose that itself names further
``knowledge/``/``skills/`` dependencies, so there is nothing to recurse
into.
"""
seen: set[Path] = set()
stack = [root_file]
Expand Down Expand Up @@ -169,6 +185,10 @@ def transitive_deps(root_file: Path, dirs: Dirs) -> list[Path]:
if sf.exists() and sf != root_file and sf not in seen:
seen.add(sf)
stack.append(sf)
for m in _SCRIPT_RE.finditer(text):
pf = dirs.scripts / f"{m.group(1)}.py"
if pf.exists() and pf not in seen:
seen.add(pf)
return sorted(seen)


Expand Down
31 changes: 31 additions & 0 deletions tests/repo/test_eval_cache.py
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,37 @@ def test_fingerprint_editing_transitive_knowledge_dep_busts_sha(corpus: Path) ->
assert a != b


def _add_transitive_script_dep(corpus: Path, body: str = "pass\n") -> Path:
scripts_dir = corpus / "plugins" / "dev-team" / "scripts"
scripts_dir.mkdir(parents=True, exist_ok=True)
script = scripts_dir / "demo_mechanics.py"
script.write_text(body)
with (corpus / "plugins" / "dev-team" / "agents" / "demo-review.md").open("a") as f:
f.write("Run `scripts/demo_mechanics.py` before Phase 1.\n")
return script


def test_fingerprint_editing_transitive_script_dep_busts_sha(corpus: Path) -> None:
"""A bare `scripts/<n>.py` reference (e.g. test-review.md's pointer to
test_review_mechanics.py, #2169) is a fingerprint contributor: a change
to the script's detection logic must bust cached results for any
agent/skill that cites its output, exactly like an edit to the agent's
own prose would."""
script = _add_transitive_script_dep(corpus)
a = _fingerprint(corpus)
script.write_text("def analyze():\n return {'changed': True}\n")
b = _fingerprint(corpus)
assert a != b


def test_fingerprint_lists_transitive_script_contributor(corpus: Path) -> None:
_add_transitive_script_dep(corpus)
res = _run(corpus, "--fingerprint", "demo::demo-review")
out = res.stdout + res.stderr
assert res.returncode == 0, out
assert "scripts/demo_mechanics.py" in out


def test_fingerprint_lists_transitive_contributors(corpus: Path) -> None:
res = _run(corpus, "--fingerprint", "demo::demo-review")
out = res.stdout + res.stderr
Expand Down
Loading