Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ SkipHow can now enable, check, and disable its own default governance when an in

This is a minor release because the setup playbook and packaged helper are new opt-in capability and the recovery rule widens what the integration playbook permits within existing authority. Existing grants, restrictions, and protected-action boundaries survive the upgrade. The public skill name and record formats are unchanged.

Deterministic checks, the September 6 audit disposition, and grader results on retained end states are recorded in [docs/evidence.md](docs/evidence.md). Every 4.2.0 behavior is `UNVERIFIED`: whether a session loads the skill from an `AGENTS.override.md` block, whether the agent operates the setup playbook as written, and whether the recovery rule changes what a run does. Host receipts for the exact 4.2.0 package are recorded in `evals/host-smoke.json` when they land.
Deterministic checks, the September 6 audit disposition, and grader results on retained end states are recorded in [docs/evidence.md](docs/evidence.md). Claude Code 2.1.261 installed and uninstalled the exact package cleanly. Codex CLI 0.153.0 installed it from the approved Git source into an isolated home, byte-identical to the committed package; three bounded sessions there show the skill enabling itself into a non-empty `AGENTS.override.md` after one confirmation, an ordinary-language request loading the kernel from that block and delivering four correct repairs to a synthetic remote with foreign work preserved, and the skill disabling itself. Each is one observation with recorded deviations, not a reliability rate. Whether the recovery rule changes what a run does, whether Claude Code loads the skill from persistent configuration, and every coordinated-delivery claim remain `UNVERIFIED`.

## 4.1.1 (2026-09-05)

Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,7 +97,7 @@ Explicit invocation remains the fallback and diagnostic path:
$skiphow The totals overlap on small screens. Find the cause and fix it.
```

What each host has actually shown is in the [dated support summary](docs/evidence.md#support-summary-as-of-2026-09-06). In short: ordinary-language loading, delivery to a synthetic remote, and native resume were observed once each on Codex with the exact 4.1.0 package in an isolated home; on Claude Code no persistent-setup run exists and one earlier bare-prompt pilot did not select the skill. No activation mode has a measured reliability, and the package ships no session hook.
What each host has actually shown is in the [dated support summary](docs/evidence.md#support-summary-as-of-2026-09-06). In short: on Codex, the current package enabled and disabled itself through the skill once each, and one ordinary-language request loaded the kernel from the written block and delivered four correct repairs to a synthetic remote; 4.1.0 showed loading, delivery, and native resume once each before that. On Claude Code no persistent-setup run exists and one earlier bare-prompt pilot did not select the skill. No activation mode has a measured reliability, and the package ships no session hook.

## Use it

Expand Down Expand Up @@ -146,7 +146,7 @@ SkipHow keeps one owner-facing entry. Critical rules stay in its kernel, while f

## What the evidence shows

Deterministic checks prove package structure; controlled runs are required for behavior claims. The behavioral observations on record were made on 2.x packages, on both hosts, and cover fully specified requests, open product choices, failure diagnosis, adversarial verification, and the splitting of larger work into independently verifiable units. On the 4.x virtual-CTO contract, retained isolated Codex diagnostics on the exact 4.1.0 package show ordinary-language loading, four correct repairs delivered to a synthetic remote with foreign work preserved, read-only behavior on analysis and unrelated requests, and native resume and compaction. A separate Claude coordination diagnostic left its synthetic remote unchanged and accepted an incorrect shipping calculation despite independent review. Coordinated cross-host delivery, failed-delegate recovery, real GitHub tracking, and every behavior of the current package remain `UNVERIFIED` until retained receipts show them.
Deterministic checks prove package structure; controlled runs are required for behavior claims. The behavioral observations on record were made on 2.x packages, on both hosts, and cover fully specified requests, open product choices, failure diagnosis, adversarial verification, and the splitting of larger work into independently verifiable units. On the 4.x virtual-CTO contract, retained isolated Codex diagnostics on the exact 4.1.0 package show ordinary-language loading, four correct repairs delivered to a synthetic remote with foreign work preserved, read-only behavior on analysis and unrelated requests, and native resume and compaction. A separate Claude coordination diagnostic left its synthetic remote unchanged and accepted an incorrect shipping calculation despite independent review. On the current package, three isolated Codex sessions show the agent-operated enable and disable paths and one ordinary-language delivery loaded through the override file, each once. Coordinated cross-host delivery, failed-delegate recovery, and real GitHub tracking remain `UNVERIFIED` until retained receipts show them.

These are observations, not a reliability rate. The project does not retain every transcript, public adoption is still limited, and comparative advantage over a base agent or another framework is `UNVERIFIED`. The [dated support summary](docs/evidence.md#support-summary-as-of-2026-09-06) says what was demonstrated on which package, host, and configuration; the rest of the [evidence ledger](docs/evidence.md) is the single home for the method, the claims each run supports, and the failures.

Expand Down
10 changes: 5 additions & 5 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ services keep their own security policies.

## Package validation, 2026-09-06

Version 4.2.0 is validated per capability in [`evals/host-smoke.json`](evals/host-smoke.json); a row without a fresh 4.2.0 receipt stays `UNVERIFIED` there. The [dated support summary](docs/evidence.md#support-summary-as-of-2026-09-06) states what each host has shown for this package. The previous 4.1.1 receipts remain in `evals/receipts/host-validation-411-20260905/`, and the isolated Codex diagnostics remain 4.1.0 observations.
Version 4.2.0 is validated per capability in [`evals/host-smoke.json`](evals/host-smoke.json): Claude Code 2.1.261 clean install and uninstall, and Codex CLI 0.153.0 clean install from the approved Git source, uninstall, persistent setup, explicit fallback, and playbook load each carry a receipt; a row without a fresh 4.2.0 receipt stays `UNVERIFIED` there. The [dated support summary](docs/evidence.md#support-summary-as-of-2026-09-06) states what each host has shown for this package. The previous 4.1.1 receipts remain in `evals/receipts/host-validation-411-20260905/`; the September 5 isolated Codex diagnostics remain 4.1.0 observations, and the [September 6 diagnostics](evals/receipts/isolated-host-420-20260906/README.md) are separate 4.2.0 observations.

The historical 4.1.0 candidate passed both host schema validators. Claude Code 2.1.261
installed all fifteen regular files byte for byte and uninstalled them in a
Expand Down Expand Up @@ -74,18 +74,18 @@ page under `learn.chatgpt.com`; the redirect target is the page actually read.
| Per-agent read-only controls | Subagent frontmatter takes a `tools` allowlist, `disallowedTools`, and `permissionMode`, whose values include `plan` for read-only exploration. | [Subagents](https://code.claude.com/docs/en/sub-agents) | 2026-09-04 | none | `UNVERIFIED` (documented) |
| Worktree isolation | `isolation: worktree` runs a subagent in a temporary git worktree. | [Subagents](https://code.claude.com/docs/en/sub-agents) | 2026-09-04 | none | `UNVERIFIED` (documented) |
| Plugin validation | Manifest `.claude-plugin/plugin.json`; `claude plugin validate <path>` validates it and `--strict` treats warnings as errors. | [Plugins](https://code.claude.com/docs/en/plugins) | 2026-09-04 | 2.1.259 | `PASS` (`scripts/check_hosts.py`, 2026-09-04) |
| Clean installation | `claude plugin marketplace add`, `claude plugin install --scope user`, `claude plugin uninstall --scope user`; `CLAUDE_CONFIG_DIR` points the host at a scratch home. | [Discover plugins](https://code.claude.com/docs/en/discover-plugins), [Skills](https://code.claude.com/docs/en/skills) | 2026-09-04 | 2.1.260 | `PASS` (`scripts/check_hosts.py --smoke`: clean home, install, 15 regular files matching exact 4.0.1 payload `3d6f359a鈥, uninstall verified, 2026-09-04) |
| Clean installation | `claude plugin marketplace add`, `claude plugin install --scope user`, `claude plugin uninstall --scope user`; `CLAUDE_CONFIG_DIR` points the host at a scratch home. | [Discover plugins](https://code.claude.com/docs/en/discover-plugins), [Skills](https://code.claude.com/docs/en/skills) | 2026-09-06 | 2.1.261 | `PASS` (`scripts/check_hosts.py --smoke`: clean home, install, 17 regular files matching exact 4.2.0 payload `5bcd09d1鈥, uninstall verified; [ledger](evals/host-smoke.json)) |

### Codex CLI

| Capability | What the source says | Source | Verified | Tested version | Status |
| --- | --- | --- | --- | --- | --- |
| Skill loading | Skills are discovered from `.agents/skills` in the current, parent, and repository-root directories, the user-level `.agents/skills` directory in the home directory, `/etc/codex/skills`, and system skills. Progressive disclosure lists name and description within 2 per cent of the context window, or 8,000 characters where that is unknown; the full file loads on selection. Explicit `$skill` invocation and implicit invocation are both available; `allow_implicit_invocation` in `agents/openai.yaml` defaults to true. | [Skills](https://developers.openai.com/codex/skills) | 2026-09-04 | none | `UNVERIFIED` (documented; no activation run on record) |
| Persistent instruction loading | Codex reads `AGENTS.override.md` in its home when that file exists and is not empty, and `AGENTS.md` otherwise, then layers project files with the same precedence per directory; `CODEX_HOME` relocates the home; empty files are skipped and the combined size is capped by `project_doc_max_bytes` (32 KiB by default). The packaged helper targets the file this rule makes effective and moves a block left in the shadowed file. | [AGENTS.md](https://learn.chatgpt.com/docs/agent-configuration/agents-md) | 2026-09-06 | 0.153.0 | `Observed` once on exact 4.1.0: the kernel loaded before edits in the isolated bootstrap diagnostic with the block in `AGENTS.md` and no override present; 4.2.0 and the override path `UNVERIFIED` until a receipt |
| Skill loading | Skills are discovered from `.agents/skills` in the current, parent, and repository-root directories, the user-level `.agents/skills` directory in the home directory, `/etc/codex/skills`, and system skills. Progressive disclosure lists name and description within 2 per cent of the context window, or 8,000 characters where that is unknown; the full file loads on selection. Explicit `$skill` invocation and implicit invocation are both available; `allow_implicit_invocation` in `agents/openai.yaml` defaults to true. | [Skills](https://developers.openai.com/codex/skills) | 2026-09-06 | 0.153.0 | `PASS` for explicit `$skiphow` invocation of exact 4.2.0 in the isolated enable and disable sessions, where the first read used a stale path before the installed copy was located; implicit selection without the activation block remains `UNVERIFIED` |
| Persistent instruction loading | Codex reads `AGENTS.override.md` in its home when that file exists and is not empty, and `AGENTS.md` otherwise, then layers project files with the same precedence per directory; `CODEX_HOME` relocates the home; empty files are skipped and the combined size is capped by `project_doc_max_bytes` (32 KiB by default). The packaged helper targets the file this rule makes effective and moves a block left in the shadowed file. | [AGENTS.md](https://learn.chatgpt.com/docs/agent-configuration/agents-md) | 2026-09-06 | 0.153.0 | `Observed` twice: on exact 4.1.0 with the block in `AGENTS.md` and no override present, and on exact 4.2.0 with the block the agent itself wrote into a non-empty `AGENTS.override.md`, the kernel loading before edits each time ([4.2.0 receipts](evals/receipts/isolated-host-420-20260906/README.md)) |
| Per-agent read-only controls | Custom agents are TOML files in the Codex home `agents/` directory or the project `.codex/agents/` and may set `sandbox_mode` per agent; the page names marking one agent read-only as the example. Absent an override, subagents inherit the parent's sandbox policy and permission mode. | [Subagents](https://developers.openai.com/codex/subagents) | 2026-09-04 | none | `UNVERIFIED` (documented; corrects the earlier claim that no declarable per-delegate profile exists) |
| Worktree isolation | The subagents page documents no worktree or separate-checkout option for a subagent. | [Subagents](https://developers.openai.com/codex/subagents) | 2026-09-04 | none | `UNVERIFIED` (not documented either way) |
| Plugin validation | Manifest `.codex-plugin/plugin.json`. There is no `codex plugin validate` subcommand; validation runs the `validate_plugin.py` script shipped with the plugin-creator system skill in the Codex repository, which CI checks out at a pinned commit. | [openai/codex plugin-creator scripts](https://github.com/openai/codex/tree/333beecd41281b1350688b417a2f20c66e2a743e/codex-rs/skills/src/assets/samples/plugin-creator/scripts) | 2026-09-04 | none locally | `UNVERIFIED` locally (validator not on this machine); required to `PASS` in CI |
| Clean installation | `codex plugin marketplace add`, `codex plugin add`, `codex plugin list --json`, `codex plugin remove` exist in `codex plugin --help`; `CODEX_HOME` relocates the host home. The plugins page documents the plugin browser and uninstall but none of these commands. | [Plugins](https://developers.openai.com/codex/plugins), `codex plugin --help` 0.153.0 | 2026-09-04 | 0.153.0 | `UNVERIFIED` (the local run was refused by a managed `/etc/codex/requirements.toml` source policy before install; nothing was installed) |
| Clean installation | `codex plugin marketplace add`, `codex plugin add`, `codex plugin list --json`, `codex plugin remove` exist in `codex plugin --help`; `CODEX_HOME` relocates the host home. The plugins page documents the plugin browser and uninstall but none of these commands. | [Plugins](https://developers.openai.com/codex/plugins), `codex plugin --help` 0.153.0 | 2026-09-06 | 0.153.0 | `PASS` for exact 4.2.0 from the approved Git source in an isolated home: 17 regular files byte-identical to the committed package, then removed ([ledger](evals/host-smoke.json)); the release runner's local marketplace is still refused by the managed `/etc/codex/requirements.toml` source policy |

### Codex surfaces

Expand Down
Loading