diff --git a/CHANGELOG.md b/CHANGELOG.md index f5bdda063a..80d6f3df9d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,42 @@ > completion state and remaining P0 gates. No version bump or release claim is > made here while that status holds. +## [1.63.0.0] - 2026-07-23 + +## **"Go register an API key" is no longer your homework.** +## **gstack asks once, drives the browser, and proves the key works.** + +Every gstack skill now recognizes the moment a plan or workflow needs something done on a third-party website: registering an API key, creating a vendor account, configuring a dashboard or OAuth app. Instead of dumping a manual step list, the skill checks for an agentic browser that can work across your logged-in accounts (Aside today), names the exact site and exact actions, and asks one question: drive it now, give me the steps, or defer. Say yes and the agent does the browsing while the parts that should stay human stay human. Passwords, payment, identity checks, and CAPTCHAs are handed to you, never faked. Captured keys go straight to an owner-only file or your secret store, never into chat output or logs. And before the skill claims success, it proves the credential works with one read-only API call. + +That last rule exists because dashboards lie. The token field on a real vendor dashboard held a masked placeholder, the real value only existed behind the Copy button, and the verify call caught it with a 401 before it could land in a .env file as garbage. + +### The numbers that matter + +Source: `bun run scripts/gstack2/run-parity.ts` on this release, plus two live vendor flows (Duffel, OpenSky Network) run end to end through the contract before shipping. + +| Metric | Before | After | +|---|---|---| +| Dispatchers that offer agentic-browser help for third-party web actions | 0 | 6 of 6 | +| Consent asked per task, never persisted or inferred | n/a | always | +| Pinned parity checks | 5,050 | 5,074 | + +The +24 checks mean nobody can quietly strip the consent gate or the secret-handling rules; the suite fails if the contract loses either. + +### What this means for builders + +The gap between "the plan says you need an OpenSky key" and "the key is in your .env and verified" used to be your browser tab, your copy-paste, and your typo. Now it is one question. Answer it and keep building. + +### Itemized changes + +### Added + +- `references/THIRD-PARTY-ACTIONS.md` generated into all six skill trees: agentic-browser detection (Aside via the existing browser-provider registry), a mandatory per-task consent question naming the exact site and actions, user-performed credential steps, no-echo secret capture, and a non-mutating verification call before any success claim. +- Dispatch protocol step directing every dispatcher to read the contract before telling the user to act on a third-party website. + +### For contributors + +- New `thirdPartyActionsContract()` in `scripts/gstack2/generate-skill-tree.ts`; 4 pinned checks per tree in `run-parity.ts` (5,050 to 5,074). + ## [1.62.0.0] - 2026-07-23 ## **Open a large repo and gstack offers to index it.** diff --git a/VERSION b/VERSION index 1042f0adfa..232d60188a 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -1.62.0.0 +1.63.0.0 diff --git a/docs/gstack-2/JUDGMENT-PARITY.md b/docs/gstack-2/JUDGMENT-PARITY.md index e5dbd3ea71..796ac91e03 100644 --- a/docs/gstack-2/JUDGMENT-PARITY.md +++ b/docs/gstack-2/JUDGMENT-PARITY.md @@ -2,7 +2,7 @@ Parity is executable, not a prose claim. Run `bun run scripts/gstack2/run-parity.ts` or the dedicated Bun tests. -The pinned release inventory passes **5,050 checks** across 55 specialist sources, 16 carved sections, 25 routing scenarios, 28 regression ports, and **78 assets**. +The pinned release inventory passes **5,074 checks** across 55 specialist sources, 16 carved sections, 25 routing scenarios, 28 regression ports, and **78 assets**. The suite verifies: diff --git a/evals/parity/transcripts/policy-units.json b/evals/parity/transcripts/policy-units.json index 64adc50e8d..2189621842 100644 --- a/evals/parity/transcripts/policy-units.json +++ b/evals/parity/transcripts/policy-units.json @@ -43,7 +43,7 @@ "prompt_sha256": "a0fcff9cad68f9305da861fa5a58990e1ef51d9d84298fe2df7ba8a2f56b1dba", "semantic_attempt_sha256": "7724c3d954f421753c597364c383bd76bc01ba9803f99fb14488287bb925aeb6" }, - "policy_sha256": "995b434e47a5355b2256544da09ccb796e299ad94dfa59ec33644e3cb76fb319", + "policy_sha256": "80221cf6e756c013db7b29979a96b5214cf9f7d1ee8e0aba35b923983caf3a52", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" @@ -88,7 +88,7 @@ "prompt_sha256": "04d82a21358f188cb823dcceeecaebea0b4ed8b9c1f8d00220ef3b499c48b1c8", "semantic_attempt_sha256": "9639406995955284517acb30f56b37dedc42a6c08142163dfa8fe0295105cbbf" }, - "policy_sha256": "c10eecd9af5dbb8e9fee6b6c03e7919ec990de8b5c73d71a80d79f076dd0702c", + "policy_sha256": "7c65e6b6f507abaf3481b2f5ce58d9ed7ec3eef2d9081349755f8e66acf969b9", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" @@ -138,7 +138,7 @@ "prompt_sha256": "8ad639f708a2ccd5c7c438b1dedf6afacdfa4eac3883a7a046ea23f29314e1c4", "semantic_attempt_sha256": "4c450039e4c622fbeaa34b7f7dfc810d3b7369aaf7e5f4d240cd6cc18a31331a" }, - "policy_sha256": "1d02974d45509033fcbaaab721afb9588f48efc6cc4e1ffcb3bc9c0d7d0440ba", + "policy_sha256": "a06a65d69f4f2df210bf925353868739351d298e06961e607ff65c17718e71eb", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" @@ -188,7 +188,7 @@ "prompt_sha256": "d2aa9aee06a116bf04cb62772e8abdf8dbd606811b6d38565389c88e86c79ede", "semantic_attempt_sha256": "64885a66bdc6688b880cbb9a59babf314834c068eb040657501bf287bfcb0d2b" }, - "policy_sha256": "995b434e47a5355b2256544da09ccb796e299ad94dfa59ec33644e3cb76fb319", + "policy_sha256": "80221cf6e756c013db7b29979a96b5214cf9f7d1ee8e0aba35b923983caf3a52", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" @@ -232,7 +232,7 @@ "prompt_sha256": "377e3a2dfc7b04272f857d27e77c9aa3c9f36c3feb63c03e5aee3ecd3dea8e2a", "semantic_attempt_sha256": "a49b1528091886749278dc22e7da4556c090800989b2bca9c0e1c09dd880fed7" }, - "policy_sha256": "da07b4bb14cb2ea3f3d2e9570a6dd7f268a549680a44926ad9dd069121d74d86", + "policy_sha256": "b1e6f4a3f0a96c4f85e5308ce1ff4b1f0da8c48143b2fb23f927e95593827de4", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" @@ -279,7 +279,7 @@ "prompt_sha256": "a01bb977de84e53d8ce3dfa427bcc93d73c6e449cd4e9c1ff63b436fd41fb0d1", "semantic_attempt_sha256": "ce5c64c62c3ff9b60949f34717dffbc908f62ce30f8a4a48ec182eba4363a206" }, - "policy_sha256": "50c01aa40fae8aed2182fc77ac94f2834c45e9fba63e81f765f66fa7bf1343cb", + "policy_sha256": "f44bdd611a10f589aa33457b209a5937c07b76e3d0c5c947dff4aeb8d25e179e", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" @@ -330,7 +330,7 @@ "prompt_sha256": "95e97e26268ad7e509527e0f54c943ed6c4105d919f250648a2ad59ce77cbdb9", "semantic_attempt_sha256": "dacd11a78e32aeb6c0065496bedd4a9770f7bbac1036e003d3768fac41eb84f0" }, - "policy_sha256": "995b434e47a5355b2256544da09ccb796e299ad94dfa59ec33644e3cb76fb319", + "policy_sha256": "80221cf6e756c013db7b29979a96b5214cf9f7d1ee8e0aba35b923983caf3a52", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" @@ -381,7 +381,7 @@ "prompt_sha256": "bdd7e8adfa7c15cf8531f84c3adaaacc725075f1a78f7487225300750547b82d", "semantic_attempt_sha256": "32c11b63412c5d873ff29dcc87c45bef1b50daa218a5c6c1075864aa24d7c3b6" }, - "policy_sha256": "995b434e47a5355b2256544da09ccb796e299ad94dfa59ec33644e3cb76fb319", + "policy_sha256": "80221cf6e756c013db7b29979a96b5214cf9f7d1ee8e0aba35b923983caf3a52", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" @@ -426,7 +426,7 @@ "prompt_sha256": "51b7d53bc342e8632e34ae31a167cbde568ab5ee2a6dd2ccc980583a30604164", "semantic_attempt_sha256": "06f77b27c9bd360562655b23ca00a0dd021bdb14ebaaf3d836845e4e0cb43d68" }, - "policy_sha256": "231071d3b07b462be19ed583f7e74a872a6f621db226de8cf6d734b4fd67c823", + "policy_sha256": "9e4f2bb725d02e9f612e17ada4a73ddaf2ea12fb7889ceb6de224b276fad2606", "policy_present": true, "prompt_is_not_authority_input": true, "verdict": "PASS" diff --git a/package.json b/package.json index fba8165572..d2a2550ed6 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "gstack", - "version": "1.62.0.0", + "version": "1.63.0.0", "description": "GStack 2 — six portable Agent Skills with an optional host-neutral runtime.", "license": "MIT", "type": "module", diff --git a/scripts/gstack2/generate-skill-tree.ts b/scripts/gstack2/generate-skill-tree.ts index 464503c220..2f26fe1248 100644 --- a/scripts/gstack2/generate-skill-tree.ts +++ b/scripts/gstack2/generate-skill-tree.ts @@ -257,7 +257,7 @@ Web context: 1. Infer the mode from product stage, surface, requested artifact, mutation authorization, evidence needs, and deployment state. Do not route by keyword alone. 2. Refine the public mode to the smallest applicable internal specialist set, then print the required execution header before any substantive output. 3. Read each active module in full from the path shown in the mode/alias tables. Its specialist body, behavioral contract, STOP gates, and appended upstream judgment ports are binding. Read a lazy specialist phase in full only when the workflow reaches its package-local reference. -4. Read \`references/EXECUTION-PROFILES.md\`, \`references/SHARED-JUDGMENT.md\`, and \`references/AUTHORITY-POLICY.md\` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read \`references/RUNTIME.md\` before capability-dependent work and \`references/WEB-CONTEXT.md\` before public-web work. When the target is a repository, read \`references/CODE-INTELLIGENCE.md\` once before substantive specialist work and follow its one-time indexing offer. +4. Read \`references/EXECUTION-PROFILES.md\`, \`references/SHARED-JUDGMENT.md\`, and \`references/AUTHORITY-POLICY.md\` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read \`references/RUNTIME.md\` before capability-dependent work and \`references/WEB-CONTEXT.md\` before public-web work. When the target is a repository, read \`references/CODE-INTELLIGENCE.md\` once before substantive specialist work and follow its one-time indexing offer. Read \`references/THIRD-PARTY-ACTIONS.md\` before directing the user to act on a third-party website (API key registration, vendor accounts, dashboards). 5. If an old asset path is unavailable, use \`references/ASSETS.md\`. If legacy prose invokes another retired skill, resolve it through \`references/COMPATIBILITY.md\` and stay inside these six dispatchers. 6. Preserve report-only versus mutation boundaries. Missing mutation authorization fails closed: do not edit merely because a specialist can fix. Commits, pushes, PRs, merges, deploys, messages, and other external mutations still require affirmative authority from the user. 7. Match the user's language. Keep code identifiers, commands, and source quotations original when translation would reduce accuracy. @@ -513,6 +513,25 @@ function webContextContract(): string { ].join('\n'); } +function thirdPartyActionsContract(): string { + return [GENERATED, + '# Third-party web actions', + '', + 'A specialist step sometimes requires action on an external website the user controls: registering an API key, creating a vendor or developer account, configuring a dashboard, webhook, OAuth app, billing plan, or domain verification. This contract governs that moment. It grants no browsing authority; `SHARED-JUDGMENT.md` and `AUTHORITY-POLICY.md` remain binding, including approval before spending.', + '', + '1. Never hand the user a manual step list for a third-party site without first checking for an agentic browser that can act across their logged-in accounts. Today that is the Aside AI browser (its section in `references/BROWSER-PROVIDERS.md` owns detection and readiness); confirm the `aside` CLI is callable, for example `aside --version`. Never install it, and never treat binary presence as consent to browse.', + '', + '2. If such a browser is available, STOP and ask one explicit question before any browsing: name the exact site and the exact actions (for example "create a test-mode API token in the Duffel dashboard"), then offer A) drive it now through the agentic browser, B) manual instructions, C) defer. The selection is per-task consent; never persist it as standing permission and never infer it from an earlier task.', + '', + '3. When driving, touch only the named site and actions. Password entry, new-account credential choice, payment, and identity verification remain user-performed steps: hand off and wait instead of acting. Prefer credential flows that never expose the secret to the agent, such as password-manager autofill or host-native copy.', + '', + '4. A captured secret (API key, token, webhook signing secret) never appears in chat output, logs, or shell history. Write it to a user-approved local file with owner-only permissions or the user\'s secret store, and keep generated-file destinations out of version control. Dashboard fields are often masked placeholders; verify the captured credential with one non-mutating API call before claiming success.', + '', + '5. If no agentic browser is available, or the user declines or defers, provide the manual steps and mark the step blocked on the user. Do not recommend or install new products to close the gap.', + '', + ].join('\n'); +} + function codeIntelligenceContract(): string { return [GENERATED, '# Optional code-intelligence indexing', @@ -621,6 +640,7 @@ function writeSharedContracts(): void { write(path.join(ROOT, 'skills', tree, 'references', 'AUTHORITY-POLICY.md'), authorityPolicyContract()); write(path.join(ROOT, 'skills', tree, 'references', 'WEB-CONTEXT.md'), webContextContract()); write(path.join(ROOT, 'skills', tree, 'references', 'CODE-INTELLIGENCE.md'), codeIntelligenceContract()); + write(path.join(ROOT, 'skills', tree, 'references', 'THIRD-PARTY-ACTIONS.md'), thirdPartyActionsContract()); write(path.join(ROOT, 'skills', tree, 'references', 'RUNTIME.md'), runtimeContract()); write(path.join(ROOT, 'skills', tree, 'references', 'BROWSER-PROVIDERS.md'), `${GENERATED}\n${renderBrowserProviderContract()}`); write(path.join(ROOT, 'skills', tree, 'references', 'support', 'runtime-bootstrap.mjs'), bootstrap); diff --git a/scripts/gstack2/run-parity.ts b/scripts/gstack2/run-parity.ts index 5505d00f4c..c5a23a63af 100644 --- a/scripts/gstack2/run-parity.ts +++ b/scripts/gstack2/run-parity.ts @@ -26,8 +26,10 @@ const ALLOWED_DISPOSITIONS = new Set(['VERBATIM_PORT', 'MECHANICAL_PORT', 'JUDGM // overlay for issue #879 (28 -> 29, targets '*', all 55 modules) added 113 // more (2 per module + 3 regression checks). The founder-resources opt-out // overlay for issue #538 (29 -> 30, office-hours only) added 5 more. The -// per-tree code-intelligence offer contract added 18 (3 per dispatcher). -export const EXPECTED_PARITY_CHECKS = 5050; +// per-tree code-intelligence offer contract added 18 (3 per dispatcher). The +// third-party web-action contract added 24 more (4 per tree: exists, marker +// format, dispatcher load, consent/secret content). +export const EXPECTED_PARITY_CHECKS = 5074; function sha256(value: string | Uint8Array): string { return createHash('sha256').update(value).digest('hex'); @@ -182,6 +184,18 @@ export function runParity(): ParityResult { `${tree} code-intelligence offer lost its silent-degrade, decline, or no-auto-install behavior`, ); } + check(dispatcher.includes('references/THIRD-PARTY-ACTIONS.md'), `${tree} dispatcher does not load the third-party web-action contract`); + const thirdPartyPath = path.join(ROOT, 'skills', tree, 'references', 'THIRD-PARTY-ACTIONS.md'); + if (fs.existsSync(thirdPartyPath)) { + const thirdParty = fs.readFileSync(thirdPartyPath, 'utf8'); + check( + thirdParty.includes('STOP and ask one explicit question before any browsing') + && thirdParty.includes('per-task consent') + && thirdParty.includes('never appears in chat output, logs, or shell history') + && thirdParty.includes('verify the captured credential with one non-mutating API call'), + `${tree} third-party web-action contract lost its consent gate or secret-handling rules`, + ); + } } const effectsPath = path.join(ROOT, 'skills', 'ship', 'references', 'EXTERNAL-EFFECTS.md'); check(fs.existsSync(effectsPath), 'Ship lacks the durable external-effect protocol'); @@ -382,7 +396,7 @@ export function runParity(): ParityResult { for (const tree of TREE_NAMES) { const compatibility = fs.readFileSync(path.join(ROOT, 'skills', tree, 'references', 'COMPATIBILITY.md'), 'utf8'); check(!compatibility.includes('../../../') && !compatibility.includes('compat/README.md'), `${tree} compatibility map escapes the selected package`); - for (const reference of ['SHARED-JUDGMENT.md', 'WEB-CONTEXT.md', 'RUNTIME.md']) { + for (const reference of ['SHARED-JUDGMENT.md', 'WEB-CONTEXT.md', 'RUNTIME.md', 'THIRD-PARTY-ACTIONS.md']) { const referencePath = path.join(ROOT, 'skills', tree, 'references', reference); check(fs.existsSync(referencePath), `${tree} lacks ${reference}`); if (fs.existsSync(referencePath)) { diff --git a/skills/debug/SKILL.md b/skills/debug/SKILL.md index b23fb6baae..56ae946825 100644 --- a/skills/debug/SKILL.md +++ b/skills/debug/SKILL.md @@ -27,7 +27,7 @@ Web context: 1. Infer the mode from product stage, surface, requested artifact, mutation authorization, evidence needs, and deployment state. Do not route by keyword alone. 2. Refine the public mode to the smallest applicable internal specialist set, then print the required execution header before any substantive output. 3. Read each active module in full from the path shown in the mode/alias tables. Its specialist body, behavioral contract, STOP gates, and appended upstream judgment ports are binding. Read a lazy specialist phase in full only when the workflow reaches its package-local reference. -4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. +4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. Read `references/THIRD-PARTY-ACTIONS.md` before directing the user to act on a third-party website (API key registration, vendor accounts, dashboards). 5. If an old asset path is unavailable, use `references/ASSETS.md`. If legacy prose invokes another retired skill, resolve it through `references/COMPATIBILITY.md` and stay inside these six dispatchers. 6. Preserve report-only versus mutation boundaries. Missing mutation authorization fails closed: do not edit merely because a specialist can fix. Commits, pushes, PRs, merges, deploys, messages, and other external mutations still require affirmative authority from the user. 7. Match the user's language. Keep code identifiers, commands, and source quotations original when translation would reduce accuracy. diff --git a/skills/debug/references/THIRD-PARTY-ACTIONS.md b/skills/debug/references/THIRD-PARTY-ACTIONS.md new file mode 100644 index 0000000000..f08e82aad7 --- /dev/null +++ b/skills/debug/references/THIRD-PARTY-ACTIONS.md @@ -0,0 +1,14 @@ + +# Third-party web actions + +A specialist step sometimes requires action on an external website the user controls: registering an API key, creating a vendor or developer account, configuring a dashboard, webhook, OAuth app, billing plan, or domain verification. This contract governs that moment. It grants no browsing authority; `SHARED-JUDGMENT.md` and `AUTHORITY-POLICY.md` remain binding, including approval before spending. + +1. Never hand the user a manual step list for a third-party site without first checking for an agentic browser that can act across their logged-in accounts. Today that is the Aside AI browser (its section in `references/BROWSER-PROVIDERS.md` owns detection and readiness); confirm the `aside` CLI is callable, for example `aside --version`. Never install it, and never treat binary presence as consent to browse. + +2. If such a browser is available, STOP and ask one explicit question before any browsing: name the exact site and the exact actions (for example "create a test-mode API token in the Duffel dashboard"), then offer A) drive it now through the agentic browser, B) manual instructions, C) defer. The selection is per-task consent; never persist it as standing permission and never infer it from an earlier task. + +3. When driving, touch only the named site and actions. Password entry, new-account credential choice, payment, and identity verification remain user-performed steps: hand off and wait instead of acting. Prefer credential flows that never expose the secret to the agent, such as password-manager autofill or host-native copy. + +4. A captured secret (API key, token, webhook signing secret) never appears in chat output, logs, or shell history. Write it to a user-approved local file with owner-only permissions or the user's secret store, and keep generated-file destinations out of version control. Dashboard fields are often masked placeholders; verify the captured credential with one non-mutating API call before claiming success. + +5. If no agentic browser is available, or the user declines or defers, provide the manual steps and mark the step blocked on the user. Do not recommend or install new products to close the gap. diff --git a/skills/design/SKILL.md b/skills/design/SKILL.md index 94a4a53a72..9508382856 100644 --- a/skills/design/SKILL.md +++ b/skills/design/SKILL.md @@ -27,7 +27,7 @@ Web context: 1. Infer the mode from product stage, surface, requested artifact, mutation authorization, evidence needs, and deployment state. Do not route by keyword alone. 2. Refine the public mode to the smallest applicable internal specialist set, then print the required execution header before any substantive output. 3. Read each active module in full from the path shown in the mode/alias tables. Its specialist body, behavioral contract, STOP gates, and appended upstream judgment ports are binding. Read a lazy specialist phase in full only when the workflow reaches its package-local reference. -4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. +4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. Read `references/THIRD-PARTY-ACTIONS.md` before directing the user to act on a third-party website (API key registration, vendor accounts, dashboards). 5. If an old asset path is unavailable, use `references/ASSETS.md`. If legacy prose invokes another retired skill, resolve it through `references/COMPATIBILITY.md` and stay inside these six dispatchers. 6. Preserve report-only versus mutation boundaries. Missing mutation authorization fails closed: do not edit merely because a specialist can fix. Commits, pushes, PRs, merges, deploys, messages, and other external mutations still require affirmative authority from the user. 7. Match the user's language. Keep code identifiers, commands, and source quotations original when translation would reduce accuracy. diff --git a/skills/design/references/THIRD-PARTY-ACTIONS.md b/skills/design/references/THIRD-PARTY-ACTIONS.md new file mode 100644 index 0000000000..f08e82aad7 --- /dev/null +++ b/skills/design/references/THIRD-PARTY-ACTIONS.md @@ -0,0 +1,14 @@ + +# Third-party web actions + +A specialist step sometimes requires action on an external website the user controls: registering an API key, creating a vendor or developer account, configuring a dashboard, webhook, OAuth app, billing plan, or domain verification. This contract governs that moment. It grants no browsing authority; `SHARED-JUDGMENT.md` and `AUTHORITY-POLICY.md` remain binding, including approval before spending. + +1. Never hand the user a manual step list for a third-party site without first checking for an agentic browser that can act across their logged-in accounts. Today that is the Aside AI browser (its section in `references/BROWSER-PROVIDERS.md` owns detection and readiness); confirm the `aside` CLI is callable, for example `aside --version`. Never install it, and never treat binary presence as consent to browse. + +2. If such a browser is available, STOP and ask one explicit question before any browsing: name the exact site and the exact actions (for example "create a test-mode API token in the Duffel dashboard"), then offer A) drive it now through the agentic browser, B) manual instructions, C) defer. The selection is per-task consent; never persist it as standing permission and never infer it from an earlier task. + +3. When driving, touch only the named site and actions. Password entry, new-account credential choice, payment, and identity verification remain user-performed steps: hand off and wait instead of acting. Prefer credential flows that never expose the secret to the agent, such as password-manager autofill or host-native copy. + +4. A captured secret (API key, token, webhook signing secret) never appears in chat output, logs, or shell history. Write it to a user-approved local file with owner-only permissions or the user's secret store, and keep generated-file destinations out of version control. Dashboard fields are often masked placeholders; verify the captured credential with one non-mutating API call before claiming success. + +5. If no agentic browser is available, or the user declines or defers, provide the manual steps and mark the step blocked on the user. Do not recommend or install new products to close the gap. diff --git a/skills/plan/SKILL.md b/skills/plan/SKILL.md index cc794cc1bc..94d3a5af73 100644 --- a/skills/plan/SKILL.md +++ b/skills/plan/SKILL.md @@ -28,7 +28,7 @@ Web context: 1. Infer the mode from product stage, surface, requested artifact, mutation authorization, evidence needs, and deployment state. Do not route by keyword alone. 2. Refine the public mode to the smallest applicable internal specialist set, then print the required execution header before any substantive output. 3. Read each active module in full from the path shown in the mode/alias tables. Its specialist body, behavioral contract, STOP gates, and appended upstream judgment ports are binding. Read a lazy specialist phase in full only when the workflow reaches its package-local reference. -4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. +4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. Read `references/THIRD-PARTY-ACTIONS.md` before directing the user to act on a third-party website (API key registration, vendor accounts, dashboards). 5. If an old asset path is unavailable, use `references/ASSETS.md`. If legacy prose invokes another retired skill, resolve it through `references/COMPATIBILITY.md` and stay inside these six dispatchers. 6. Preserve report-only versus mutation boundaries. Missing mutation authorization fails closed: do not edit merely because a specialist can fix. Commits, pushes, PRs, merges, deploys, messages, and other external mutations still require affirmative authority from the user. 7. Match the user's language. Keep code identifiers, commands, and source quotations original when translation would reduce accuracy. diff --git a/skills/plan/references/THIRD-PARTY-ACTIONS.md b/skills/plan/references/THIRD-PARTY-ACTIONS.md new file mode 100644 index 0000000000..f08e82aad7 --- /dev/null +++ b/skills/plan/references/THIRD-PARTY-ACTIONS.md @@ -0,0 +1,14 @@ + +# Third-party web actions + +A specialist step sometimes requires action on an external website the user controls: registering an API key, creating a vendor or developer account, configuring a dashboard, webhook, OAuth app, billing plan, or domain verification. This contract governs that moment. It grants no browsing authority; `SHARED-JUDGMENT.md` and `AUTHORITY-POLICY.md` remain binding, including approval before spending. + +1. Never hand the user a manual step list for a third-party site without first checking for an agentic browser that can act across their logged-in accounts. Today that is the Aside AI browser (its section in `references/BROWSER-PROVIDERS.md` owns detection and readiness); confirm the `aside` CLI is callable, for example `aside --version`. Never install it, and never treat binary presence as consent to browse. + +2. If such a browser is available, STOP and ask one explicit question before any browsing: name the exact site and the exact actions (for example "create a test-mode API token in the Duffel dashboard"), then offer A) drive it now through the agentic browser, B) manual instructions, C) defer. The selection is per-task consent; never persist it as standing permission and never infer it from an earlier task. + +3. When driving, touch only the named site and actions. Password entry, new-account credential choice, payment, and identity verification remain user-performed steps: hand off and wait instead of acting. Prefer credential flows that never expose the secret to the agent, such as password-manager autofill or host-native copy. + +4. A captured secret (API key, token, webhook signing secret) never appears in chat output, logs, or shell history. Write it to a user-approved local file with owner-only permissions or the user's secret store, and keep generated-file destinations out of version control. Dashboard fields are often masked placeholders; verify the captured credential with one non-mutating API call before claiming success. + +5. If no agentic browser is available, or the user declines or defers, provide the manual steps and mark the step blocked on the user. Do not recommend or install new products to close the gap. diff --git a/skills/qa/SKILL.md b/skills/qa/SKILL.md index eb823eebe7..a2068eca77 100644 --- a/skills/qa/SKILL.md +++ b/skills/qa/SKILL.md @@ -27,7 +27,7 @@ Web context: 1. Infer the mode from product stage, surface, requested artifact, mutation authorization, evidence needs, and deployment state. Do not route by keyword alone. 2. Refine the public mode to the smallest applicable internal specialist set, then print the required execution header before any substantive output. 3. Read each active module in full from the path shown in the mode/alias tables. Its specialist body, behavioral contract, STOP gates, and appended upstream judgment ports are binding. Read a lazy specialist phase in full only when the workflow reaches its package-local reference. -4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. +4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. Read `references/THIRD-PARTY-ACTIONS.md` before directing the user to act on a third-party website (API key registration, vendor accounts, dashboards). 5. If an old asset path is unavailable, use `references/ASSETS.md`. If legacy prose invokes another retired skill, resolve it through `references/COMPATIBILITY.md` and stay inside these six dispatchers. 6. Preserve report-only versus mutation boundaries. Missing mutation authorization fails closed: do not edit merely because a specialist can fix. Commits, pushes, PRs, merges, deploys, messages, and other external mutations still require affirmative authority from the user. 7. Match the user's language. Keep code identifiers, commands, and source quotations original when translation would reduce accuracy. diff --git a/skills/qa/references/THIRD-PARTY-ACTIONS.md b/skills/qa/references/THIRD-PARTY-ACTIONS.md new file mode 100644 index 0000000000..f08e82aad7 --- /dev/null +++ b/skills/qa/references/THIRD-PARTY-ACTIONS.md @@ -0,0 +1,14 @@ + +# Third-party web actions + +A specialist step sometimes requires action on an external website the user controls: registering an API key, creating a vendor or developer account, configuring a dashboard, webhook, OAuth app, billing plan, or domain verification. This contract governs that moment. It grants no browsing authority; `SHARED-JUDGMENT.md` and `AUTHORITY-POLICY.md` remain binding, including approval before spending. + +1. Never hand the user a manual step list for a third-party site without first checking for an agentic browser that can act across their logged-in accounts. Today that is the Aside AI browser (its section in `references/BROWSER-PROVIDERS.md` owns detection and readiness); confirm the `aside` CLI is callable, for example `aside --version`. Never install it, and never treat binary presence as consent to browse. + +2. If such a browser is available, STOP and ask one explicit question before any browsing: name the exact site and the exact actions (for example "create a test-mode API token in the Duffel dashboard"), then offer A) drive it now through the agentic browser, B) manual instructions, C) defer. The selection is per-task consent; never persist it as standing permission and never infer it from an earlier task. + +3. When driving, touch only the named site and actions. Password entry, new-account credential choice, payment, and identity verification remain user-performed steps: hand off and wait instead of acting. Prefer credential flows that never expose the secret to the agent, such as password-manager autofill or host-native copy. + +4. A captured secret (API key, token, webhook signing secret) never appears in chat output, logs, or shell history. Write it to a user-approved local file with owner-only permissions or the user's secret store, and keep generated-file destinations out of version control. Dashboard fields are often masked placeholders; verify the captured credential with one non-mutating API call before claiming success. + +5. If no agentic browser is available, or the user declines or defers, provide the manual steps and mark the step blocked on the user. Do not recommend or install new products to close the gap. diff --git a/skills/review/SKILL.md b/skills/review/SKILL.md index 7fde491c35..4f6047ad4c 100644 --- a/skills/review/SKILL.md +++ b/skills/review/SKILL.md @@ -27,7 +27,7 @@ Web context: 1. Infer the mode from product stage, surface, requested artifact, mutation authorization, evidence needs, and deployment state. Do not route by keyword alone. 2. Refine the public mode to the smallest applicable internal specialist set, then print the required execution header before any substantive output. 3. Read each active module in full from the path shown in the mode/alias tables. Its specialist body, behavioral contract, STOP gates, and appended upstream judgment ports are binding. Read a lazy specialist phase in full only when the workflow reaches its package-local reference. -4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. +4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. Read `references/THIRD-PARTY-ACTIONS.md` before directing the user to act on a third-party website (API key registration, vendor accounts, dashboards). 5. If an old asset path is unavailable, use `references/ASSETS.md`. If legacy prose invokes another retired skill, resolve it through `references/COMPATIBILITY.md` and stay inside these six dispatchers. 6. Preserve report-only versus mutation boundaries. Missing mutation authorization fails closed: do not edit merely because a specialist can fix. Commits, pushes, PRs, merges, deploys, messages, and other external mutations still require affirmative authority from the user. 7. Match the user's language. Keep code identifiers, commands, and source quotations original when translation would reduce accuracy. diff --git a/skills/review/references/THIRD-PARTY-ACTIONS.md b/skills/review/references/THIRD-PARTY-ACTIONS.md new file mode 100644 index 0000000000..f08e82aad7 --- /dev/null +++ b/skills/review/references/THIRD-PARTY-ACTIONS.md @@ -0,0 +1,14 @@ + +# Third-party web actions + +A specialist step sometimes requires action on an external website the user controls: registering an API key, creating a vendor or developer account, configuring a dashboard, webhook, OAuth app, billing plan, or domain verification. This contract governs that moment. It grants no browsing authority; `SHARED-JUDGMENT.md` and `AUTHORITY-POLICY.md` remain binding, including approval before spending. + +1. Never hand the user a manual step list for a third-party site without first checking for an agentic browser that can act across their logged-in accounts. Today that is the Aside AI browser (its section in `references/BROWSER-PROVIDERS.md` owns detection and readiness); confirm the `aside` CLI is callable, for example `aside --version`. Never install it, and never treat binary presence as consent to browse. + +2. If such a browser is available, STOP and ask one explicit question before any browsing: name the exact site and the exact actions (for example "create a test-mode API token in the Duffel dashboard"), then offer A) drive it now through the agentic browser, B) manual instructions, C) defer. The selection is per-task consent; never persist it as standing permission and never infer it from an earlier task. + +3. When driving, touch only the named site and actions. Password entry, new-account credential choice, payment, and identity verification remain user-performed steps: hand off and wait instead of acting. Prefer credential flows that never expose the secret to the agent, such as password-manager autofill or host-native copy. + +4. A captured secret (API key, token, webhook signing secret) never appears in chat output, logs, or shell history. Write it to a user-approved local file with owner-only permissions or the user's secret store, and keep generated-file destinations out of version control. Dashboard fields are often masked placeholders; verify the captured credential with one non-mutating API call before claiming success. + +5. If no agentic browser is available, or the user declines or defers, provide the manual steps and mark the step blocked on the user. Do not recommend or install new products to close the gap. diff --git a/skills/ship/SKILL.md b/skills/ship/SKILL.md index c7a6b2fe09..f2ea03a6cb 100644 --- a/skills/ship/SKILL.md +++ b/skills/ship/SKILL.md @@ -27,7 +27,7 @@ Web context: 1. Infer the mode from product stage, surface, requested artifact, mutation authorization, evidence needs, and deployment state. Do not route by keyword alone. 2. Refine the public mode to the smallest applicable internal specialist set, then print the required execution header before any substantive output. 3. Read each active module in full from the path shown in the mode/alias tables. Its specialist body, behavioral contract, STOP gates, and appended upstream judgment ports are binding. Read a lazy specialist phase in full only when the workflow reaches its package-local reference. -4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. +4. Read `references/EXECUTION-PROFILES.md`, `references/SHARED-JUDGMENT.md`, and `references/AUTHORITY-POLICY.md` for every invocation. Infer Depth from structured operating conditions, then obey its mandatory modules, legal skips, artifacts, and claim limits. Read `references/RUNTIME.md` before capability-dependent work and `references/WEB-CONTEXT.md` before public-web work. When the target is a repository, read `references/CODE-INTELLIGENCE.md` once before substantive specialist work and follow its one-time indexing offer. Read `references/THIRD-PARTY-ACTIONS.md` before directing the user to act on a third-party website (API key registration, vendor accounts, dashboards). 5. If an old asset path is unavailable, use `references/ASSETS.md`. If legacy prose invokes another retired skill, resolve it through `references/COMPATIBILITY.md` and stay inside these six dispatchers. 6. Preserve report-only versus mutation boundaries. Missing mutation authorization fails closed: do not edit merely because a specialist can fix. Commits, pushes, PRs, merges, deploys, messages, and other external mutations still require affirmative authority from the user. 7. Match the user's language. Keep code identifiers, commands, and source quotations original when translation would reduce accuracy. diff --git a/skills/ship/references/THIRD-PARTY-ACTIONS.md b/skills/ship/references/THIRD-PARTY-ACTIONS.md new file mode 100644 index 0000000000..f08e82aad7 --- /dev/null +++ b/skills/ship/references/THIRD-PARTY-ACTIONS.md @@ -0,0 +1,14 @@ + +# Third-party web actions + +A specialist step sometimes requires action on an external website the user controls: registering an API key, creating a vendor or developer account, configuring a dashboard, webhook, OAuth app, billing plan, or domain verification. This contract governs that moment. It grants no browsing authority; `SHARED-JUDGMENT.md` and `AUTHORITY-POLICY.md` remain binding, including approval before spending. + +1. Never hand the user a manual step list for a third-party site without first checking for an agentic browser that can act across their logged-in accounts. Today that is the Aside AI browser (its section in `references/BROWSER-PROVIDERS.md` owns detection and readiness); confirm the `aside` CLI is callable, for example `aside --version`. Never install it, and never treat binary presence as consent to browse. + +2. If such a browser is available, STOP and ask one explicit question before any browsing: name the exact site and the exact actions (for example "create a test-mode API token in the Duffel dashboard"), then offer A) drive it now through the agentic browser, B) manual instructions, C) defer. The selection is per-task consent; never persist it as standing permission and never infer it from an earlier task. + +3. When driving, touch only the named site and actions. Password entry, new-account credential choice, payment, and identity verification remain user-performed steps: hand off and wait instead of acting. Prefer credential flows that never expose the secret to the agent, such as password-manager autofill or host-native copy. + +4. A captured secret (API key, token, webhook signing secret) never appears in chat output, logs, or shell history. Write it to a user-approved local file with owner-only permissions or the user's secret store, and keep generated-file destinations out of version control. Dashboard fields are often masked placeholders; verify the captured credential with one non-mutating API call before claiming success. + +5. If no agentic browser is available, or the user declines or defers, provide the manual steps and mark the step blocked on the user. Do not recommend or install new products to close the gap.