docs(rfc): RFC-5 polyglot task protocol — ecosystem citizenship for foreign-language tasks - #1687
docs(rfc): RFC-5 polyglot task protocol — ecosystem citizenship for foreign-language tasks#1687rickylabs wants to merge 26 commits into
Conversation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
|
[PHASE: RESEARCH] Run 5 opened — RFC-5 polyglot task protocol (ecosystem citizenship).
Next phase comment lands with Generated by Claude Code |
… lane split (run 5) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
…ps + engine audit Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
…ping, callback surface, restate/conformance raws Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
…tension register, defect map Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
|
[PHASE: RESEARCH] Research complete — two workflow rounds, 25/25 agents, corpus committed ( Round 1 (8 source groups + engine audit + adversarial critic): Faktory/Sidekiq · Celery v2/BullMQ sandboxed · Lambda Runtime API/LSP · Temporal + Hatchet/Inngest/Restate · sd_notify/gRPC-health/CloudEvents/W3C-trace-context · oRPC v1+v2 · OpenAPI/NSwag codegen · FFI callback channels. The engine audit upgraded our caveats into a 14-item defect register with file:line cites (new finds: full parent env leaked to every subprocess D-9, abort never kills the spawned child D-6, attempt hardcoded 0 D-5, timeout collapsed to failed on persist D-8, JS Round 2 (critic-driven): security/token scoping corpus (Lambda per-invoke creds, Temporal task tokens, K8s bound SA tokens, RFC 8693, macaroons/Biscuit) — previously a zero-evidence pillar; task→engine callback-surface ruling (attempt-bound verbs on the task channel, everything else on an authenticated loopback oRPC surface — even Tier-0 tasks get full ecosystem access); Restate steals ratified (yanked-version registry hard-confirmed; terminal-frame discipline double-attested; cancellation settled as cooperative-only-on-duplex — T0/T1 keep OS signals, stated honestly); conformance harness designed from gRPC-interop/Autobahn/CloudEvents precedent ( Eight ecosystems converge on the same skeleton (research.md §2): protocol bytes off the payload channel, context in the envelope, structured errors with coordinator-decided retry, terminal-frame discipline, capabilities over version handshakes, heartbeat replies as the control channel, forward-only lifecycle DAG, absolute deadlines. Next: plan.md locking the §6 tension register (5 spikes flagged: frame transport, token delivery, loopback TCP-vs-UDS + sandbox survival, stdin buffering, protocol overhead vs the run-1–4 baselines), with pre-registered verdict criteria and the out-of-scope register — then PLAN-EVAL dispatch (required this run; Generated by Claude Code |
…-scope register Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
|
[PHASE: PLAN] Plan locked (2a8fb8b) — PLAN-EVAL dispatch requested via
Six spikes (L8, K1–K6) carry pre-registered decision criteria — headline: K4 protocol-overhead bar (T1 framing adds ≤1.0ms to the Go subject's 6.2ms exec-wall p50, ≤5% e2e at c=16) so the protocol cannot silently eat the execution wins of #1678–#1686; K5 may honestly fail (then T1 cancellation stays signal-only, with the Restate citation licensing exactly that). L9 pre-registers the RFC's own verdict criteria, and the out-of-scope register (9 items) applies the completeness-probe correction from the series miss — anything a reviewer thinks belongs in scope is a FAIL_RESCOPE conversation, not a silent change. Hard stop is in effect: no spike slice starts before PLAN-EVAL PASS. Evaluator inputs: Generated by Claude Code |
|
@openhands-agent model=openrouter/minimax/minimax-m3 output=pr-comment iterations=500 phase=plan head=2a8fb8b5b9f079088694bd1d1746b2dbf76257fb Trusted base SHA: 9634735 use harness SKILL
Act as the formal PLAN-EVAL session for this pull request. Do not edit files, create commits, push, Return concise findings with exact evidence and required amendments. End with exactly one supported |
OpenHands Agent — CompletedOPENHANDS_VERDICT: FAIL_PLAN Model: Summary — PLAN-EVAL for RFC-5 polyglot task protocol (PR #1687)SummaryFormal PLAN-EVAL session for
Changes
Validation
Responses to review comments or issue commentsNone received on PR #1687 during this session. The PR is a draft still; the Remaining risks
VerdictOPENHANDS_VERDICT: FAIL_PLAN Run: https://github.com/rickylabs/netscript/actions/runs/32340551406 |
|
OPENHANDS_VERDICT: FAIL_PLAN PLAN-EVAL checklist (gates/plan-gate.md):
The plan is substantively excellent — research is thorough, L1–L7 are locked with corpus Open-decision sweep (evaluator-run):
Required fixes to unblock PASS:
After Fix 1–6 land the plan re-submits via openhands + status:plan-eval on PR #1687. The Full deliverable: |
…mit slices, risk register Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
|
[PHASE: PLAN — cycle-1 fixes applied] All six PLAN-EVAL required fixes landed in
Cycle-1 verdict and dispositions mirrored in the run dir's Generated by Claude Code |
|
@openhands-agent model=openrouter/minimax/minimax-m3 output=pr-comment iterations=500 phase=plan head=f4ae0898e7d95427373af8efc0571da67bd8e629 Trusted base SHA: 9634735 use harness SKILL
Act as the formal PLAN-EVAL session for this pull request. Do not edit files, create commits, push, Return concise findings with exact evidence and required amendments. End with exactly one supported |
|
@openhands-agent model=openrouter/deepseek/deepseek-v4-flash-0731 output=pr-comment iterations=800 phase=impl head=45b134aadb477e6bcc5c0821fae28b2f7dea8f9d Trusted base SHA: 9634735 use harness SKILL
Act as the formal IMPL-EVAL session for this pull request. Do not edit files, create commits, push, Return concise, severity-ranked findings with exact evidence and required action. End with exactly |
OpenHands Agent — CompletedOPENHANDS_VERDICT: FAIL_FIX Model: OPENHANDS_VERDICT: FAIL_FIX IMPL-EVAL — RFC-5 polyglot task protocol (PR #1687) — re-dispatch on S10 headSummaryFormal IMPL-EVAL (separate evaluator session) for the RFC-5 polyglot-task-protocol deliverable at The revised RFC is substantively sound and reference-complete: PLAN-EVAL chain on record (cycle-1 However, I found concrete documentation/bookkeeping defects that must be fixed before the Changes
Validation
Findings (severity-ranked)
No FAIL_RESCOPE / FAIL_DEBT / FAIL_PLAN — plan is valid and in-scope; only small doc fixes are owed. Responses to review commentsNone required — no open/outdated review threads were unanswered. Remaining risks
VerdictOPENHANDS_VERDICT: FAIL_FIX Run: https://github.com/rickylabs/netscript/actions/runs/32358375889 |
…text-pack, seam cite) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
|
[PHASE: IMPL-EVAL — cycle-2 fixes applied] All four FAIL_FIX findings from run 32358375889 dispositioned in
Worth noting from the cycle-2 report: the evaluator judged the S10 design itself "substantively sound and reference-complete" — it hand-tested the Zod 4.4.3 normative snippets, verified every engine-seam file:line citation against source, and re-checked the K1–K6 figures; the findings were bookkeeping only. Eval-loop accounting: first FAIL_FIX of the IMPL-EVAL loop (two-failure limit). Mirrored in Generated by Claude Code |
|
@openhands-agent model=openrouter/deepseek/deepseek-v4-flash-0731 output=pr-comment iterations=800 phase=impl head=c45d6c1cf6cbb7f79d047b9dc694d8d20d406e31 Trusted base SHA: 9634735 use harness SKILL
Act as the formal IMPL-EVAL session for this pull request. Do not edit files, create commits, push, Return concise, severity-ranked findings with exact evidence and required action. End with exactly |
OpenHands Agent — CompletedOPENHANDS_VERDICT: FAIL_FIX Model: IMPL-EVAL cycle-3 (OpenHands) — RFC-5 polyglot task protocolTrusted base: 9634735 SummaryThe sole changed deliverable is
Validation
Findings (severity-ranked)F-1 [moderate — doc accuracy, FAIL_FIX] Summary misstates the plan's retired-defect register
Remaining (minor, non-blocking, consistent with R5 records, no action):
Responses to review commentsBlocking per instruction (evaluator session; do not repair). This cycle reports FAIL_FIX with the Remaining risks
OPENHANDS_VERDICT: FAIL_FIX Run: https://github.com/rickylabs/netscript/actions/runs/32359981617 |
…aligned with wave tables Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
|
[PHASE: IMPL-EVAL — cycle-3 fix applied; escalating per the two-failure rule] Cycle 3 (run 32359981617) confirmed all four cycle-2 fixes and re-verified every seam citation and spike figure, then returned FAIL_FIX on one moderate finding: the Summary's headline claim said the five-wave plan retires "D-1..D-10/D-12/D-13", contradicting the RFC's own wave/seam tables (which retire D-1..D-9, D-12, D-13, D-14 — D-10 is not wave-retired). The evaluator was right on both counts. Fixed in Escalation: this was the second FAIL_FIX of the IMPL-EVAL loop, which reaches the harness's two-failure limit — so per
Until you rule, the DoD verdict box stays open and Generated by Claude Code |
…ia OpenHands (R5-D-8) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
|
@openhands-agent model=openrouter/x-ai/grok-4.6 output=pr-comment iterations=800 use harness SKILL
ROLEAct as the OWNER-COMMISSIONED ADVERSARIAL REVIEWER for RFC-5 on PR #1687, at HIGH reasoning effort throughout. This is a third pass ON TOP of the formal gates — do not re-run their checklists, and do not re-discover their findings: PLAN-EVAL cycle-2 PASS (run 32343592955); IMPL-EVAL cycle-1 PASS on the pre-revision head (superseded by the S10 redesign), cycle-2 FAIL_FIX (4 doc fixes, applied in c45d6c1), cycle-3 FAIL_FIX (1 doc fix — defect-retirement claim — applied in 02b1c6e). No design finding has survived any cycle; every applied fix is verifiable in the commit trail. Your job is to BREAK the design and the argument. Read-only: do not edit files, commit, push, or repair findings; never mutate CONTEXTRun 5 of the polyglot RFC series (#1678 scriptc, #1683 rust, #1685 dotnet, #1686 golang — merged execution-layer runs). RFC-5 specifies the NetScript Task Protocol (NTP): ecosystem citizenship for polyglot tasks — interoperability, observability, communication layer, error & lifecycle management. Read in order:
Core decisions to attack: three conformance tiers computed by a conformance suite (T0 legacy-forever / T1 structured one-shot / T2 duplex worker); two surfaces (attempt-bound task-channel verbs vs authenticated loopback oRPC citizen surface, TCP 127.0.0.1, per-attempt opaque capability tokens invalidated on retry); sentinel-NDJSON stdout framing ( ACCEPTED caveats — do NOT report as discoveries, but DO attack whether accepting them was sound: K6 measured on a replica (R5-D-3); fd-3 infeasible on the Deno host (R5-D-2); UDS demoted (R5-D-4); Docker/Aspire/Windows loopback survival untested (R5-D-5); sentinel-forgery residual (RFC Drawbacks #2); engine defects D-1..D-14 deliberately not fixed in this docs-only PR; D-10/D-11 deliberately outside the five-wave retirement. ATTACK SURFACE (in priority order)
OUTPUTWrite the full deliverable to Generated by Claude Code |
…R5-D-8 update) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
…wen3.8-max R5-D-8 resolved: owner chose the allowlisted open-model fallback after the Grok 4.6 dispatch was rejected by the main-ref workflow allowlist. No workflow or config change; dispatch proceeds on the refreshed head. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
|
@openhands-agent model=openrouter/qwen/qwen3.8-max output=pr-comment iterations=800 use harness SKILL
ROLEAct as the OWNER-COMMISSIONED ADVERSARIAL REVIEWER for RFC-5 on PR #1687, at HIGH reasoning effort throughout. This is a third pass ON TOP of the formal gates — do not re-run their checklists, and do not re-discover their findings: PLAN-EVAL cycle-2 PASS (run 32343592955); IMPL-EVAL cycle-1 PASS on the pre-revision head (superseded by the S10 redesign), cycle-2 FAIL_FIX (4 doc fixes, applied in c45d6c1), cycle-3 FAIL_FIX (1 doc fix — defect-retirement claim — applied in 02b1c6e). No design finding has survived any cycle; every applied fix is verifiable in the commit trail. This pass also serves as the closing verification for the IMPL-EVAL two-failure escalation (owner ruling in drift R5-D-8). Your job is to BREAK the design and the argument. Read-only: do not edit files, commit, push, or repair findings; never mutate CONTEXTRun 5 of the polyglot RFC series (#1678 scriptc, #1683 rust, #1685 dotnet, #1686 golang — merged execution-layer runs). RFC-5 specifies the NetScript Task Protocol (NTP): ecosystem citizenship for polyglot tasks — interoperability, observability, communication layer, error & lifecycle management. Read in order:
Core decisions to attack: three conformance tiers computed by a conformance suite (T0 legacy-forever / T1 structured one-shot / T2 duplex worker); two surfaces (attempt-bound task-channel verbs vs authenticated loopback oRPC citizen surface, TCP 127.0.0.1, per-attempt opaque capability tokens invalidated on retry); sentinel-NDJSON stdout framing ( ACCEPTED caveats — do NOT report as discoveries, but DO attack whether accepting them was sound: K6 measured on a replica (R5-D-3); fd-3 infeasible on the Deno host (R5-D-2); UDS demoted (R5-D-4); Docker/Aspire/Windows loopback survival untested (R5-D-5); sentinel-forgery residual (RFC Drawbacks #2); engine defects D-1..D-14 deliberately not fixed in this docs-only PR; D-10/D-11 deliberately outside the five-wave retirement. ATTACK SURFACE (in priority order)
OUTPUTWrite the full deliverable to Generated by Claude Code |
OpenHands Agent — CompletedOPENHANDS_VERDICT: FAIL_FIX Model: Adversarial review — RFC-5 NetScript Task Protocol (PR #1687)Pass: owner-commissioned adversarial review (third pass on top of formal gates; closing Non-duplication: PLAN-EVAL c2 PASS and IMPL-EVAL c1–c3 findings were read before this pass. What survived attack (not findings): sentinel-scan demux soundness (K1 corpus, partial-sentinel
Findings (severity-ranked)MAJOR-1 — Frame grammar cannot represent a >~4 KB task result: terminal
|
|
@openhands-agent model=openrouter/qwen/qwen3.8-max output=pr-comment iterations=800 use harness RE-DISPATCH NOTE: attempt 1 (run 32398533310) stalled mid-analysis with no log output for 100+ minutes and was abandoned; this is a fresh attempt with an identical brief. If that run ever posts a late verdict, THIS run's verdict is the one of record. SKILL
ROLEAct as the OWNER-COMMISSIONED ADVERSARIAL REVIEWER for RFC-5 on PR #1687, at HIGH reasoning effort throughout. This is a third pass ON TOP of the formal gates — do not re-run their checklists, and do not re-discover their findings: PLAN-EVAL cycle-2 PASS (run 32343592955); IMPL-EVAL cycle-1 PASS on the pre-revision head (superseded by the S10 redesign), cycle-2 FAIL_FIX (4 doc fixes, applied in c45d6c1), cycle-3 FAIL_FIX (1 doc fix — defect-retirement claim — applied in 02b1c6e). No design finding has survived any cycle; every applied fix is verifiable in the commit trail. This pass also serves as the closing verification for the IMPL-EVAL two-failure escalation (owner ruling in drift R5-D-8). Your job is to BREAK the design and the argument. Read-only: do not edit files, commit, push, or repair findings; never mutate CONTEXTRun 5 of the polyglot RFC series (#1678 scriptc, #1683 rust, #1685 dotnet, #1686 golang — merged execution-layer runs). RFC-5 specifies the NetScript Task Protocol (NTP): ecosystem citizenship for polyglot tasks — interoperability, observability, communication layer, error & lifecycle management. Read in order:
Core decisions to attack: three conformance tiers computed by a conformance suite (T0 legacy-forever / T1 structured one-shot / T2 duplex worker); two surfaces (attempt-bound task-channel verbs vs authenticated loopback oRPC citizen surface, TCP 127.0.0.1, per-attempt opaque capability tokens invalidated on retry); sentinel-NDJSON stdout framing ( ACCEPTED caveats — do NOT report as discoveries, but DO attack whether accepting them was sound: K6 measured on a replica (R5-D-3); fd-3 infeasible on the Deno host (R5-D-2); UDS demoted (R5-D-4); Docker/Aspire/Windows loopback survival untested (R5-D-5); sentinel-forgery residual (RFC Drawbacks #2); engine defects D-1..D-14 deliberately not fixed in this docs-only PR; D-10/D-11 deliberately outside the five-wave retirement. ATTACK SURFACE (in priority order)
OUTPUTWrite the full deliverable to Generated by Claude Code |
…h (R5-D-8) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
OpenHands Agent — CompletedOPENHANDS_VERDICT: FAIL_FIX Model: RFC-5 (PR #1687) - Owner-commissioned adversarial pass (attempt 2, run 32416122194)Role: owner-commissioned adversarial reviewer, third pass ON TOP of the formal gates; closing Bottom line: the architecture survives the attack. Tiers, two surfaces, sentinel framing, Not re-reported: PLAN-EVAL cycle-2 PASS (run 32343592955) and all IMPL-EVAL findings (cycle-2 IMPL-EVAL two-failure escalation - closing verification (brief-mandated)
MAJOR findingsA1 - The retry/redelivery driver is load-bearing and owned by no wave (migration realism)Cites: RFC lifecycle section lines 571-588 ( The hole: the RFC specifies an attempt state machine whose RETRY, PAUSE, checkpoint-redelivery, Concrete scenario: W2-W3 land. A Go task emits Fix direction: either (a) add an explicit wave/scope row owning the retry driver (behavior A2 - Redelivery identity semantics unspecified; token fencing does not cover the case it promises (wire + security)Cites: RFC edge-case row line 659 ("Queue redelivery races a live attempt | attempt tokens The hole: the RFC never says whether a redelivered message reuses the execution record Concrete scenario: task A runs 40 s under the default 300 s timeout. The KV visibility timeout is Fix direction: normative edge rows: (1) redelivery identity - reuse executionId, increment A3 - Bootstrap-token lifetime and caller binding unspecified -> confused-deputy (security)Cites: RFC line 457 ("bound to 127.0.0.1: with a per-boot SSRF bearer (spike K3 The hole: the RFC contradicts itself on bootstrap scope ("per-boot" line 457 vs per-task env Concrete scenario (per-boot reading): task A is declared with capabilities(['kv:reports/']); Fix direction: one sentence, normative: bootstrap tokens are per-execution, and A4 - Frame size limits are transport-dependent but legislated universally; envelope size is unbounded (frame grammar)Cites: RFC line 105 (Frame = single write(2) <= PIPE_BUF), lines 343-344 The hole, in three clauses: (1) T2 checkpoint routing "8-256 KB via inbound frame" contradicts Concrete scenario: a C# SDK author targets Windows and caps payloads at 24 KB; a Rust author Fix direction: define (a) an envelope size budget + spill rule (payloads over budget -> artifact A5 - resultSchema validation failure has no semantics (spec ambiguity)Cites: RFC line 313 (value ... validated against resultSchema when 'ok'); line 451 (T0 fallback The hole: the RFC specifies the failure semantics of payload validation (authoring error, FAIL, Concrete scenario: task declares resultSchema = z.object({url: z.string().url()}), emits Fix direction: one edge-case row: result-validation failure -> attempt ends failed with A6 - Normative progress throttle contradicts the spike it cites and the W4 acceptance bar (evidence validity + spec)Cites: RFC line 454 ("Latest-wins throttle (min(0.8x timeout, 30 s), Temporal shape) lives in The hole: min(0.8x timeout, 30 s) evaluates to a 30 s flush interval at the default 300 s Concrete scenario: W4 implements line 454 literally; the in-plugin K6 re-measurement reports Fix direction: correct line 454 to the measured design (state the flush interval constant, e.g. MINOR findingsM1 - Cancellation ladder gaps (grace defaults, kill timing, drain-after-kill). Lines 589-591, M2 - Duplicate-terminal vocabulary mismatch. Line 655: duplicate result frames log M3 - T2 init silence: demotion vs zombie. Lines 603-606 ("no echo -> T0/T1 detection") vs lines M4 - Conformance exclusion authority undefined. Lines 630-631: "Tier N achieved iff all M5 - Invalidation trigger list incomplete (lease requeue). Lines 608-609: missed-ping budget M6 - Attempt arithmetic undefined. Lines 262, 550 (z.int().min(0)), audit D-5 (hardcoded 0). M7 - "Closed envelope" vs open-world MUST. Summary line 22 ("closed, versioned envelope") OBSERVATIONS (including accepted-caveat soundness attacks)O1 - K4 evidence: verdict stands, presentation overclaims at c=16. Negative deltas O2 - Accepted caveats were soundly accepted. K6-on-replica (R5-D-3): sound conditional on A6 - O3 - "None at T0" vs the env allowlist. Lines 737-745 admit the allowlist is "the one O4 - Doctrine judgment on the package split: COMPLIANT. The in-core protocol/ module with O5 - Completeness probe: no FAIL_RESCOPE. The owner's ask - tasks "actually part of the O6 - Deadline/clock model is sound. Line 660's resolution (host-authored, host-enforced, What I could NOT break
Fix priorityA3 (security hole) > A1 (unimplementable lifecycle) > A4 (unrepresentable payloads) > A2 (races)
Validation performed this pass (read-only)
ADVERSARIAL_VERDICT: CONCERNS Run: https://github.com/rickylabs/netscript/actions/runs/32416122194 |
…-1 stall-diagnosis correction Attempt 1 (qwen3.8-max, run 32398533310) completed after 4h21m with the full review on head 4242a46; 4 MAJOR + 4 MINOR findings, all MAJORs supervisor-verified against RFC/spike sources. Attempt-2 of-record clause superseded; disposition returned to owner with S11 proposal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
…nsolidated S11 scope Attempt 2 (run 32416122194) completed uncancelled with CONCERNS/FAIL_FIX; of-record status unchanged (attempt 1). Novel findings A2-A5, M1-M7 and three observation notes folded into the proposed S11 scope; A1/A6 converge with the of-record MAJOR-2/MAJOR-3. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5
Summary
Run 5 of the polyglot series — and a deliberate pivot. Runs 1–4 (#1678 scriptc, #1683 rust-workers, #1685 dotnet, #1686 golang) measured execution and inherited the polyglot contract ("
TASK_ID/TASK_PAYLOADenv in, last JSON line of stdout out") as a fixed seam. This run challenges the seam itself (owner brief, 2026-08-20): evolve polyglot tasks from foreign-language runners into ecosystem citizens — interoperability, observability, communication layer, error & lifecycle management — as one generic, cross-language protocol.Delivered:
rfcs/0000-polyglot-task-protocol.md— the NetScript Task Protocol (NTP), revised in S10 (owner content review, drift R5-D-7) from a findings report into a full architectural design proposal (803 lines, 26 code fences, at the accepted-RFC 0001 bar): normative Zod wire schemas (EnvelopeV1,TaskFrameV1/HostFrameV1unions,StructuredErrorV1withRETRY|PAUSE|FAIL); TypeScript port interfaces on the auth-core blueprint (compositeTaskProtocolBackendPort+ narrow ports + typed unsupported errors + named registry); a file-level package map (plugin-workers-core/src/protocol/,testing/conformance/, citizen contracts,withTaskProtocoldecorator, KV token store, scaffold-asset shims); a per-seam engine-integration table (file:line, refactored-vs-preserved, defect linkage D-1..D-14); bound builder generics + theTaskOutcomeV1union; the citizen-surface oRPC contract with a route×capability table and numbered token-verification algorithm; lifecycle state machine; T2 handshake with zombie rules; conformance case grammar; six-axis extension model; edge-case resolutions; and a five-wave staged implementation plan (W1 protocol-core → W5 T2 worker) with acceptance bars tied to the spike numbers. Three conformance tiers (T0 legacy-forever / T1 structured / T2 duplex) are computed by the conformance suite; no TaskType vocabulary changes (series precedent, 5×).Evidence (Appendix A): 32-file ratified research corpus (two workflow rounds, 25/25 agents), engine audit with file:line defect register, and six pre-registered spikes all passing their primary criteria — headline: the full T1 contract costs +0.41 ms through the real dispatch path (1920/1920 executions), sandbox-scoped loopback at 0.49 ms p50, in-band cancel acked at 3.5/30 ms p50 (Go/python3).
Scope
rfcs/0000-polyglot-task-protocol.md— the RFC (docs only)..llm/runs/claude-harness-profile-rfc-benchmark-shzhgv--polyglot-protocol-rfc/— run artifacts: supervisor/research/plan/plan-eval/evaluate/worklog/drift/context-pack,research-sources/corpus (32 files), spike code K1–K6, raw JSONL + generatedresults-spikes.md.packages/orplugins/source is touched (G4; independently re-confirmed by both evaluator cycles).Slices
S1 bootstrap · S2 research round 1 · S3 round 2 + synthesis · S4 plan + PLAN-EVAL (cycle-1 FAIL_PLAN → fixes → cycle-2 PASS, run 32343592955) · S5 K1+K2 · S6 K3+K5 · S7 K4+K6 · S8 spike synthesis + drift · S9 RFC (first authoring) · S10 owner-review revision (report → architectural design; R5-D-7) · S10b IMPL-EVAL cycle-2 fixes F1–F4. Per-slice commits + PR comments are the trail.
Validation
ci:skip-e2e+ci:skip-scaffold).Harness
.llm/runs/claude-harness-profile-rfc-benchmark-shzhgv--polyglot-protocol-rfc/(ARCHETYPE-3 + SCOPE-docs).plan-eval.md).70d101a(run 32346098261 — superseded by the S10 rewrite, kept as record); cycle 2 on S10 head45b134a(run 32358375889) — FAIL_FIX with design judged "substantively sound and reference-complete" and four doc fixes (F1 corpus count, F2 context-pack, F3 this body, F4 seam cite path), all applied inc45d6c1; cycle 3 re-dispatched on the fixed head. Full chain mirrored inevaluate.md.Drift / Debt
R5-D-1 PLAN-EVAL cycle-1 fixes · R5-D-2 fd-3 infeasible on Deno host · R5-D-3 K6 replica · R5-D-4 UDS constraints · R5-D-5 container/Aspire/Windows untested · R5-D-6 zod lock-alias reverted · R5-D-7 owner content review → S10 redesign. Full detail in
drift.md.Definition of Done
Refs #1679, #1684 (protocol intersects the monty sandbox spike and the web-worker pool).
🤖 Generated with Claude Code
https://claude.ai/code/session_013H2FUAx1v6BbP6PgLTNqH5