Skip to content

feat(dashboard): visual revamp to reference bar (adversarial-gated) - #780

Draft
rickylabs wants to merge 14 commits into
mainfrom
feat/dashboard-visual-revamp
Draft

feat(dashboard): visual revamp to reference bar (adversarial-gated)#780
rickylabs wants to merge 14 commits into
mainfrom
feat/dashboard-visual-revamp

Conversation

@rickylabs

Copy link
Copy Markdown
Owner

Dashboard visual revamp — to the reference bar, adversarially gated

Raises the NetScript Dev Dashboard prototype to a premium operator-console quality bar (studied
against real product references), in NS One identity (warm-cream/dark, hard press shadow,
copper/teal/amber, --ns-* tokens, ns-* markup that round-trips to Fresh). Visual + layout only —
no routes/logic/data/copy changes.

What's here (this commit)

  • Baseline references + DESIGN-LANGUAGE.md, ROLLOUT-DOCTRINE.md, HOME-SPEC.md — the bar and
    the per-screen doctrine (each screen bespoke; mine a different reference each pass; no two screens
    alike).
  • DS-UPLIFT-BACKLOG.md — every new component/token/variant/option/mobile pattern the redesign
    needs, accumulated per pass. Drives a NS One design-system uplift after the prototype is signed
    off, before building the real product.
  • Per-pass reports (render/_visual-reports/*.md): V2b Home, V3 investigation spine, V4 consoles.
  • Adversarial gate: Kimi K2.6 (via OpenCode) — structural (code) + vision (screenshots) — checks
    each pass for distinctness + variety-leverage before acceptance (visual/_evals/).

Cadence

The evolving prototype source (render/prototype.dc.html + assets/*.css) and each new pass's report
are committed here as each visual pass completes (never mid-edit), with a PR comment per slice +
before/after screenshots. Passes so far: Home, spine, Workers (excellent), Sagas (being re-done as an
n8n-style node canvas). Rollout continues: control plane, AI, extensions, final polish.

The prototype is a Claude Design working record; it round-trips to @netscript/fresh-ui source later.

Baseline references, design-language + rollout doctrine, Home spec, DS-uplift
ledger, and per-pass reports (V2b Home, V3 investigation spine, V4 consoles).
Prototype source + per-pass snapshots land as each visual pass completes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Sagas (V4b) — n8n node-canvas redesign

The rejected 3-column console template is replaced by a canvas-dominant node graph: a forward-path lane + a distinct compensation/rollback lane, wired by JS-measured bezier connectors (no SVG {{ }} holes), click-a-node → properties sidebar (touching edges recolor copper), mobile → bottom sheet, fully data-driven (the no-compensation export instance renders only the forward lane). NS One identity only (--ns-* tokens, ns-* markup, warm-cream/dark, hard press shadow, copper/teal/amber); routes/data/copy-meaning unchanged.

Adversarial gate — Kimi K2.6 (vision, native OpenCode), reading the shots against the n8n flow-builder + Kafka-console references: ACCEPT-WITH-FIXES, bespoke 74/100. Confirmed genuinely saga-tailored. One real defect: ~75–80% of the canvas is empty dotted grid → fails the "dense, no dead space" bar.

Gate-directed fix pass now in flight:

  • Fill the void (mined from ref-21 n8n / ref-11 console): definition mini-map with active-path overlay, a per-lane compensation-health stat strip (avg latency / retry rate / compensation success %), richer edge error badges (e.g. E_TIMEOUT on the reserving → refund wire), a slim step execution rail, and ghost-node outlines for queued/terminal states so the graph owns the canvas.
  • Variety-leverage: replace the airy step-timeline bullet list with a dense sortable step-history table (ref-11), tab the right detail panel (State / Transitions / Grounded / Actions), and give nodes stronger color identity (forward teal-tinted, compensation copper-tinted, terminal neutral) per ref-21 — within NS One tokens.

Source snapshot + refreshed screenshots land after the fix pass. Report: V4b-sagas-canvas.md · gate transcript: V4b-adversarial-vision.md.

Sagas capability screen rebuilt as a canvas-dominant n8n-style node graph
(forward + compensation lanes, JS-measured bezier connectors, click-to-open
tabbed inspector, mobile bottom sheet, fully data-driven), then density-fixed
and every section made bespoke:

- canvas coverage 29%→47%: definition mini-map, per-lane compensation-health
  stat strips, E_TIMEOUT edge badges, step-execution rail, ghost-node outlines
- 4 plain metric cards → saga-health band (comp-rate donut, in-flight/terminal
  split, instances-by-state channel bar, forward/rollback edge ratio)
- airy step list → dense sortable step-history table; transitions → typed
  transition stream (advance→ / compensate↩ / terminal◼), grouped by instance

Adversarial vision gate (Kimi K2.6, native OpenCode) across the n8n +
dev-console + pm-dashboard references: ACCEPT-WITH-FIXES, bespoke 82/100,
canvas density now passing. Residual chrome-polish prescriptions logged for
the final sweep. NS One identity only; routes/data/copy unchanged; no SVG holes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Sagas — accepted at bespoke 82/100 (committed b33a1338)

Density fix + whole-screen-bespoke pass landed. The empty-grid dead-space is resolved (canvas coverage 29%→47%; the stronger gate now rates canvas density passing, <5% empty). Surrounding chrome rebuilt bespoke:

  • 4 plain metric cards → saga-health band (compensation-rate donut, in-flight/terminal split, instances-by-state channel bar, forward/rollback edge ratio)
  • per-lane compensation-health stat strips, E_TIMEOUT edge badges, step-execution rail, ghost-node outlines, definition mini-map + legend
  • airy step list → dense sortable step-history table; transitions → typed transition stream (advance→ / compensate↩ / terminal◼)
  • tabbed node inspector (State / Transitions / Evidence / Actions); mobile bottom sheet

Gate — now emits a concrete component-prescription table (≥8 rows) naming exact elements in named references (n8n flow-builder / Kafka dev-console / pm-dashboard), fed a diverse ref set via native OpenCode Kimi K2.6 vision: ACCEPT-WITH-FIXES, bespoke 82/100 (↑ from 74). Transcript: visual/_evals/V4b-fix-adversarial-vision.md.

Residual is chrome polish (stat-cards → click-to-filter bar; transition-stream → timeline connectors; payload → syntax-highlighted editor; retry → action palette) — logged for the final polish sweep rather than gold-plating one screen. Moving to the next screen.

…5→62)

Triggers capability screen rebuilt with its own composition (distinct from the
Sagas node-canvas and Workers registry):

- tinted category-health band (Schedule/Webhook/Event/Manual) with per-type
  sparkbars, armed/silenced pills, click-to-filter
- WHEN→DO rule builder (event → condition chip → action) + action-type palette
- schedule day-strip (7-day next-fire horizon) + temporal next-fire cards with
  progress-to-fire bars
- dense sortable firing-history table with a wired query toolbar (quick-search,
  outcome filter, show-limit) and per-row 24h sparklines + inline arm switches
- iconified metric strip (shield/flame/alert + micro-sparkline + trend)
- dead-letters as a compact mini-table (error snippet, age, retry-count, retry/
  inspect) + depth summary
- tabbed detail drawer (action-chain / payload / schedule) → mobile bottom sheet

Adversarial vision gate (Kimi K2.6, native OpenCode) across n8n / dev-console /
pm-dashboard refs: first pass 55 → focused fix pass resolved all five concrete
defects → 62, ACCEPT-WITH-FIXES. Residual gate asks (node-graph builder, code-
editor test box, segmented row bars) are subjective/gold-plating and are logged
for the final polish sweep. NS One tokens only; routes/data/copy unchanged; no
SVG holes; 16-route regression clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Triggers — accepted (committed e0da07c4)

Bespoke rules-and-firings console — its own composition, not a Sagas/Workers clone:

  • category-health band (Schedule/Webhook/Event/Manual: per-type sparkbars, armed/silenced pills, click-to-filter)
  • WHEN→DO rule builder (event → condition chip → action) + action-type palette
  • schedule day-strip (7-day next-fire horizon) + temporal next-fire cards with progress-to-fire bars
  • dense firing-history table + a wired query toolbar (quick-search, outcome filter, show-limit) + per-row 24h sparklines + inline arm switches
  • iconified metric strip (shield/flame/alert + micro-sparkline + trend)
  • dead-letters mini-table (error snippet, age, retry-count, inline retry/inspect) + depth summary
  • tabbed detail drawer (action-chain / payload / schedule) → mobile bottom sheet

Gate (native OpenCode Kimi K2.6 vision, diverse refs): first pass 55 → focused fix pass resolved all 5 concrete defects → 62, ACCEPT-WITH-FIXES. Remaining gate asks (turn the builder into a full node-graph, a code editor in the webhook test box, segmented row bars) are subjective/gold-plating — logged for the final polish sweep rather than looping one screen against an insatiable gate. Transcripts: V5-triggers-adversarial-vision.md, V5-triggers-fix-adversarial-vision.md.

Next: Streams.

… 72)

Streams rebuilt as a topic console — every section stream-native and distinct
from all neighbours:

- streams-health composite header (throughput sparkline / consumer lag + back-
  pressure / partition pips / retention) — not plain number cards
- live throughput hero: JS-measured teal area chart + hover crosshair + floating
  value readout + peak/avg/floor footer + backpressure pill
- producer → partitioned-log → consumer topology fan-bus (per-group lag)
- partition & offset ledger with inline lag heat-bars
- partition × time lag heatmap
- consumer-group radial lag gauges (Consumers tab)
- live event tail (monospace records) + fan-out delivery drawer (per-consumer
  attempt/outcome) → mobile bottom sheet
- Config tab: KV + line-numbered Avro schema

Adversarial vision gate (Kimi K2.6, native OpenCode) vs Kafka-console + finance
refs: 72/100 ACCEPT-WITH-FIXES; strong bespoke bones confirmed. Residual polish
(throughput time-range pills, live-tail filter bar, denser message panel) logged
to visual/_evals/POLISH-SWEEP.md for the final sweep. NS One tokens only; routes/
data/copy unchanged; no SVG holes; 16-route regression clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Streams — accepted at gate 72/100 (committed 1c50218b)

Bespoke durable-stream topic console — every section stream-native, distinct from all neighbours:

  • streams-health composite header (throughput sparkline / consumer-lag + backpressure / partition pips / retention)
  • live throughput hero — JS-measured area chart + hover crosshair + value readout + peak/avg/floor footer + backpressure pill
  • producer → log → consumer topology fan-bus (per-group lag)
  • partition & offset ledger with inline lag heat-bars + partition × time lag heatmap
  • consumer-group radial lag gauges (Consumers tab)
  • live event tail + fan-out delivery drawer (per-consumer attempt/outcome) → mobile bottom sheet
  • Config tab: KV + line-numbered Avro schema

Gate: 72/100 ACCEPT-WITH-FIXES — bones confirmed genuinely bespoke (topology, ledger, heatmap, fan-out). Residual polish (throughput time-range pills, live-tail filter bar, denser message panel) logged to visual/_evals/POLISH-SWEEP.md for the final sweep. Transcript: V6-streams-adversarial-vision.md.

Next: Config Resolution.

…xed (gate 68)

Config Resolution rebuilt around a novel, signature precedence-waterfall
metaphor (distinct from the four consoles):

- precedence waterfall: framework → package → env·profile → runtime override on
  a numbered rail; the winning layer is elevated with the hard press-shadow,
  shadowed values are struck through with source:line provenance, and non-
  contributing (not-set / passes-through) layers collapse to slim pips so
  partial-resolution states carry no dead space (silent-layer space −49%)
- effective-config table with namespace filters + winning-layer chips; effective
  values in DM Mono with type-color (numbers amber / strings teal / booleans copper)
- KPI header with supporting micro-viz (namespace ticks, override pips, a
  shadowed/total meter, precedence diamonds)
- framed side-by-side override diff (context lines + pink/green gutter, DM Mono),
  shown only when a runtime override wins
- resolution trail + detail drawer → mobile bottom sheet; fully data-driven per key

Adversarial vision gate (Kimi K2.6, native OpenCode) vs code-diff / console /
flow refs: 68/100 → focused density fix resolved the concrete hits (collapse
not-set layers, denser KPIs, framed diff, mono type-color). Deferred asks (node-
graph rail, inline override editor) logged to POLISH-SWEEP.md. NS One tokens only;
routes/data/copy unchanged; no SVG holes; 16-route regression clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Config Resolution — accepted (committed ee8f8ace)

A novel precedence-waterfall metaphor — visually distinct from all four consoles:

  • precedence waterfall: framework → package → env·profile → runtime override on a numbered rail; the winning layer is elevated (press-shadow), shadowed values struck through with source:line provenance, and non-contributing (not-set / passes-through) layers collapse to slim pips so partial-resolution states carry no dead space (silent-layer space −49%)
  • effective-config table with namespace filters + winning-layer chips; values in DM Mono with type-color (numbers amber / strings teal / booleans copper)
  • KPI header with micro-viz (namespace ticks, override pips, shadowed/total meter, precedence diamonds)
  • framed side-by-side override diff (context lines + pink/green gutter), shown only when a runtime override wins
  • resolution trail + detail drawer → mobile bottom sheet; data-driven per key (override-wins vs package-wins verified)

Gate: 68/100 → focused density fix resolved the concrete hits (collapse not-set layers, denser KPIs, framed diff, mono type-color). Deferred asks (node-graph rail, inline override editor) → POLISH-SWEEP.md. Transcript: V7-config-adversarial-vision.md.

Next: Migrations.

…-fixed (gate 62)

Migrations rebuilt around a bespoke migration version-chain timeline (distinct
from the Config waterfall and the capability consoles):

- version-chain timeline: applied spine (v1→v2→v3) → a slim LIVE DB · HEAD drift
  divider → a dashed-copper pending queue; each node carries an inline metadata
  row (objects-changed · rows-affected · duration + applied-at) so it earns its
  height with data, not padding
- schema-drift half-ring arc-gauge (n drifted / m-of-k objects in sync) + per-
  object drift deltas (has → wants, struck-through)
- header composite with supporting micro-viz (applied ratio bar, pending/drift
  pips, relative-time hint) — not plain number cards
- "1 migration pending" apply card with a what-will-change preview + CLI
- dense migration ledger table; per-migration detail drawer with the DDL diff
  correctly DEMOTED to a secondary tab → mobile bottom sheet

Adversarial vision gate (Kimi K2.6, native OpenCode) vs timeline / console /
diff refs: 62/100 → focused density fix (compact nodes w/ metadata, denser
header, slim divider). Gold-plating asks (ledger inline-diff, docs-tabs, DDL
editor chrome, invented author/step data) declined → POLISH-SWEEP.md. Also
front-loaded density defaults into ROLLOUT-DOCTRINE for future first passes.
NS One tokens only; routes/data/copy unchanged; no SVG holes; 16-route clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Migrations — accepted (committed e6568825)

A bespoke migration version-chain timeline — distinct from the Config waterfall and the consoles:

  • version-chain timeline: applied spine (v1→v2→v3) → a slim LIVE DB · HEAD drift divider → a dashed-copper pending queue; each node carries an inline metadata row (objects-changed · rows-affected · duration + applied-at) so it earns its height with data, not padding
  • schema-drift arc-gauge (n drifted / m-of-k in sync) + per-object drift deltas (has → wants, struck)
  • header composite with micro-viz (applied ratio bar, pending/drift pips, relative time)
  • apply card with a what-will-change preview + CLI; dense migration ledger; detail drawer with the DDL diff correctly demoted to a secondary tab → mobile sheet

Gate: 62/100 → focused density fix (compact nodes w/ metadata, denser header, slim divider). Gold-plating (ledger inline-diff, docs-tabs, DDL editor, invented author/step data) declined → POLISH-SWEEP.md. I also front-loaded the gate's recurring density feedback into ROLLOUT-DOCTRINE.md so future first passes land denser. Transcript: V8-migrations-adversarial-vision.md.

Next: Dead-Letter Queues.

…(gate 62, accepted on verified density)

DLQ rebuilt as an operational triage console — distinct from every prior screen:

- error-signature CLUSTER list (the signature): failures grouped by error mode
  (E_CONN ×3, E_TIMEOUT ×2 …) with source chips, a death sparkbar, age range, and
  per-cluster Retry-all / Purge + click-to-filter — not a flat list
- triage strip: composite tiles with a by-queue split micro-bar, retry-exhausted
  pips, and a red arrival sparkbar (no plain number cards)
- queue-health fill-meters (redis ELEVATED / postgres LOW / kv CLEAR) with depth-
  vs-capacity + trend sparkbar + oldest age (horizontal meters kept deliberately
  distinct from Migrations' arc gauge)
- dense failed-message table with retry pips + severity rails + bulk-select sticky
  bar; message inspector (payload JSON + retry-history timeline + gated reprocess/
  purge) → mobile bottom sheet
- data-driven: populated, all-clear, and filtered-empty states all render

Adversarial vision gate: 62/100. Accepted on verified density (9/10) — the gate's
major asks were already present (micro-viz, cluster severity borders), functionally
wrong (remove bulk-select checkboxes), or variety-declined (arc gauges). Two trivial
residuals (frame payload JSON, connect retry timeline) logged to POLISH-SWEEP.md.
NS One tokens only; routes/data/copy unchanged; no SVG holes; 16-route clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Dead-Letter Queues — accepted on verified density (committed 9b3d5427)

An operational triage console, distinct from every prior screen:

  • error-signature clusters (the signature): failures grouped by mode (E_CONN ×3, E_TIMEOUT ×2…) with source chips, death sparkbar, age range, per-cluster Retry-all/Purge + click-to-filter
  • triage strip with by-queue split micro-bar, retry-exhausted pips, arrival sparkbar
  • queue-health fill-meters (redis ELEVATED / postgres LOW / kv CLEAR) — kept horizontal on purpose so DLQ ≠ Migrations' arc gauge
  • dense failed-message table (retry pips + severity rails + bulk-select) + message inspector (payload JSON + retry-history timeline + gated reprocess/purge) → mobile sheet

Gate: 62/100 — accepted on verified density (9/10). The gate over-reached: its major asks were already present (micro-viz, cluster severity borders), functionally wrong (it wanted to remove the bulk-select checkboxes that make triage work), or variety-declined (arc gauges). Two trivial residuals (frame payload JSON, connect retry timeline) → POLISH-SWEEP.md. This is where the gate's absolute score stops being reliable and judgment takes over. Transcript: V9-dlq-adversarial-vision.md.

Next: Catalog.

…ted on verified quality)

Catalog rebuilt as a contract-registry explorer — distinct from every monitoring
screen (browse → drill into a contract, not a console):

- header composite: by-kind distribution channel-bar (click-to-filter) + describe()
  coverage ratio meter + REST/RPC/SDK transport-reach tri-meter (no plain cards)
- filter toolbar: search + kind/method/coverage chips + reset; live-filters the
  rail, hides empty groups, has a no-results state
- grouped registry rail: kind-grouped, method-tinted selectable contract cards
  (method chip + mono proc + input→output preview + duality tri-pip + coverage dot)
  + per-group coverage micro-bar + gated "install crons" affordance
- contract detail: version + compat chip, method/transport strip, a line-numbered
  TYPE-COLORED signature + secondary schema, and a producers→consumers strip
- mobile bottom sheet with the full signature/schema/relationships

Adversarial vision gate: 62/100. Accepted on verified quality (9/10). The gate
over-reached with OFF-DOCTRINE asks — raw-hex green buttons, a centered modal to
replace the intentional mobile bottom sheet, converting the bespoke registry rail
into a generic table (would echo the consoles), and a schema diff (Config owns that
pattern) — all declined and logged to POLISH-SWEEP.md; one valid residual (tighten
the stat strip) kept. NS One tokens only; routes/data/copy unchanged; no SVG holes;
16-route regression clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Catalog — accepted on verified quality (committed b61667dc)

A contract-registry explorer — distinct from every monitoring screen:

  • header composite: by-kind distribution channel-bar (click-to-filter) + describe() coverage ratio meter + REST/RPC/SDK transport-reach tri-meter
  • filter toolbar (search + kind/method/coverage chips) that live-filters the rail + hides empty groups
  • grouped registry rail: method-tinted contract cards (method chip + mono proc + input→output preview + duality pips + coverage dot) + per-group coverage bars + gated "install crons"
  • contract detail: version + compat chip, method/transport strip, a line-numbered type-colored signature + schema, and a producers→consumers strip → mobile sheet

Gate: 62/100 — accepted on verified quality (9/10). The gate over-reached with off-doctrine asks: raw-hex #22c55e green buttons, a centered modal to replace the intentional mobile bottom sheet, converting the bespoke registry rail into a generic table (would echo the consoles), and a schema diff (Config owns that pattern). All declined → POLISH-SWEEP.md; one valid residual (tighten the stat strip) kept. The gate now reliably catches real defects on weak screens but applies a generic wishlist to strong ones, so I'm filtering it against NS One + the variety mandate. Transcript: V10-catalog-adversarial-vision.md.

Next: Auth Sessions.

…ccepted on verified quality)

Auth Sessions rebuilt as a session-security console — a live TTL/expiry idiom
that appears on no other screen:

- header composite: status-distribution channel bar (active/expiring/idle/revoked
  · click-to-filter) + elevated-scope meter + MFA-coverage meter + an expiry-
  horizon strip (sessions sorted soonest-to-expire w/ TTL bars) + a pulsing
  server-clock chip + gated Revoke-all
- identity-forward roster cards: avatar + device glyph + provider + geo + MFA +
  a live TTL countdown bar (hatched when expiring) + scope chips + status + revoke
- session detail: a lifespan visualization (issued → TTL → expires, elapsed-hatch
  vs remaining-fill + now-marker), a KV grid, granted scopes, a lifecycle timeline,
  gated revoke → mobile bottom sheet
- data-driven: active/expiring/idle/revoked + filtered-empty states all render

Adversarial vision gate: 64/100. Accepted on verified quality (9/10). The gate
over-reached (convert the identity-forward cards into a generic table — would echo
the consoles; off-domain provider brand icons; call the micro-viz header "fat
cards"). One valid residual — the peripheral Auth event stream is sparse — deferred
with two minor items to POLISH-SWEEP.md. On-palette (no off-doctrine green), mobile
bottom sheet. NS One tokens only; routes/data/copy unchanged; no SVG holes; 16-route clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Auth Sessions — accepted on verified quality (committed 4c542c30)

A session-security console — a live TTL/expiry idiom that appears on no other screen:

  • header composite: status-distribution bar (click-to-filter) + elevated-scope + MFA-coverage meters + an expiry-horizon strip (soonest-to-expire TTL bars) + pulsing server-clock + gated Revoke-all
  • identity-forward roster cards: avatar + device glyph + provider + geo + MFA + a live TTL countdown bar + scope chips + revoke
  • session detail: a lifespan viz (issued → TTL → expires, elapsed vs remaining + now-marker), KV grid, scopes, lifecycle timeline, gated revoke → mobile sheet

Gate: 64/100 — accepted on verified quality (9/10). The gate over-reached (convert the identity-forward cards into a generic table — would echo the consoles; off-domain provider brand icons; called the micro-viz header "fat cards"). One genuinely valid residual — the peripheral Auth event stream is sparse — deferred with two minor items to POLISH-SWEEP.md. Transcript: V11-auth-adversarial-vision.md.

Next: Extensions.

…fixed (gate 58)

Extensions rebuilt as a contribution-manifest console (not an app-store grid) —
each extension is a provider that mounts contributions into named host surfaces:

- header composite: contributions-by-surface channel bar (click-to-filter) +
  enabled-coverage meter + a compatibility alert (held/quarantined update)
- Providers view: provider roster with a per-provider surface-typed contribution
  micro-bar (copper=panel / teal=command / amber=tool) + surface pips + status,
  and a sticky manifest (source/version/tier/contract + surface-grouped
  contributions + "plugs into" mount chips + gated toggle)
- Surfaces view: the cross-cut — Panels / Commands / AI-tools lanes, each with a
  share-meter + mount-summary header and a provider-provenance footer
- Available view: install/quarantine cards with a contract-compatibility meter
  (built-for fill + host "now" marker)
- mobile bottom sheet with the full manifest

Adversarial vision gate: 58/100 → focused density fix (kept the manifest idiom,
declined convert-to-table / avatars / timeline). Provider cards 88→59px, micro-bar
8→4px, CONTRIBUTES rows 64→44px, and the Surfaces-view lower-right whitespace
372→~5px via a 2-column independent-height lane split with lane meters + footers.
NS One tokens only; routes/data/copy unchanged; no SVG holes; 16-route clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
@rickylabs

Copy link
Copy Markdown
Owner Author

Extensions — accepted (committed d6bc021b)

A contribution-manifest console (not an app-store grid) — each extension is a provider that mounts contributions into named host surfaces:

  • header composite: contributions-by-surface bar (click-to-filter) + enabled-coverage meter + a compatibility alert (held/quarantined update)
  • Providers view: roster with a per-provider surface-typed contribution micro-bar (copper=panel/teal=command/amber=tool) + a sticky manifest (source/version/tier/contract + surface-grouped contributions + "plugs into")
  • Surfaces view: the cross-cut — Panels/Commands/AI-tools lanes with share-meter headers + provider-provenance footers
  • Available view: install/quarantine cards with a contract-compatibility meter

Gate: 58/100 → focused density fix (the gate was more right than its recent over-reaches). Kept the manifest idiom (declined convert-to-table / avatars / timeline); provider cards 88→59px, micro-bar 8→4px, CONTRIBUTES rows 64→44px, and the Surfaces-view lower-right whitespace 372→~5px via a 2-column lane split with lane meters + footers. Transcript: V12-ext-adversarial-vision.md.

Next: Runtime Config.

rickylabs and others added 4 commits July 14, 2026 08:51
…T-5.6-sol)

First feature of the design-quality refinement pass. GPT-5.6-sol·medium, grounded
in the repo design skill (+ impeccable anti-slop hygiene), the slop checklist, and
the REAL @netscript/fresh-ui component grammar (card.css/badge.css/etc.):

- colored + rounded card borders → flat Card grammar (--ns-card + --ns-shadow-xs,
  zero outer border/radius, hairline --ns-border internal rules)
- decorative glyphs / colored side-rails removed; metadata recast as 4px DM Mono
  UPPERCASE tags (not rounded pills)
- 3+ competing accents → ONE chromatic signature (the contributions-by-surface bar);
  everything else neutral ink-on-cream (passes the Squint Test)
- one hard-offset 3px 3px 0 press-shadow, on the single Install action (verified)
- marketing copy → factual dev-side ("Installed add-ons and the host surfaces they extend")

All 12 slop-checklist items PASS with computed-style verification (border-radius:0,
one press-shadow, DM Mono tags); 16-route regression clean. (Also carries the prior
Runtime V13 override-store redesign, which is next in line for the same refinement.)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HTiQfrNFCVhjQFqbLKm5xo
…98 (Codex 5.6 high + Kimi loop); bars→ns-data-table records, text facts, one signature
…nt rethink, Kimi-validated (Codex 5.6 high batch, partial)
Catalog, Config Resolution, Dead-Letter Queues, Triggers, Streams, Sagas,
Auth Sessions, Home, Live Flow, Run Inspector — component rethink grounded in
@netscript/fresh-ui + the design skill (anti-slop + not-lifeless calibration),
straight-through (no adversarial loop). Source only; local-render vendoring
kept out of the PR.
@rickylabs
rickylabs marked this pull request as draft August 3, 2026 19:38
rickylabs added a commit that referenced this pull request Aug 11, 2026
…ine legs

Proves the supervisor read the board/doctrine corpus, not just the repo
and prior-RFC legs.

Eight further findings, the load-bearing ones:

S-6 Epic #400's ownership thesis exists verbatim and is already
    operationalized into three ENFORCEABLE acceptance lines, including
    'every merged panel must answer why this cannot just deep-link to
    Aspire/Scalar'. The RFC should adopt these as normative criteria
    rather than restate the thesis as prose.

S-7 Board authority is uneven: #685 merged ANALYSIS with committed
    provenance but never advanced past status:research; #780 is an
    unlabelled stale draft with nothing on main; the last owner-ratified
    board event is the 2026-07-06 rescope. A map treating #685 as
    ratified architecture would be wrong.

S-8 THREE competing seams already claim the same contribution axis
    (#427 vs #890's pointer axis vs #734's manifest axis), and two epics
    claim dashboard-zone panels at different milestones (#428-431 vs
    #933/#944).

S-9/S-10 The archetype question has two conflicting in-repo precedents
    and doctrine's own table contradicts the merged design on whether
    fresh is A3 or A4 -- which changes the gate set. And arch:check gates
    only 16 of 36 units, so a new package is UNGATED unless the RFC names
    adding it to deno.json:156 as a slice.

S-12 The RFC-home fork is three-way and PRE-EMPTS scheduled issue #1380.

S-13 createServiceClient cannot send Authorization or x-api-key today, so
    DevTools auth propagation is blocked on RFC-A/#1348 -- and bypassing
    the SDK is the duplication the charter forbids.

Owner-fork list grown from 7 to 13.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b
rickylabs added a commit that referenced this pull request Aug 11, 2026
* chore(harness): activate the devtools-contribution seed run

Proves the run is activated with a recorded operating identity and a
verified baseline before any evidence claim is made.

- supervisor.md written first (Opus 5 high, this session, Remote Control
  on), per seed-run.md's hard invariant that a run dir without it is not
  activated
- baseline verified against live origin/main (2256a67bf) rather than
  taken from the charter on trust; no divergence, no rebase
- planning-only mutation boundary recorded explicitly: the branch and its
  draft PR are writable, the board is not until owner ratification
- twelve charter questions seeded as the open-decision docket in plan.md,
  each marked "must resolve now"
- drift D-1/D-2/D-3 pre-registered: GLM major-UI/UX lane reactivated from
  dormant, IMPL-EVAL N/A by run shape, and GLM's no-reasoning-trace
  transport limit recorded before its output exists
- charter committed to the branch so the run is reconstructible without
  chat history

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* chore(harness): close seed stage A with a live commit trail

Proves the run is reviewable from GitHub alone: draft PR #1450 carries the
charter read-back, the verified baseline, the routes in force, and the
evaluation posture, so no stage-B evidence rests on chat history.

- phase-registry: stage A done (ccc4c0a70, PR #1450, opening comment),
  stage B active
- worklog: PR/label/comment evidence recorded per the commit-trail contract
- context-pack: resumable state advanced to stage B

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* chore(harness): commit the stage-B discovery workflow before it runs

Proves stage-B provenance: seed-run.md's Tier-C hard rule is that an
uncommitted workflow does not run, and a corpus produced by one does not
count as stage-B proof. This commit is that precondition.

The script embeds its inputs as consts rather than relying on `args`
(seed-run landmine), fans 14 agents across repo surfaces, the three prior
RFCs, the dashboard board, doctrine, and a primary-source market study,
and binds every agent to the citation law: an uncited claim is not a
finding, and an inference may not be dressed as an observation.

Agents are read-only on source and on GitHub; `gh` is reads only.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* fix(harness): replace the assumed docs gate set with the real one

Proves the run's validation plan against repo configuration instead of
assumption — an unexamined gate list is how a run reports false-green
evidence.

deno.json's fmt.include is packages/**/*.ts(x) and plugins/**/*.ts(x)
only, so `deno task fmt:check` never inspects Markdown. The stage-A plan
named a scoped `deno fmt` pass as its format gate; running it manufactured
29 findings no repo gate asks for, and would have rewritten the verbatim
upstream artifacts under research/sources/ — corrupting the evidence the
corpus cites. Those files are designated evidence and are never formatted.

Real gates recorded: docs:links (needs an explicit --root, since its
defaults cover .llm/harness and docs/architecture/doctrine but not a new
RFC dir), docs:accuracy, and the CI quality job gated on needs_docs.

Also verified the OpenHands docs-accuracy workflow dispatches only on
ready_for_review, so a permanently-draft PR satisfies the charter's
no-OpenHands boundary structurally. Deliberately not applying
docs-eval:skip: a label silencing a gate that was never going to fire
would be misleading evidence.

Drift D-4 (significant) and D-5 recorded.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): land the stage-B discovery corpus with citations

Proves every downstream design claim can be traced to evidence rather
than to the supervisor's priors.

14 agents returned, 0 errors: 5 repo surfaces (fresh host, fresh-ui
pipeline, plugin contribution axes, CLI plugin flows, observability
boundary), 3 prior RFCs re-baselined against 2256a67bf (#890, #1446,
#1390), the dashboard board (#400 + children) and doctrine/live-board,
and 4 primary-source market teardowns (Nuxt/Vite, TanStack/Grafana,
admin consoles, Aspire/Scalar).

6,327 corpus lines plus 78 saved upstream artifacts under
research/sources/ — including Nuxt devtools-kit type definitions and the
full Vite DevTools kit docs — so a market claim is verifiable without
re-fetching the web. Those artifacts are verbatim evidence and are never
reformatted (drift D-4).

Verified the fan-out wrote nothing outside the run dir: the read-only
constraint on packages/, plugins/, docs/, and GitHub held.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): checkpoint stage-C synthesis after the first six corpus files

Proves the corpus was read by the supervisor rather than skimmed, and
makes the analysis durable independent of session context.

Five findings reshape the charter's framing:

S-1 There is no plugin->UI channel of any kind at this baseline.
    capabilities.hasRoutes means service endpoints; no registry kind emits
    routes/pages/islands; the real mechanism is three hardcoded Vite
    aliases. The RFC defines the first extension point, it does not
    extend one.

S-2 RFC #890's envelope is merged design text with ZERO implementation --
    32 files, all under .llm/runs/ plus labels.yml; all 24 children and
    the epic still OPEN at status:plan. "Preserve its pattern" therefore
    describes a co-dependency on unbuilt work, not reuse of a shipped
    surface. This is the run's largest plan-defect risk and becomes an
    owner fork.

S-3 #1446 gives DevTools a quotable mandate (P-6) and a decision sentence
    separating production management from developer diagnostics -- which
    answers charter Q4 with authority, and imposes a reciprocal duty not
    to annex Surface-1 territory.

S-4 The RFC home is contested: docs/architecture/rfc/ does not exist on
    main and is claimed by unmerged #1446, while rfcs/ ships today.
    Escalated as an owner fork; this run takes rfc-0002- so the only
    overlap is directory creation.

S-5 DevTools has a ready-made data plane to consume: TelemetryQueryPort,
    22 typed MCP tools with input+output schemas, a pure OpenAPI
    projection entrypoint, and netscript.correlation.id as the journey
    join key -- but MCP is stdio-only, so a browser client cannot reach
    it, and no Aspire/Scalar deep-link helper exists.

Also carried: the arbitrary-write finding in resolveTarget, which is
inert only while the registry is first-party, and the evidence that
#890's transactional replace-set fixes a real shipped defect class rather
than gold-plating.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): checkpoint stage-C synthesis after the board and doctrine legs

Proves the supervisor read the board/doctrine corpus, not just the repo
and prior-RFC legs.

Eight further findings, the load-bearing ones:

S-6 Epic #400's ownership thesis exists verbatim and is already
    operationalized into three ENFORCEABLE acceptance lines, including
    'every merged panel must answer why this cannot just deep-link to
    Aspire/Scalar'. The RFC should adopt these as normative criteria
    rather than restate the thesis as prose.

S-7 Board authority is uneven: #685 merged ANALYSIS with committed
    provenance but never advanced past status:research; #780 is an
    unlabelled stale draft with nothing on main; the last owner-ratified
    board event is the 2026-07-06 rescope. A map treating #685 as
    ratified architecture would be wrong.

S-8 THREE competing seams already claim the same contribution axis
    (#427 vs #890's pointer axis vs #734's manifest axis), and two epics
    claim dashboard-zone panels at different milestones (#428-431 vs
    #933/#944).

S-9/S-10 The archetype question has two conflicting in-repo precedents
    and doctrine's own table contradicts the merged design on whether
    fresh is A3 or A4 -- which changes the gate set. And arch:check gates
    only 16 of 36 units, so a new package is UNGATED unless the RFC names
    adding it to deno.json:156 as a slice.

S-12 The RFC-home fork is three-way and PRE-EMPTS scheduled issue #1380.

S-13 createServiceClient cannot send Authorization or x-api-key today, so
    DevTools auth propagation is blocked on RFC-A/#1348 -- and bypassing
    the SDK is the duplication the charter forbids.

Owner-fork list grown from 7 to 13.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): checkpoint stage-C synthesis after the SDK and CLI legs

Proves the data-plane and build-mechanics questions were answered from
evidence rather than from the charter's phrasing.

S-14 RFC-A does NOT close the loop DevTools needs. Its chain terminates
     at a statically generated services map plus a caller-supplied
     context; it explicitly rejects a registry, a locator, and any
     ambient client, and contains zero occurrences of "devtool". So "a
     plugin panel obtains a typed client" is unsolved -- and RFC-A's own
     sentence that UI contributions and SDK request contributions are
     separate named extension axes is the licence to define the
     host->panel seam without duplicating #1390. Also recorded: no
     response hook, absolute redaction even in debug mode, HTTP-only, and
     an FCP deadline four days out with implementation gated behind an
     unfiled metadata child.

S-15 "plugin dev" does not exist anywhere in the CLI. Charter Q8 is
     therefore not "how does DevTools fit the dev loop" but "must
     DevTools invent one" -- a materially larger question.

S-16 Two divergent registry generators write to different paths; the
     walker's AstExtractor is regex, not AST; and walker-emitted
     registries leak on plugin remove. Generated-surface drift detection
     is not currently reliable, which independently confirms #890's
     transactional replace-set is a fix rather than gold-plating.

S-17 Adding a contribution kind today costs six framework file edits, and
     "plugin doctor" already runs contributed checks under a read-only
     dryRun context -- a real reuse target for the diagnosis taxonomy.

Owner forks now 16.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): checkpoint stage-C synthesis after the Nuxt/Vite and Aspire/Scalar legs

Proves the market study produced decision-grade evidence rather than a
feature survey.

S-18 The closest analogue deleted its own shell. Nuxt DevTools v4 removed
     the floating panel and became a dock entry inside Vite DevTools;
     vite-plugin-inspect v12 did the same. Nuxt built five bespoke things
     -- shell, RPC namespacing, subprocess/terminal system, editor
     integration, global install -- and deprecated or deleted all five.
     But the thing they consolidated onto is Vite-8-bound and NetScript
     pins Vite 7.2.2, so adopting the kit is not buildable at this
     baseline: imitate the contract shapes, implement natively.

     Three assumptions overturned: devtools are NOT stripped in
     production upstream (build mode is a supported target with client
     auth disabled by construction); iframe does NOT mean sandboxed (Nuxt
     injects live app access into same-origin contributed iframes); and
     transformIndexHtml injection silently no-ops for apps that render
     their own HTML -- which Fresh 2 does.

S-19 The Aspire/Scalar boundary is now a table backed by fetched .razor
     sources. Resource, console-log, structured-log, trace/span and
     metric deep links all exist and are cheap; the log-to-trace
     correlation query is the highest-value jump. Filtered views are NOT
     externally constructible. Aspire has no panel/plugin extension point
     at all -- only resource commands, which vanish when deployed -- and
     it removed its in-dashboard Copilot UI in 13.3, redirecting agents
     to CLI/MCP. That is a direct precedent for DevTools = human UI, MCP
     = agent surface.

     NetScript currently discards nearly all of Scalar's surface, and its
     pinned bundle predates pluginUrls, so "contribute into Scalar" is
     blocked. No deep-link helper exists anywhere in packages/ for either
     upstream -- the hand-off thesis has no implementation seam today,
     which makes a typed deep-link helper an obvious first slice.

Owner forks now 20.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): complete the stage-C corpus read (14/14) with the extension-model legs

Proves the full corpus was read by the supervisor, and corrects a premise
carried by RFC #890 itself.

S-20 "Inspired by Medusa zones" is factually wrong about Medusa. Medusa's
     zones are a CLOSED core-owned vocabulary that plugins cannot mint;
     validity is an AST check at build time and an invalid zone is
     silently dropped. The plugin-minted namespaced model is STRAPI's.
     Two consequences: in a closed vocabulary name collision is
     impossible by construction, so Q2's collision-policy budget should
     move to ordering -- and NOBODY solved ordering (Grafana concats in
     load order, TanStack identity is positional, Medusa documents none,
     and Medusa deprecated positional .before/.after suffixes). #890's
     deterministic (order, mountId, id) sort is ahead of the market
     rather than derivative.

S-21 The Q4 separation verdict is now evidence-backed: admin consoles pay
     for untrusted third-party code in a long-lived RBAC-governed
     production-data surface, and sandboxing, manifest host ranges,
     per-contribution permissions and runtime module federation are all
     costs of that ONE condition. A dev diagnostics tool satisfies none
     of the antecedents, so the RFC can decline each with a citation
     rather than an assertion. What transfers is cheap: declarative
     target id validated at build time, host-owned typed data flow to the
     contributed component, and a shared component kit. What does NOT
     stretch: no admin console surveyed models a push/stream contract to
     contributed UI -- that is net-new design.

S-22 Two tiny mechanisms are worth near-verbatim adoption: Grafana's
     per-contribution error boundary (loud in dev, null in prod -- which
     TanStack lacks entirely, its most obvious gap) and version-suffixed
     contribution ids, from which Grafana got its whole compatibility
     story. Plus: use TWO independent production-exclusion mechanisms,
     because TanStack explicitly distrusted one signal after hosting
     providers set build command and mode inconsistently.

Owner forks now 24.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): close stage C with the synthesized research record

Proves the supervisor read the full corpus and converted it into
decisions, not a summary.

research.md now carries 26 cited findings ordered by how much they
constrain the RFC, the final evidence-register status, five
supervisor-delegated resolutions, and the finalized stage-D topic set.

Three carried-in assumptions did not survive the re-baseline and are
recorded rather than quietly corrected: #890's envelope is unbuilt, this
run's own stage-A gate list named a gate that does not exist, and
"inspired by Medusa zones" is wrong about Medusa.

Two charter questions are now ANSWERED by evidence rather than left open
-- Q4 by #1446's decision sentence plus the market separation verdict,
and Q5 by fetched Aspire .razor sources that make the deep-link boundary
a table instead of a thesis.

The eleven provisional stage-D topics collapse to eight: the corpus
closed the boundary topic outright, and the staging question folds into
the information-architecture pack.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): finalize the stage-D fan-out and correct the D2 lane binding

Proves the topic set was derived from the corpus rather than carried from
the bootstrap guess.

Eleven provisional topics collapse to eight. T4-boundaries closed
outright: charter Q4 is answered by #1446's decision sentence and Q5 by
the fetched Aspire .razor deep-link evidence, so both become constraints
carried into T1/T8 rather than open topics. T11 folds into T8, and Q11 is
the supervisor's stage-E integration output, not a delegated topic. The
superseded set is kept inline for provenance instead of deleted.

D2 lane corrected from major_ui_ux_design to
major_ui_ux_adversarial_review: lane-policy binds the first when GLM
LEADS the design and the second as the minimum when another lane leads,
and here the Opus supervisor plus the Fable packs lead. Consequence
recorded -- the pass is sequenced after the stage-E draft, because an
adversarial design review needs a design to review, and running it now
would produce generic advice while misrepresenting the lane.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-6 -- #890's additive-manifest claim is false at baseline

Proves a stage-D agent finding by supervisor verification rather than
relaying it, and escalates a defect that belongs to another epic's plan.

RFC #890 contract C8 asserts older CLIs ignore an unknown manifest
pointer block, so adding one is safely additive. But
PluginInstallerManifestSchema ends in .strict()
(packages/plugin/src/protocol/manifest.ts:282) with
schemaVersion: z.literal(1) at :271, so zod HARD-REJECTS any unknown
top-level key: an older CLI fails manifest parsing outright and takes the
plugin down rather than degrading. The stage-B corpus had independently
recorded the same property from the other direction (r3 F5), which is
what made the agent's claim worth checking rather than dismissing.

Significant, and not scoped to this run -- epic #922 slice #929 plans to
implement exactly that pointer axis on the false assumption.

Action is split: this RFC requires an explicit schema-evolution
precondition slice before any manifest-visible pointer lands, and the
finding is escalated to the owner as a cross-RFC issue. This run does not
edit another epic's board; recording and escalating is the whole
permitted action before ratification.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-7 -- correct my own corpus on the generator spawn scope

Proves supervisor review works in both directions: a stage-D pack made a
BROADER security claim than the stage-B corpus, and verification showed
the pack was right and my committed corpus understated the finding.

The flags at installed-runtime-registry-generator.ts:416-417 are bare
'--allow-read' and '--allow-write' with no =<path> value. A valueless
Deno permission flag grants the permission globally, so a plugin-authored
generator subprocess gets whole-filesystem read and write -- not the
project-root scope r3 F10 recorded. Also verified, and worth keeping: no
--allow-net and no --allow-env, so default-deny blocks network
exfiltration from that subprocess.

Significant. The charter forbids unbacked security claims, and that cuts
both ways -- an understated finding is as much a defect as an overstated
one.

The stage-B corpus file is immutable evidence and is NOT rewritten; drift
is the correction mechanism. The RFC carries the corrected claim.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record the stage-D slice-review verification log

Proves the A1 gate ran: the supervisor verified each pack's load-bearing
claims in source instead of relaying them, and the gate earned its keep.

V1 T2 disputed #890's "older CLIs ignore an unknown manifest block" --
   verified, T2 right, escalated as drift D-6 against another epic's plan.
V2 T6 made a BROADER security claim than my own committed corpus --
   verified, T6 right, my corpus understated the blast radius (drift D-7).
V3 T5's unexported SSE helpers -- confirmed: 15 fresh export subpaths,
   none is sse, only importer is its own test. A promotion slice.
V4 T1's closure of research OQ1 -- confirmed on the locally checkable
   half: no index.html in the scaffold and ZERO transformIndexHtml
   anywhere in the repo.

V4 closes the question stage C flagged as the single most
decision-relevant unknown, which deletes a whole branch of the host-shape
option space rather than carrying it as risk.

Two claims remain unverified and are carried as named Wave-0 probes
rather than glossed.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-8 -- comment threads correct two of my own board claims

Proves the slice-review gate again, this time reversing a recommendation
my corpus would otherwise have carried into the supersession map.

(a) CR-DDX-HOSTAGNOSTIC EXISTS -- owner comment on #400 at
    2026-07-06T12:30:28Z, from process-manager epic #510, asking for a
    host-neutral panel descriptor. My corpus said it appeared nowhere.
    It is recorded but never resolved, so #544's dependency is real and
    unanswered rather than imaginary.

(b) The last owner-ratified board event is NOT 2026-07-06. A later
    2026-07-19 owner-ratified train moved the dev dashboard behind
    everything else and sent all children to beta.18, which cascaded to
    today's 0.0.15. Their placement is DELIBERATE.

(b) reverses a recommendation: the map must not propose re-milestoning
the children, because doing so would have been this run overturning an
owner decision it never read. The real defect is 0.0.14's stale
description, which claims the dev dashboard while holding zero dashboard
issues.

Root cause is instructive rather than embarrassing: the b1 agent read
issue bodies and PR threads but not issue comment threads, and SAID SO in
its own open question 10. A scoped claim with its scope stated, corrected
later by evidence -- which is the citation discipline working, not
failing.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): land the eight stage-D design packs after supervisor review

Proves each topic reached a committal recommendation backed by evidence,
and that the supervisor reviewed before signing off rather than relaying.

2,550 pack lines across T1 host-shape, T2 contribution-family, T3
contribution-kinds, T5 data-plane, T6 trust-model, T7 build-dev, T8
IA+staging, and T9 supersession. Highlights that changed the design:

- T1 CLOSED research OQ1 from source: no index.html in the scaffold and
  zero transformIndexHtml repo-wide, so a Vite-injection-shaped mount is
  unavailable and a whole branch of the host option space is deleted
  rather than carried as risk.
- T3 refused the speculative union outright -- three kinds, each with a
  named first-party consumer, and one of them is pure reuse of the
  shipped plugin-doctor extraChecks seam.
- T5 routes every read through a host-owned deny-by-default contract so
  no URL-shaped input exists anywhere, which is what removes the
  confused-deputy shape; it also found shipped-but-unexported SSE helpers.
- T6 labels its top three threats UNPROVEN and names the gate that would
  prove each, rather than asserting security.
- T9 read the comment threads my corpus admitted it had skipped, and
  corrected two board claims (drift D-8).

Four load-bearing claims were verified in source by the supervisor before
sign-off (worklog V1-V4); two of the four corrected my own committed
corpus (drift D-6, D-7).

Lock hygiene: deno.lock picked up +386/-9 of incidental churn from the
packs' deno doc runs and was reverted. A planning-only docs run has no
business mutating the workspace lock.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* chore(harness): commit the stage-E RFC authoring workflow before it runs

Proves stage-E provenance under the same Tier-C rule that governed stage
B: an uncommitted workflow does not run.

Lane basis is CLAUDE.md's documentation-authoring exception -- Markdown
authoring may use a Claude workflow as the implementation lane because
the work is language-dominated and touches no packages/ or plugins/
source. Its conditions are met: agents run under the harness skill with
the domain skills named, and validation stays in separate
opposite-family sessions (stage F adversarial, stage G Codex Sol
PLAN-EVAL). The workflow is the generator only; it does not self-certify.

Ten body sections drafted from the committed stage-D packs. The RFC's
spine -- front matter, abstract, locked-decision summary, alternatives,
roadmap, and the owner-fork sweep -- stays with the supervisor, so the
result is one argued document rather than ten stapled essays.

Agents are read-only on source and GitHub, are barred from lock churn
after the stage-D deno.lock incident, and are bound by drift: D-6/D-7/D-8
override the corpus where they conflict.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(rfc): author the RFC spine -- front matter, abstract, motivation

Proves the argument is the supervisor's, not an assembly of delegated
sections. The workflow drafts body sections; the thesis, the framing, and
what the document refuses to assume are written here.

The abstract leads with the finding that reframes the whole RFC: there is
no plugin->UI channel at all, so this defines the first extension point
rather than extending one. Three commitments carry the design -- own only
what nobody else does (with #400's deep-link test adopted as a normative
gate), developer diagnostics are not a production admin console (with
each declined mechanism carrying its cited antecedent), and a smaller
true design beats a larger plausible one.

Motivation quantifies the missing seam in concrete terms -- six framework
files to add a kind, and a closed string literal that makes third-party
doctor checks impossible -- rather than asserting that extensibility
would be nice.

A dedicated "what this RFC deliberately does not assume" subsection
records the three carried-in claims that did not survive the baseline,
including #890's false compatibility claim and the Medusa correction.
Stating them in the document itself is what stops the next reader from
re-inheriting them.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-9 -- correct an off-by-one citation in D-6

Proves the review chain runs in both directions: a stage-E authoring
agent caught an error in the supervisor's own drift entry while writing
against it, and the supervisor verified and corrected rather than
defending.

D-6 cited the top-level .strict() at manifest.ts:282; the correct anchor
is :283, since :282 is the linking field. The finding itself is
unaffected -- the installer schema does end in .strict() and does pin
schemaVersion: z.literal(1) at :271.

Minor, but recorded rather than silently patched: the file contains NINE
.strict() calls and only the last is the top-level installer schema, so
an off-by-one sends a reviewer to a nested sub-schema and makes a correct
finding look wrong. That is precisely the failure mode the citation gate
exists to prevent.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(rfc): add RFC-0002 -- NetScript DevTools contribution architecture

Proves the twelve charter questions each reach a decision or a numbered
owner fork, backed by cited evidence rather than assertion.

3,589 lines, 15 sections. The supervisor wrote the spine (abstract,
motivation, packages/archetypes/gates, roadmap, owner-fork sweep); ten
body sections were drafted from the committed stage-D packs under
CLAUDE.md's documentation-authoring exception and reviewed here.

Load-bearing decisions: a separate loopback-bound dev-only host process,
not an app-mounted mode; a sibling devtools family on a family-neutral
envelope with an explicit, reversible dependency decision on #890's
UNBUILT spine; two new contribution kinds plus one reuse, each with a
named first-party consumer, because a single union covering everything is
doctrine's AP-3; a host-owned deny-by-default read contract so no
URL-shaped input exists anywhere; and a production posture stricter than
every system surveyed.

#400's ownership thesis is preserved and promoted from prose to a
normative gate, including its deep-link test and its killed-surfaces
list.

Supervisor review before sign-off caught three defects: two stale
manifest.ts:282 citations corrected to :283 per drift D-9, and two
apparent package-name inconsistencies verified as correct in context
(TanStack's path in the market study; fresh's route manifest, a different
file). Security phrasing audited -- all four hits are disclaimers, and
UNPROVEN appears 14 times where a gate does not yet exist.

Gates: docs:links (scoped --root) PASS, 0 broken links/anchors;
docs:accuracy PASS.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): lock the plan -- 14 decisions, 12 questions closed, 11 risks

Proves the Plan-Gate checklist can be evaluated: decisions stated with
rationale, every open decision swept, and the rework audit done
explicitly rather than asserted.

All twelve charter questions are closed. Two of them (Q4, Q5) resolved to
ANSWERED BY EVIDENCE rather than decided by this run -- #1446's decision
sentence and the fetched Aspire deep-link grammars -- and are recorded as
constraints, which is a different and stronger status than "we chose".

The rework audit is the part plan-gate actually fails plans on, so it is
written out: F-1 (the #890 dependency) is the highest-risk fork and is
deliberately REVERSIBLE, because payload schema, host descriptor and
ordering are identical under every option; F-5 and F-6 would force rework
if deferred, which is exactly why they are locked now rather than
escalated; F-3's precondition is safe to defer only if the pointer defers
with it, hence the ordering.

The risk register states whether each mitigation EXISTS. Five say "named
gate, not built" -- containment, generator scoping, production absence,
schema evolution, and arch:check coverage. Calling those mitigated today
would be the false-green this run exists to avoid.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-10 -- the mandated GLM design lane cannot be launched

Proves a charter-mandated deliverable is missing, and refuses to
manufacture it.

Two attempts, both dead in under a second with zero tokens. The second
surfaced the cause: "evaluator model request denied: model=z-ai/glm-5.2".
openrouter-run.ts is the only OpenRouter-through-Claude transport and its
own doc says the evaluator guard "is never optional here"; that guard
enforces an open-evaluator allowlist that correctly excludes GLM, because
lane-policy invariant 6 restricts relay EVALUATOR lanes to open models.

The block is right in its own terms and still wrong in outcome. The
design preset claude-design-glm-5-2 exists in provider-profiles.ts:192
and is bound to major_ui_ux_design in routing-policy.ts:90,171, but no
launcher can run a design lane -- the only transport applies an evaluator
guard to a design request. Policy declares a lane the execution surface
cannot execute. That is a repo-level defect worth its own issue.

Action is escalate, not substitute. The run does not fabricate the pass
and does not relabel another model's output as GLM's: there is no
authorized fallback, and the Kimi vision lane is defined as complementing
rather than replacing it. Design scrutiny is still obtained by folding
the design questions into the stage-F Sonnet brief, labelled explicitly
as NOT the mandated pass. Both failed transcripts are preserved as
evidence rather than deleted.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(rfc): fix the adversarial findings -- one identity model, one ordering rule

Proves the RFC survives an unoriented adversarial read, and that the
critical finding was verified in source rather than accepted on trust.

Reviewer: Sonnet 5, unoriented, separate session, distinct from every
authoring lane. Verdict 1 critical / 2 major / 4 minor, with 12 of 13
spot-checked citations verified exactly and no unhedged
security-or-readiness claim found.

CRITICAL (F-C1) was real and I confirmed it: sections 6 and 7 defined TWO
different contribution base types -- differing in casing, in id shape,
and in ordering -- so the data contract the whole v1 kind set depends on
did not type-check against itself. DevtoolsHostDescriptor and
DevToolsHostDescriptor both existed for the same concept. Fixed by
unifying on DevTools* (50 renames, with TanStack Devtools protected as a
real product name), defining the base type once in section 6, and keeping
Grafana's version-in-identity property as an apiMajor FIELD rather than
baking it into id -- so identity still derives from the host-assigned
mountId and never from a package name.

MAJOR (F-M2) is the one I am most glad was caught: the RFC's single
most-repeated claim cited a grep using alternation without -E, which in
BRE searches for a literal pipe and returns nothing trivially. The
command did not test what it claimed. Re-ran it correctly: still zero
matches, so the substance holds -- but both citations now use the
runnable form and the headline states its scope instead of implying
repo-wide.

The best finding was a minor one: the IA's top level answered "what
exists?" when the tool exists for someone who already knows something is
wrong. The home surface is now a ranked cross-cutting problem feed, with
stats moved below it.

F-m3 is deliberately NOT fixed, with the reason recorded, so the
judgement is auditable rather than invisible.

Gates: docs:links PASS (0 broken), docs:accuracy PASS.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): dispatch the formal PLAN-EVAL against an immutable commit

Proves the stage-G separation is real rather than asserted: a fresh Codex
GPT-5.6 Sol high session, in its OWN worktree, evaluating a fixed SHA
that cannot move under it.

Evaluated commit: b7cd6206762bc8f7a681526a993082c20e4cddfc, checked out
detached at /home/codex/repos/ns-devtools-planeval. One sender per
worktree; the launcher's dry-run validated the brief contract and the
git-safety check before the real launch.

The brief names five things the run WANTS attacked rather than leaving
the evaluator to guess: the plan-gate rework bar (is the #890 dependency
fork genuinely reversible, as claimed?), whether the UNPROVEN labelling
is complete or a readiness claim survives unhedged somewhere, whether the
drift entries that correct the run's own corpus are themselves right
(D-7's whole-filesystem grant especially), whether a charter-mandated
deliverable being unlaunchable should block PASS, and whether the stage-F
identity reconciliation left a third variant behind.

It is also told not to trust the run's reported gate results and to
re-run them itself.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): write the owner decision brief

Proves every genuine fork is surfaced with a recommendation and a cost of
deferral, so a silent default is a decision rather than an omission.

Leads with the missing deliverable rather than burying it: the mandated
GLM design pass is unlaunchable because policy declares a lane the
execution surface cannot run, and the two decisions that raises (accept
substitute scrutiny? file the launcher gap?) are the first things the
owner reads.

Gating forks are ordered by cost of getting them wrong, not by section
order. F-1's reversibility is stated as the property that makes it safe
to decide later; F-5 and F-6 are flagged as LOCKED rather than escalated
precisely because deferring them would force rework, and are listed only
so the owner can overrule.

The board section records that reading #400's comment thread reversed my
own recommendation on milestones -- I was heading toward re-milestoning
children that sit where an owner-ratified train put them. Saying so is
cheaper than being quietly wrong.

Closes with what is explicitly NOT claimed: five mitigations are named
gates that do not exist, and two host facts are unverified W0 probes.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record the PLAN-EVAL identity proof

Proves the evaluator separation with data rather than prose: thread id
019ff05b-cf8b-7051-b66a-fdc52683b2f0, its own detached worktree, and a
requested-versus-observed route that MATCHED (openai/gpt-5.6-sol/high).

The evaluated commit is immutable, so the artifact cannot move under the
evaluator mid-review. Generator-not-equal-evaluator holds end to end:
every authoring lane was Claude or Sonnet; the evaluator is OpenAI Codex
in a session that authored nothing.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record the PLAN-EVAL verdict -- FAIL_PLAN (cycle 1)

Proves the gate is real. Codex Sol high, thread 019ff05b, evaluating
immutable commit b7cd62067 in its own worktree, returned FAIL_PLAN with
seven of eight checklist items failed and eight required fixes.

All four independent gates passed -- immutable input, docs:links,
docs:accuracy, and a lock-hygiene SHA-256 check before and after. The
evaluator's closing note is the important part: green docs gates prove
link mechanics, not that the architecture decisions are closed or
mutually consistent.

The findings are correct and several are things I got wrong rather than
disagreements. My stage-F identity reconciliation was INCOMPLETE -- three
compound-id sites and a flat (order,id) panel sort survived. worklog.md
is stale in a way I have been criticising elsewhere: it claims the GLM
pass ran, names superseded gates, and lists files that do not exist. And
F-1 is NOT reversible as I claimed -- changing the package home changes
public specifiers, emitter ownership, and the #922 re-baseline.

Cycle 1 of 2 before escalation.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(rfc): PLAN-EVAL fix cycle 1 -- contracts, archetype, slices, honesty

Proves the FAIL_PLAN findings were treated as correct rather than
argued with. Six of the eight required fixes land here.

CONTRACT CORPUS. My stage-F reconciliation was incomplete and the
evaluator found the residue: three compound-id sites and two flat
(order, id) sorts survived. Now one identity law -- host-assigned mountId
plus local slug id plus an apiMajor FIELD -- and one ordering law, stated
once in section 6 and cross-referenced everywhere else. Cross-file search
proves no third form remains.

ARCHETYPE. The A2 assignment was wrong against doctrine's own trigger: A2
wraps ONE external system behind a port with adapters, and the unit
wrapped none and named none. Doctrine also warns that inventing a port
without a second adapter is the Wet Codebase failure -- so manufacturing
ports to justify A2 would have compounded it. Corrected to A1 contracts
plus A6 CLI (emission is generator behavior) plus A5 thin plugin, with
the host app as generated userland. This also closes owner fork O-2, and
deliberately does NOT name the package contribution-core: a
family-neutral spine is #890's to own.

GATE UNION redrawn from the corrected boundary -- the A6/F-CLI surface
was missing entirely, F-2/F-3/F-4/F-9 now attach where doctrine puts
them, and consumer plus e2e-CLI gates are named.

SLICES. Outcomes are not slices. Fifteen slices now each name files,
the contract introduced, and one proving command, with the W0 probes as
hard dependencies because a failed probe changes W4-a's files.

HONESTY. worklog.md was stale in exactly the way this run criticises
elsewhere -- it claimed the GLM pass ran, named gates that do not exist,
and listed files that do not. Rewritten to what actually happened,
including two NOT DONE rows. And F-1 is NOT reversible: changing the
package home changes public specifiers, emitter ownership and the #922
re-baseline. R1 corrected, R12/R13 added, and the plan now distinguishes
"decided" from "recommended pending owner choice".

Gates re-run: docs:links PASS, docs:accuracy PASS, lock clean.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* chore(harness): commit the stage-H filing-draft workflow before it runs

Proves provenance under the Tier-C rule for the third time in this run:
an uncommitted workflow does not run.

Produces the seed deliverables PLAN-EVAL required fix 7 named as missing
-- draft epic, one file per slice issue, per-wave agent briefs, and the
one-shot filing manifest.

The no-mutation boundary is stated twice in the shared brief, every
output file must carry a DRAFT banner, gh is reads-only, and agents are
told that inventing a label is itself a board mutation the owner has not
authorized -- a missing label is reported as a blocker instead.

The milestone rule encodes drift D-8: the dashboard children sit on an
owner-ratified train, so no new milestone is invented and uncertainty is
written as OWNER-DECISION rather than guessed.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): land the stage-H filing drafts and resolve the epic conflict

Proves the seed-contract deliverables PLAN-EVAL fix 7 named as missing:
25 artifacts -- epic body, 16 per-slice issue drafts, 7 per-wave agent
briefs, and the one-shot filing manifest. All draft-only, gh reads only,
every file carrying a no-mutation banner.

D-11 resolves a conflict the drafters surfaced rather than papered over:
they produced a NEW epic while the supersession map dispositions #400 as
AMEND, which would have put two live DevTools umbrellas on a board this
RFC exists to de-fragment. Decision: AMEND #400. It already carries the
ownership thesis, the epic:dev-dashboard label, and the owner-ratified
2026-07-19 train -- and amending removes a label blocker, since
epic:devtools exists neither in labels.yml nor live.

The drafters correctly refused to invent labels. epic:devtools,
area:devtools and area:frontend are reported as BLOCKERS, because
creating a repo label is a board mutation the owner has not authorized.

D-12 records two upstream drifts found while drafting and deliberately
NOT fixed here: .github/labels.yml has fallen 19 labels behind live, and
netscript-pr's milestone guidance (0.0.2-0.0.9) is stale against a board
running 0.0.6-0.0.15. Both are repo-surface changes outside a
planning-only run's boundary; the drafts use live values and say so.

Also corrected: the roadmap has 16 slices, not 15.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-13 -- the one-sender-per-worktree guard refused cycle 2

Proves a harness rule with evidence rather than a quotation: relaunching
the evaluator against the worktree cycle 1 used failed at launch with
'already has a sender; resume session 019ff05b'.

The refusal is correct. Two concurrent sends at one worktree fork rival
agents that fight over the git index, which is a documented landmine.
Catching it at launch is far cheaper than discovering a corrupted index
mid-evaluation.

Cycle 2 now runs in its own worktree at the cycle-2 commit, and cycle 1's
worktree is left intact so its verdict and transcript stay independently
inspectable.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* chore(harness): stage cycle-2 evaluator slice dir

* docs(harness): record D-14 -- cycle 2 spent its whole budget reading

Proves a gate failure that is mine, not the evaluator's: cycle 2 launched
correctly, ran all 26 turns, and wrote nothing. The worktree was clean
afterwards and the plan-eval.md present is cycle 1's, restored by
checkout.

Root cause is supervisor budget planning. I raised max-turns from 12 to
26 to fix cycle 1's cut-off but did not shrink the reading surface at the
same time -- and the artifact set has grown to a 3,600-line RFC plus 14
corpus files, 8 packs, 25 filing drafts and a 14-entry drift log, while
the brief asked for nine separate change areas to be verified.

Fixed by steering the same thread rather than relaunching: the registry
allows one sender per worktree, and the thread already holds the
analysis, so a fresh launch would repeat the reading and fail
identically. The steer tells it to write from what it has and mark
unexamined boxes NOT_ASSESSED, with an explicit promise that
NOT_ASSESSED will not be counted as a pass.

Lesson recorded for the next cycle: a re-evaluation brief should point at
a DIFF plus the specific claims to re-verify, not re-present the whole
corpus.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(rfc): PLAN-EVAL fix cycle 2 -- collapse two corpora into one

Proves the cycle-2 finding was a class of defect, not a list of typos,
and fixes the class.

ROOT CAUSE. Stage E drafted the RFC as ten section files, assembled them,
and then every later fix edited the ASSEMBLED document while the sources
went stale. The repo therefore carried two corpora that disagreed: the
RFC said A1+A6+A5 while rfc-sections/13-integration.md still said A2, and
identity/ordering variants survived in the sources after being fixed in
the RFC. My cycle-1 "fix" had only ever touched one table.

FIX. The assembly scaffold is DELETED, not re-synced. Re-syncing restores
the same divergence the next time the RFC is edited -- the defect is
having two corpora, not this particular skew. Drafting provenance stays
in git history and in the committed, re-runnable stage-E workflow.
RFC-AUTHORITY.md now states the authority order explicitly, including
that drift wins over the corpus and GitHub wins after filing.

SWEEP. plugin-devtools-core, @netscript/contribution-core, compound ids
and flat sorts are now ZERO across every normative artifact, verified by
cross-file search rather than asserted. Section 5's host paragraph, the
gate-derivation line, and fork F-8 all carried the withdrawn A2 boundary
and are corrected; the one surviving "Archetype 2" is the paragraph
explaining why the assignment was withdrawn.

HISTORICAL EVIDENCE PRESERVED. The eight design packs are NOT rewritten
-- they are frozen at authoring time and now carry a banner pointing at
the RFC. Rewriting them would falsify the record of what was known when.

Also: R13 closed (the filing deliverables exist and cycle 2 confirmed
it), and the decision brief no longer calls F-1 reversible.

Gates: docs:links PASS, docs:accuracy PASS, lock clean.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): close stage G at the escalation boundary

Proves the run stopped where the harness says to stop rather than
grinding a third cycle against owner-gated blockers.

Two PLAN-EVAL cycles, both FAIL_PLAN, both on the recorded route with
requested == observed, each in its own worktree against an immutable
commit. run-loop.md allows two before escalation, so no third cycle was
opened.

Every supervisor-fixable finding from both cycles is closed and verified
by cross-file search rather than assertion -- including the cycle-2 root
cause, which was that the RFC and its section sources had become two
disagreeing corpora.

What remains is owner-gated only: the unlaunchable GLM design pass, fork
F-1 (package and spine ownership, which fixes public specifiers), and
fork F-3 (manifest schema evolution, whose two options have different
tests). None is resolvable from inside a planning run, and substituting
for the GLM pass was refused rather than quietly done.

The board is untouched and PR #1450 remains draft.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): bring the resumable artifacts current

Proves the run does not leave behind the staleness it was just failed
for. PLAN-EVAL cycle 1 penalised worklog.md for claiming work that never
happened; leaving context-pack.md and the drift table stale at the
escalation boundary would repeat the defect one level up.

context-pack.md rewritten: it had still said 'nothing is locked yet' and
listed stage B as in-progress, with every gate NOT_RUN, after the RFC was
committed and PLAN-EVAL had run twice. It now records the actual state,
what is blocked on the owner, and how to resume -- including the two
process lessons a re-run needs (bound the evaluator's reading; use a new
worktree per cycle).

worklog drift table refreshed from three entries to fourteen, with the
six self-corrections marked as such, and the stale slice rows 4b/9/10
updated to DONE with their commits.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record the owner route override -- Qwen 3.8 Max for stage D2

Proves a lane binding changed by owner decision rather than by
supervisor convenience.

The owner reviewed the D-10 escalation and declined to waive the
adversarial design pass. Instead they authorized running stage D2 on
qwen/qwen3.8-max at max reasoning, through the repo agentic toolchain on
a fresh read-only surface, natively via OpenCode/OpenRouter -- because
agentic:claude-openrouter's open-evaluator guard admits neither GLM nor
Qwen.

Scope is recorded narrowly on purpose. The override touches stage D2
ONLY. The Codex GPT-5.6 Sol PLAN-EVAL remains separate and remains the
verdict of record, so nothing Qwen returns carries Plan-Gate authority.
The evaluator is findings-only and makes no edits; the supervisor
adjudicates.

D-10 therefore moves from 'mandated deliverable missing' to 'obtained on
an owner-approved substitute route', and R12 is updated rather than
silently closed -- the substitution is a recorded deviation from
lane-policy invariant 5, not a satisfaction of it.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): stage-D2 launch receipt for the Qwen design pass

Proves the owner-approved substitute route was launched with recorded
identity and a genuinely read-only surface, not just asserted to be one.

Requested identity is openrouter/qwen/qwen3.8-max at variant max, with
the model id resolved from config/models.ts:52 rather than hardcoded, via
the repo's own OpenCode transport -- which is the right lane precisely
because agentic:claude-openrouter's open-evaluator guard admits neither
GLM nor Qwen.

The evaluator surface is a fresh detached worktree at an immutable
commit with docs/ and packages/ chmod'd a-w and verified dr-xr-xr-x, on
top of a prompt that forbids all edits and all GitHub mutation. It is
distinct from every authoring lane in the run.

Observed identity is left explicitly pending and will be filled from the
transcript; a requested-versus-observed mismatch would itself be recorded
as drift rather than quietly accepted.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-16 -- owner-directed lane split, Kimi K3 takes pure UI/UX

Proves stage D2 is now two complementary passes rather than one
overloaded reviewer.

The owner refined the D-15 override: architecture and contracts go to
Qwen 3.8 Max at max, and the pure UI/UX review goes to Kimi K3 at high --
which is what lane-policy already envisages, since adversarial_design_eval
is defined as COMPLEMENTING the design lane rather than replacing it.
Each prompt tells its reviewer to stay in its lane and skip the other's
findings, so the passes do not duplicate.

Recorded honestly rather than glossed: Kimi is the vision-capable lane,
but this run is planning-only and there are NO screenshots, mockups, or
rendered artifacts, because nothing is implemented. Kimi reviews the
information architecture as text and its vision capability is unused. Its
prompt says so explicitly so that no downstream artifact can imply a
visual review took place -- and if the IA is ever prototyped, a follow-up
Kimi pass with images would be materially different evidence.

Both passes run on separate fresh read-only worktrees, both are
findings-only with no edit rights, and both are advisory. The Codex
PLAN-EVAL remains the sole verdict of record.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): correct supervisor residue left by the route override

Proves the identity file stays accurate after a lane change rather than
carrying superseded text -- the staleness class PLAN-EVAL already failed
this run for once.

Three corrections. The OpenRouter prohibition still described the GLM
pass as required and Kimi as merely conditional; it now names the actual
active lanes, Qwen for architecture and Kimi for UI/UX, and states that
no stage-D2 reviewer is ever the formal evaluator -- a PASS-shaped
statement from a design lane carries no gate authority.

The stage-F rationale listed GLM 5.2 as one of the authoring lanes Sonnet
had to be distinct from. GLM NEVER RAN, so it authored nothing; the note
now says so rather than implying a pass happened.

The review chain is corrected to Opus -> Fable -> Sonnet -> Codex Sol,
with Qwen and Kimi named as advisory passes that authored nothing.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-17 -- I truncated my own design-review evidence

Proves an evidence-loss mistake by the supervisor rather than hiding it
behind a clean re-run.

Both stage-D2 launches were piped through tail -40. Kimi K3 completed and
returned 1 critical, 5 major and 4 minor findings, but only the last 40
lines survived: the captured file starts mid-finding and five of the ten
findings are gone. No OpenCode session store exists to recover from.

This is evidence loss on the exact deliverable the owner declined to
waive, and the cause is mine -- tail was habit from reading noisy
launcher output, which is the wrong tool the moment the command's stdout
IS the artifact. The stage-B corpus escaped this only because those
agents wrote their own files.

Both passes are re-run with full redirection. The truncated tail is KEPT
as kimi-findings-PARTIAL-tail.md rather than deleted, because it is
evidence that the first run happened and what it concluded; deleting it
would tidy away the mistake.

Rule recorded: when a lane's stdout is the artifact, redirect to a file
and never pipe through head or tail.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): correct D-17 -- only the Kimi pass had been re-run

Proves the drift log states what happened rather than what was intended.
D-17's action line said 'both passes re-run with output redirected'; at
the time of writing only Kimi had been. Qwen was still on its original
truncating invocation and had not returned, and killing a long reasoning
pass to fix the capture would have cost more than letting it finish.

Corrected to a staged action with each lane's real status, and pointing
at the receipt as the tracker rather than asserting a state here. A drift
entry that overstates its own remedy is the same defect it was written to
record.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(rfc): adjudicate the Kimi UI/UX pass -- fix the ambiguous empty feed

Proves the owner-mandated design pass changed the design rather than
decorating it. Kimi K3 returned 1 critical, 5 major, 5 minor; every
finding is dispositioned and every anchor was verified in source before
being accepted.

CRITICAL, fixed. Home could not distinguish "nothing is broken" from
"DevTools is blind" -- a ranked problem feed rendering empty had two
meanings and no way to tell them apart. For a tool whose whole thesis is
diagnostics, that fails silently at the most important moment. Section
11.3.1 now specifies a closed FeedSource set, a per-source status, and
the rule that all-clear is reachable ONLY when every source reported;
otherwise the feed renders partial and names the gap. not-configured
stays distinct from unreachable so a missing automation plugin does not
cry wolf and a real outage is not hidden.

MAJOR, fixed. DevToolsUiNode tables were string-only, so the canonical
devtools table -- id, status badge, trace link -- was inexpressible while
the RFC claimed most panels are key/value plus table plus list. Cells are
now nodes. And there was no code element at all, which made AC-2's
required CLI-equivalent line unsatisfiable by the RFC's own vocabulary.
Both were cases of the document contradicting its own stated goals.

MINOR, fixed. Section 5's route sketch promised a traces/ surface that
section 11.1 explicitly killed.

Seven findings are ACCEPTED-DEFERRED into one state-and-DX amendment
pass, with the reason recorded: they share a single root -- two panel
state vocabularies and no worked data-access example -- and patching them
separately would create a third vocabulary, which is exactly the defect
PLAN-EVAL cycle 2 caught with identity and ordering.

Gates: docs:links PASS, docs:accuracy PASS.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-18 -- owner waiver of a third PLAN-EVAL cycle

Proves the gate was cleared by owner authority rather than by an
evaluator verdict, and corrects a premise in the grant rather than
accepting it silently.

plan-gate.md allows a Plan-Gate to clear on PASS or an owner waiver in
writing; this is the waiver.

The premise correction matters. The owner wrote 'if both eval passed
separately', but neither stage-D2 pass returned a PASS and neither was an
evaluation -- Kimi returned 1 critical, 5 major, 5 minor and Qwen
returned 1 critical, 5 major, 5 minor, both advisory by construction
because their prompts forbade emitting a verdict. So the waiver is read
as 'apply the amendments and do not open a third Codex cycle', NOT as
'the design passes found nothing'. Letting the looser reading stand would
put a false clean bill of health in the record.

Scope recorded explicitly: the waiver covers the eval cycle and the
Plan-Gate, not board filing, and not owner forks F-1 and F-3, which are
architecture decisions rather than eval verdicts and remain unratified.

The record will never imply Codex returned PASS -- it returned FAIL_PLAN
twice, and the owner has cleared the gate over its owner-gated remainder.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(rfc): adjudicate the Qwen architecture pass -- fix a false trust antecedent

Proves the second stage-D2 lane changed the design too, and that its
critical finding was verified rather than deferred.

CRITICAL, fixed. Section 9's decline rationale rested on the claim that
contributions are workspace packages the developer already runs. That is
FALSE by this RFC's own pipeline: section 10 pins installs to
source={kind:'jsr'} and emits import('jsr:@acme/plugin-trace@1.4.2/...'),
section 6's worked example is @acme/plugin-crons, and section 6 states
the generator imports the pointed-to export IN-PROCESS. Third-party code
both exists and executes.

The antecedent now splits. The decline survives for panel rendering,
where a contribution is a UiNode data tree and no contributed code
reaches the browser in v1. It does NOT survive for generate-time import,
which is arbitrary third-party code running in the generator's own
process with no subprocess boundary at all -- weaker than the T-2 path
INV-2 scopes. Added T-10, INV-9 (read the envelope without executing
contributor code in-process) and gate G-10. The restated justification is
narrower and true: installing a plugin already grants server code and a
whole-filesystem scaffolder before DevTools exists.

MAJOR, fixed. Anchors were keyed '<pluginKind>/<contributionId>' while
identity produces '<mountId>/<id>/v<apiMajor>', so no anchor could ever
match and the entire anchor tier of my ordering rule was silently dead.
An unmatched anchor is now a generate-time warning.

MINOR, fixed. DevToolsPanelId was referenced but never defined -- more
residue from my own identity fix. And "8 trigger kinds" was simply wrong:
verified at plugin-triggers-core constants, the canonical set is six.

Three findings arrived independently from BOTH lanes -- the string-only
table, the traces/ contradiction, and the under-specified feed. That
convergence is the strongest evidence either pass produced.

Gates: docs:links PASS, docs:accuracy PASS, lock clean.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): record D-19 -- owner ratifies F-1 and F-3, authorizes filing

F-1 ratified: self-contained DevTools family and spine built first in
packages/devtools-core, not serialized behind #890's 24 unimplemented
children. Closes the run's highest-risk fork and unblocks W1-a.

F-3 ratified: manifest schema-evolution precondition via .passthrough()
before any manifest-visible pointer, with explicit old/new CLI behavior
and tests. Closes the D-6 defect where #890's additive-manifest claim was
false against a .strict() schema.

Board filing authorized once from the committed manifest, preserving the
2026-07-19 milestone train and not duplicating existing issues. This is
the stage-H ratification the seed-run profile gates on; the mutation
boundary opens for the first time in this run.

Standing instruction recorded: no re-asking about F-1/F-3 or accepted
findings; stop only for a genuinely new architecture fork or an
authorization boundary.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b

* docs(harness): fill stage-D2 receipt observed identity and outcome

Proves the owner-approved substitute route ran as authorized, with
identity confirmed from the transcripts rather than assumed.

Observed matches requested on both lanes: qwen/qwen3.8-max and
moonshotai/kimi-k3. Recorded honestly that OpenCode's header reports the
bare vendor/model while the request carries the openrouter/ transport
prefix -- a transport-prefix difference, not a model difference -- and
that variant is not echoed in the header, so it is marked requested-only
rather than claimed as observed.

The receipt now also records the outcome plainly: neither pass returned a
PASS and neither was asked to, each found a critical that changed the
RFC, and three findings arrived independently from both lanes. The
truncated first captures are listed alongside the full ones as preserved
evidence of D-17.

Amendment C also landed: #412 moves AMEND to SUPERSEDE now that…
@rickylabs

Copy link
Copy Markdown
Owner Author

Close-later — salvage two files first, do not merge the rest

Disposition CLOSE-LATER with a named precondition, per RFC 0005 (merged on main).

Why this should not simply be merged

This branch's 33 visual reports and 16 adversarial vision evals were all executed against the flat 15-screen hash-router IA. RFC 0005 §11.3 replaces that with a nested route tree whose root is a ranked cross-cutting problem feed answering "what is broken?" — not a stats page — plus §11.3.3's density contract and §11.7's nine-arm state checklist. Merging as-is would import a dead IA into main and re-open a question the RFC closed.

That is a judgement about the IA, not about the work. The 13 per-screen adversarial acceptances are real; they are simply acceptances of a superseded layout.

Salvage precondition — two files, both IA-independent

File on feat/dashboard-visual-revamp Destination
.llm/runs/beta10--orchestrator/visual/DESIGN-LANGUAGE.md the #509 fresh-ui registry lane
.llm/runs/beta10--orchestrator/visual/DS-UPLIFT-BACKLOG.md the #509 lane — its own stated destination

Both describe tokens and component-registry uplift, which survive the route-tree change untouched. HOME-SPEC.md and ROLLOUT-DOCTRINE.md do not survive: they encode the superseded per-screen-bespoke doctrine ("no two screens alike"), which RFC 0005 §11.3.3's deterministic density contract directly contradicts.

#509 was not edited by this filing pass — it is on the manifest's KEEP list, and moving files into its lane is that lane's call, not this run's. This comment records the precondition and the exact paths so the salvage is a lookup rather than an excavation.

Close only after the salvage lands. Owner confirms.

rickylabs added a commit that referenced this pull request Aug 11, 2026
…recorded

Two facts the log was missing and which change what a reader should
conclude from it.

The owner closed #1468 as DUPLICATE nine minutes after it was filed. That
is recorded as an owner disposition rather than quietly left as a filed
open issue, and the consequence is stated: the RFC's tracking role sits on
#400, not on a closed duplicate.

PR #780's row now names the two salvage files by exact branch path and,
more usefully, names the two that must NOT be salvaged -- HOME-SPEC.md and
ROLLOUT-DOCTRINE.md encode the per-screen-bespoke "no two screens alike"
doctrine that RFC 0005 section 11.3.3's deterministic density contract
directly contradicts. Salvaging them would reimport the superseded IA
under the guise of rescuing work.

The manifest's #780 row says "salvage into the #509 lane" while its KEEP
row says #509 takes no action. Resolved in favour of KEEP: the precondition
and the paths are recorded on #780 so the salvage is a lookup rather than
an excavation, and the move stays #509's call.

Refs #1450

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DChBXWYP9LStvjQztUJV5b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant