fix(tmux-restore): timestamped plan generations β a bad save no longer destroys the claude session bindings - #1383
Conversation
β¦longer destroys the bindings
MEASURED data loss, workbench 2026-09-06:
21:47:17 a good plan β 47 entries, 46 carrying a bound session id
21:54:21 the tmux server died, taking 47 claude conversations
22:09:40 a continuum autosave fired on the DEGRADED post-crash workspace
and `cmd_save` overwrote ~/.config/initiatives/restore-plan.json
with 10 entries. The cheat-sheet went in the same second.
NO BACKUP EXISTED.
The conversations were recovered only because tmux-resurrect keeps its saves
TIMESTAMPED (`tmux_resurrect_<ts>.txt` + a `last` symlink) and each pane line
happens to carry a full `claude --resume <id>`; 24 were rebuilt from that file.
THAT ASYMMETRY WAS THE BUG: a bad save cost the LAYOUT nothing and the BINDINGS
everything.
So mirror what resurrect already does. Each `save` writes an immutable
generation into `restore-plans/restore-plan_<ts>.json` (+ the cheat-sheet, which
had the identical defect from the identical writer) and repoints
`restore-plan.json` / `restore-cheatsheet.md` at it as relative symlinks. Every
existing reader β `cmd_restore`, `plan_staleness_hours`, `cmd_show`, `--plan`,
`tmux-restore-observe.sh` β goes on reading the same two paths; `stat()` and
`Path.exists()` follow symlinks, so `basis=layout` and the observe script's
mtime reads are unchanged.
What this deliberately does NOT do: refuse a shrinking save. The operator
closing windows is ordinary use, so a guard on a falling entry count would fire
on ordinary use and train everyone to bypass it. The degraded save is still
written β it is simply no longer the only copy. What a shrink gets is a WARNING
keyed on BOUND SESSION IDS (not the entry count: dropping five unbound windows
loses nothing resumable and must stay quiet) naming the previous generation and
the exact `restore --plan <path>` that recovers it.
RETENTION: 192 generations, a COUNT rather than an age. The writer is hook-
driven, so an age bound gives no bound on disk at all; a count bounds disk
whatever the cadence does, at the price of a cadence-dependent span, which is
stated rather than hidden. 192 is 48h at continuum's 15-min interval β long
enough to outlast a crash noticed the following evening (~20h) or over a
weekend (~40h). Cost: the live 10-entry plan measures 3,820 B + 3,267 B, so a
47-entry generation pair is ~33 KB and 192 of them ~6.3 MB.
Also fixed, same writer, same shape:
* `_write_atomic` β `write_text` truncates first, so a save killed mid-write
left a zero-byte plan. Temp file + `os.replace`.
* `free_generation_stamp` β the stamp has one-second resolution and the
incident was a same-second write. A manual `save` racing the 15-min hook
would have shared a stamp and clobbered the previous generation from inside
the mechanism built to stop that. It must also be strictly NEWER than
everything present, not merely free: a first draft returned the slot pruning
had just freed, the new generation sorted OLDEST, and the run reported
`4 kept (max 3), 0 pruned` β the retention cap silently unenforced. Caught
by its own test; a backwards clock reproduces it with no race.
* `adopt_pre_generation_files` β the DEPLOY of this change must not itself be
the bad save. A host still on the old writer has a regular file at the
pointer path holding possibly the only good plan; it is copied in as a
generation before the pointer moves.
* `prune_generations(protect=β¦)` β pruning is the only code here that deletes,
so it is the only code that can recreate the defect. It refuses to unlink
the generation the pointer was just aimed at, whatever the ordering
argument does.
TESTS β 13 new, red/green matrix measured at base c5e425c:
RED at base on the DATA LOSS (not on an AttributeError; the assertions use the
test module's own `_bound_ids` and walk the state dir implementation-blind, so
they measure "is the binding still on disk"):
a_shrinking_save_does_not_destroy_the_previous_bindings 46 of 46 ids lost
a_shrinking_save_does_not_destroy_the_previous_cheat_sheet no resume cmd left
a_pre_generations_plan_is_preserved_by_the_first_new_save 46 of 46 ids lost
a_save_dropping_bound_ids_names_the_recovery_command no report at all
plus 8 more red at base, all 13 green at HEAD (102 in the file).
GREEN AT BASE BY DESIGN, labelled as invariant guards not regression coverage:
a_shrinking_save_is_written_not_refused (the constraint on the fix)
a_shrink_that_drops_no_bindings_stays_quiet (control: ids, not counts)
the_pointer_still_reads_as_the_current_plan (reader compatibility)
MUTATION SWEEP, 13 mutants under PYTHONDONTWRITEBYTECODE=1, fresh tree each,
every patch verified to have applied (a `str.replace` matching nothing scores
SURVIVED), each kill confirmed to carry THAT guard's own assertion text:
12 KILLED β no-adoption, protect-ignored, prune-wrong-end, prune-never,
report-on-count, report-silent, stamp-free-not-newest, raw-second-stamp,
non-atomic-write, empty-plan-spends-a-slot, generations-dir-hardcoded, and
the full revert to the in-place overwrite.
1 SURVIVED and is explained, not waved away: removing ONLY the generation
writes leaves adoption running, and adoption alone then provides a one-deep
backup β the mutant fails to break the property rather than the test failing
to see it. Reverting both together (M1b) kills both regression tests with
their own messages, and the base-tree measurement above is the definitive
form of that same mutant.
Positive control for retention: 5 saves at keep=3 leave exactly the newest 3,
with `0 pruned` asserted below the cap and `1 pruned` on the run that first
exceeds it β a pruner wired to nothing cannot pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019V1ysK1JBgdNyfqAt6gDKC
Claude-Session-Id: 5542cd95-4967-4463-8fe2-0f0a75194e9d
β¦and the plan is now a symlink
Found by verifying the compatibility claim instead of asserting it. The previous
commit's message said the observe script's mtime reads were "unchanged" because
`stat()` follows symlinks. That is true of Python's `Path.stat()` β so
`plan_staleness_hours` really was unaffected β and FALSE of GNU `stat(1)`, which
uses **lstat** by default.
MEASURED on a link whose target was stamped 12:00:00:
stat -c '%y' <link> -> 2026-09-07 22:41:47 (the LINK was repointed then)
stat -Lc '%y' <link> -> 2026-09-07 12:00:00 (the PLAN was written then)
Both are timestamps, both render plausibly, and nothing in the output says which
one you are looking at β so `plan_mtime=` would have silently become "when the
pointer moved" and `plan_layout_skew_seconds` would have compared the pointer's
repoint time against the layout's write time: two writers that are not the two
that line claims to compare. In practice the link is repointed in the same
second the generation is written, so the numbers would usually have looked
right, which is what makes it worth pinning rather than shrugging at.
Both `$PLAN` sites now dereference. The layout sites are deliberately left
alone: `replayed_layout`/`newest_layout` glob `tmux_resurrect_*.txt`, which are
always real files β `<resurrect-dir>/last` is a symlink but is never what those
reach. `[ -f "$PLAN" ]` needs no flag either; `test` dereferences already.
TEST: test_the_plan_mtime_is_the_PLANS_not_the_symlinks β a real generations
dir plus a pointer into it, layout mtime 1 and generation mtime 1001.
RED at base c5e425c on the skew: "the plan/layout skew was not computed from
the plan GENERATION's mtime (1001) against the layout's (1) β `stat` read the
symlink". Green at HEAD (44 in that file).
The skew is asserted as the exact number 1000 rather than as "present" because
the two candidate answers differ by ~1.8e9 seconds β arithmetic separates them
and a presence check cannot. The `plan_mtime=` expectation is DERIVED by running
`stat -c %y` on the generation rather than written as a literal: epoch 1001
renders as 1969-12-31 or 1970-01-01 depending on the host offset, and the first
draft asserted `1970-` and failed at -0600. A test that passes in one timezone
is the config-blind suite, not coverage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019V1ysK1JBgdNyfqAt6gDKC
Claude-Session-Id: 5542cd95-4967-4463-8fe2-0f0a75194e9d
|
Retention arithmetic upgraded from extrapolation to measurement.
The operator's workspace has since recovered, and the live files now measure: 18,953 B + 15,131 B = 34,084 B per generation pair against the ~33 KB the source predicts β so 192 generations is ~6.5 MB, and the number in (Read-only |
β¦nd 46 of 47 conversations came back A subagent dispatched by this session ran `TMUX_TMPDIR=... tmux kill-server` at 21:54:15.802 CDT against the operator's live server. 42 tmux-spawn scopes tore down 1.2s later, destroying 47 live claude conversations. The OOM hypothesis I reported earlier is REFUTED, with controls: no kernel oom-kill for that boot (positive control β the same grep matches real OOM kills from 2026-08-28), systemd-oomd not installed, no tmux core, and the alarming 57.7G/38.2G peaks were per-scope LIFETIME high-water marks printed at teardown over 7-23h wall clocks, not concurrent usage. Recovery method, recorded because it is reusable: tmux-resurrect's window lines carry a layout string whose every cell ends in that pane's numeric id, and the agent-ledger keys one record per pane on that id carrying the session id. Join on the pane id. It carried its own positive control -- 24 windows had independently recorded their own `claude --resume <id>`, and the join reproduced 21 of 21 where both sources spoke, zero disagreements. The rest were matched from restored pane scrollback via find-session, each disambiguated by a string unique to one candidate with a positive control that the term existed at all. Also records: PRs #1376 and #1383 gated TOGETHER on an integration branch (both touch tmux-session-restore.py, and a clean git merge is not a clean merge), and BLOCKED by a client-subdomain leak on main that neither PR caused. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ApQvw3A9KbFEUXVjtTAk4j Claude-Session-Id: 5542cd95-4967-4463-8fe2-0f0a75194e9d
β¦-generations Claude-Session-Id: 097b404c-db17-4472-bd37-dc90cf8fa675
β¦-back could overwrite a generation Round-1 audit findings, all four, in one commit. The π΄ is a data loss inside the mechanism this PR exists to build. π΄ 1 β THE STAMP IS NOW UTC. `list_generations` sorts on the NAME and pruning deletes from the older end, so every ordering guarantee here rests on the stamp being monotonic. Local time is not: it repeats an hour at every DST fall-back. Measured on this host's zone (America/Winnipeg), the two instants 2026-11-01 06:00 and 07:00 UTC are BOTH `20261101T010000` local β one stamp, two saves. The anchor then parsed that ambiguous string with `time.mktime`, `stamp > newest` could never be satisfied, and the fall-through returned an OCCUPIED stamp: a bound session id and its cheat-sheet destroyed at rc 0 with nothing printed, while `prune_generations` reported `0 pruned`. The docstring called that "a bounded, VISIBLE loss"; it was not visible by any means. `generation_stamp` uses `time.gmtime`, and the anchor uses `calendar.timegm` to match. The resurrect-style format is unchanged β only the clock. Nothing parses these names as local time. π΄ 1b β AN EXHAUSTED SEARCH NOW RAISES instead of returning an occupied stamp. A save that fails loudly costs one save; a save that clobbers costs the bound plan it existed to protect. π‘ 2 β THE CONCURRENT-SAVE RACE THE DOCSTRING NAMED IS CLOSED. `scripts/tmux-post-save.sh:21` backgrounds and disowns `save` with no lock, so a manual save genuinely races the 15-minute hook. What was implemented covered the SEQUENTIAL same-second case only. Two processes in the same second both saw the slot free via `exists()`, both took the stamp, then interleaved over FIXED temp names β measured: `FileNotFoundError` out of `os.replace`, a generation holding one process's bytes under the other's rename, and `FileExistsError` out of `os.symlink`. Now: the free-check and the claim are one `O_CREAT|O_EXCL` operation, and `_write_atomic`/`_point_at` use per-process temp names. `_write_atomic` also removes its own temp on failure β a leftover `.tmp` does not match `_GEN_PLAN_RE`, so pruning would never reap it. π‘ 3 β `prune_generations` REPORTS ONLY WHAT IT DELETED. The unlink `OSError` was swallowed and the stamp appended regardless. Measured: `2 kept (max 1), 1 pruned` while NOTHING had been pruned. A persistent unlink failure gives unbounded growth reported as healthy retention on every save. π‘ 4 β THE SHRINK REPORT CANNOT NAME A FILE THE SAME SAVE PRUNED. `protect=(stamp,)` did not cover `previous_gen` β the file the recovery command names. Measured at KEEP_GENERATIONS=1: the report named a path whose `exists()` was False. This file's own rule is that a warning pointing at the wrong file is worse than none. RED AT BASE / GREEN AT HEAD β the matrix, per test Base = this branch merged with main (d732106), source reverted, tests kept. RED at base, green at HEAD (regression coverage): test_a_generation_stamp_is_monotonic_across_a_DST_fall_back[America/Winnipeg] test_an_exhausted_stamp_search_REFUSES_instead_of_overwriting test_a_concurrent_save_cannot_take_a_stamp_another_save_claimed test_write_atomic_temp_names_are_per_process test_prune_reports_only_what_it_actually_deleted test_the_shrink_report_never_names_a_file_this_save_just_pruned GREEN at base β labelled INVARIANT GUARDS in their own docstrings, not counted as regression coverage: test_a_generation_stamp_is_monotonic_across_a_DST_fall_back[UTC] the non-vacuity control: in a zone without DST there is no collision, so this is what proves the Winnipeg parametrisation is doing the catching. test_a_save_inside_a_repeated_local_hour_does_not_destroy_a_generation[both] the end-to-end loss is NOT deterministic from a test: glibc's `mktime` tie-break for a repeated hour is unspecified, and in this harness it resolves so the base anchor steps forward and no loss occurs. The audit measured it going both ways. Kept because it pins the property on the real `cmd_save` path; labelled so nobody reads it as evidence the bug is caught. TWO FIXTURE ERRORS OF MY OWN, FOUND BY WATCHING THE TESTS AT BASE Both would have shipped as coverage that catches nothing: * The DST tests first used epoch `1793440800`, which is 2026-10-31 β a day off the transition. The fixture never entered the repeated hour and BOTH parametrisations passed at base. The constant is now DERIVED, with the derivation recorded in the test. * `test_an_exhausted_stamp_searchβ¦` first occupied a contiguous run of stamps, which does not exhaust the search at all: the anchor JUMPS PAST `newest`, so the slot after the newest generation is always free. It scored DID NOT RAISE β the guard was unreachable, not working. It now holds the listing empty while the files exist, which is the actual race (another process claimed them between this caller's listing and its claim). Change-scoped: 187 passed across test_tmux_session_restore.py, test_tmux_restore_observe.py and test_tmux_restore_trigger.py. The #1415 kill-mention ledger still passes with its own positive control (this change adds no new file). Per CLAUDE.md as of today, no full-tier run: CI is advisory and the local full-suite ritual is retired. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013jdbmhCKa6edhTmiADsziR Claude-Session-Id: 097b404c-db17-4472-bd37-dc90cf8fa675
|
Round-1 fix-pass claims for #1383, in the fenced form β Posted after the PR merged (squash π΄ Five of round 1's nine findings were NOT fixed before merge, and are not claimed above.
|
β¦r incident), and stop the staleness gate counting the crash's own damage (#1586) * fix(tmux-restore): two refusals compounded into 8.5h of silence β make one loud and stop the other counting the crash's own damage MEASURED 2026-09-11: the tmux server died at 02:36:44 with ~52 live claude conversations. The restore chain then failed to recover them for 8.5 HOURS, because two individually-correct refusals compounded. 02:37:14 no tmux server is running β REFUSING to restore. (rc 0, silent) β¦nothing for 8.5 hoursβ¦ 11:00:02 restore plan is out of step with the saved layout by 8.6h (limit 2.0h, basis=layout) β too stale, skipping. (a) A MID-SESSION no-server refusal is now LOUD; a cold-boot one stays quiet. The refusal itself is unchanged and still correct β creating a server inside the unit's cgroup is what lost 43 conversations on 2026-09-06. What was missing is the discriminator: `no_server_is_mid_session()` compares the POINTER's mtime against uptime. `cmd_save` writes a generation only when it found live claude panes, so a plan written INSIDE this boot proves a workspace existed during this boot and does not now β something a cold boot cannot produce. Uptime alone cannot do it (a fact about the host, not the workspace) and resurrect's `last` is the wrong witness for the same reason (b) is about: the crash refreshes it. The channel is the EXIT CODE, deliberately, and the exit-0 argument survives for the cold boot. `OnFailure=notify-failure@%n` bypasses do-not-disturb, and `nix/home.nix`'s rule for that bypass is precise about what abuses it: "any unit that can fail on a STANDING condition breaches it again". A cold boot IS such a condition β hence rc 0, unchanged. A mid-session death is an EVENT, fires at most once per death (`PathChanged=` is an event on one watched name), and is exactly the class the bypass exists for. A journal line or a marker file is a reader an operator has to RUN, which is the thing that did not happen for 8.5 hours. New rc 75 (EX_TEMPFAIL), not 1: `ExecMainStatus` already means three unrelated things at 1 here, and `tmux-restore-observe.sh` records that field. (b) The staleness gate no longer reads a post-crash layout save as evidence the plan is stale. The crash brought up a new EMPTY server; continuum autosaved that degraded layout at 10:59; the good 02:23 plan then looked 8.6h "out of step" and the gate refused. The guard against restoring a stale plan blocked recovery at exactly the moment it was needed β the inverted-basis shape #1317 already fixed once for powered-off time. `plan_staleness_hours` now cross-checks the "the plan stopped being refreshed" story (`state > mt`) against a witness the crash CANNOT refresh: the agent ledger, written by the claude/opencode processes themselves. Where the panes stopped when the plan did, the layout's extra hours were bought by a dead workspace and are discounted (new basis `ledger`). Where work continued β the real 2026-08-05 outage, plan frozen Jul 5 while resurrect ran to Jul 29 β the refusal is unchanged. PANE records only: `claude-s-<id>.json` records are written by agents with no $TMUX_PANE, i.e. by the investigation of the very outage, and counting those would refuse the recovery all over again. MEASURED from the incident's preserved evidence (`~/.cache/tmux-crash2-20260911T110407/`, a `cp -a` of the ledger): plan 02:23:15, newest pane record `claude-p28.json` 02:36:40, crash 02:36:44, boot 2026-09-06 18:01:46. So at 11:00:02 the plan was 8.6h old inside a 113h boot (=> mid-session, loud) and work outlived it by 0.22h (=> 0.22h vs the 2.0h limit, the recovery runs). The basis vocabulary is now FIVE arms (layout/ledger/liveness/skew/wall) and the `why` dict in `cmd_restore` covers all five β a shorter contract KeyErrors in the refusal path. The documented inert case (the liveness term when `uptime <= limit`) is DELIBERATELY left alone; this change does not subsume it. Tests: red at 337114e, green at HEAD. Every new test pins BOTH the pointer's mtime and uptime via `_boot_shape`, and the ledger witness via `_staleness_fixture(pane_age_s=β¦)`, because the dev-host tier has a real fresh plan and a populated ledger while the nix sandbox has neither β unpinned, they assert opposite arms per tier. 14/14 mutants killed, green control before and after. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session-Id: 2fd1ba4f-a498-421f-bda1-30bf1657519e * feat(tmux-restore): generation retention 48h -> 7 days (KEEP_GENERATIONS 192 -> 672) The operator asked for 7-day retention on 2026-09-11. β PREMISE, RESTATED CORRECTLY: "store all saves instead of just the most recent" is ALREADY SHIPPED β #1383 introduced generations and they are live on both hosts. This change is PURELY the retention window. Nothing about what gets written changes. WHY 48h WAS NOT ENOUGH. The 192 constant was argued from a worst case of "a Friday-night crash noticed Sunday" (~40h). 2026-09-11 refuted that worst case: the tmux server died TWICE in 24h, and the second recovery leaned on generations from before the first. A window a single bad weekend can consume leaves no margin for a second incident inside it. MEASURED AT THE NEW BOUND, not extrapolated. 2026-09-12, 672 synthesised generations of the live 53-entry shape (~34 KiB each, plan + cheat-sheet), host at load ~78: n disk richer_generation (warm) prune (no-op) 138 4.6 MiB 5-8 ms 0.2 ms 192 6.5 MiB 7-8 ms 0.3 ms 672 22.6 MiB 25-38 ms 0.6-0.9 ms 22.6 MiB matches the 22 MiB the ask predicted. The scan cost is fine: it is a once-per-`restore` call on a path that already waits up to 30s for a tmux server. β Those figures are WARM β the directory had just been written. The previously recorded 0.26s at 138 was a cold page cache; the measured 138->672 ratio (~3-7x) puts a cold scan at 672 around 1-2s. That is an extrapolation from a measured ratio, not a measured number, and it is said that way in the source. Pruning is unaffected in the steady state β one save prunes exactly one generation whatever the cap is. Tests: * `test_retention_spans_a_week_at_continuums_cadence` β red at 192 (48.0h vs the 168h asked for), green at HEAD. Pinned as a DERIVED SPAN rather than as `== 672`, because the literal is a consequence of the ask and continuum's 15-minute cadence, and `== 672` would go green for a cadence change that silently halved the window. * `test_pruning_still_behaves_at_the_FULL_retention_bound` β labelled an INVARIANT GUARD, not regression coverage: it derives its fixture size from the constant, so it is green at 192 and 672 alike and cannot go red on this change. It is the mechanical confirmation the pruner still behaves at the real bound, which nothing exercised before (every other pruning test runs at a fixture-sized 3). 4/4 mutants killed, green control before and after. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session-Id: 2fd1ba4f-a498-421f-bda1-30bf1657519e * revert: move retention 192->672 to its own PR (#1588) Not a change of mind β a SPLIT. The round-1 audit of #1586 returned two π΄ against the restore-refusal rework and nothing against retention. Retention is the operator's explicit ask, is independently mergeable, and should not wait on a rework it has no dependency on. The change is carried unaltered (plus the audit's π’-13/π’-14 corrections) on fix/retention-7d, off current origin/main, as #1588. This branch keeps the two refusal fixes only. This reverts commit d42e78a. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session-Id: 2fd1ba4f-a498-421f-bda1-30bf1657519e * fix(tmux-restore): round 1 β latch the loud arm, and scope the ledger witness to the plan's tmux generation Both round-1 π΄s were re-measured here before being fixed. Both held. π΄-1 β THE EXIT-75 ARM FIRED ON A STANDING CONDITION. A tmux socket OUTLIVES its server. Measured with a private socket: tmux -L rgprobe1 new-session -d ... -> socket present, has-session rc 0 tmux -L rgprobe1 kill-server ls /run/user/1000/tmux-1000/rgprobe1 -> STILL THERE tmux -L rgprobe1 has-session -> rc 1, "no server running on ..." and 7 of the 8 sockets in the operator's runtime dir were orphans. So after a death `ConditionPathExists=` keeps passing, `wait_for_tmux_server()` takes its 30s timeout arm, and `plan-written-this-boot` stays true for the rest of the boot β every activation would fire a DND-defeating `notify-failure@` toast, with `X-Restart-Triggers` restarting the unit on every home-manager switch (14 activations on 2026-09-11, four inside 9s). The source paragraph that claimed "`PathChanged=` is an event, so it fires at most once per death" rested on the socket disappearing; it does not, and that paragraph is rewritten. FIX: `claim_midsession_alert()` β an `O_CREAT|O_EXCL` marker in `$XDG_RUNTIME_DIR` (a tmpfs the OS destroys with the login session, so it is self-scoping and needs no cleanup path that could fail). KEYED ON THE PLAN'S MTIME, not the boot: 2026-09-11 had TWO deaths, 00:05 and 02:36, and a boot-scoped latch would have silenced the second. No runtime dir, or an unwritable one, WITHHOLDS the alarm rather than firing it β "once" cannot be promised there, and an unbounded toast through the bypass is worse than a missed one. `SuccessExitStatus=75` was considered and rejected: it suppresses the toast entirely, which is the silence this arm exists to end. π΄-2 β THE WITNESS WAS REFUTED BY THE EVIDENCE DIRECTORY IT CITED. Calling the shipped `newest_pane_ledger_activity()` on `~/.cache/tmux-crash2-20260911T110407/agent-ledger` returns 11:03:34 (`claude-p2.json`, a RECOVERY pane), not the 02:36:40 the docstring claimed: `worked` 8.672h vs `gap` 8.6h, so `worked < gap` was FALSE and the gate still refused. The stated invariant β "when the server dies they all stop writing at once" β is false: an agent in a pane of the NEW server writes `claude-p<N>.json` like any other, and pane ids restart at %0 so it overwrites the dead generation's file. FIX: scope the witness to the plan's own TMUX GENERATION, which this file already understands (`ledger_binding`'s generation check). `build_plan` now records `tmux_pid` on each entry; `newest_pane_ledger_activity(generation, β¦)` reads each record and keeps only that generation's. A post-crash agent is in a different generation and CANNOT move the witness. Grouping the snapshot by `tmux_pid`: 4025325 60 recs newest 2026-09-06 17:25:30 4063373 56 recs newest 2026-09-11 02:36:40 <- the plan's; died 02:36:44 1111077 31 recs newest 2026-09-11 00:05:46 (crash 1) 627687 9 recs newest 2026-09-07 21:54:13 1820124 2 recs newest 2026-09-11 11:03:34 <- post-crash β SCOPE, STATED IN THE SOURCE: a plan that does not record its generation β every plan written before this change, INCLUDING the preserved 02:23 plan β measures nothing and the gate behaves exactly as before. Recovering vs a legacy plan is `restore --plan <generation>`, which bypasses the gate. Recovering the generation by session-id vote was measured and REJECTED as a heuristic: the votes on the real snapshot are 54/11/4/1, contaminated by conversations resumed across generations. `record_generation()` is now one reader shared with `ledger_binding`, rather than the field being spelled at two sites. BEFORE/AFTER ON IDENTICAL INPUTS (fixture shaped like the snapshot): origin/main 8.580h basis=layout REFUSES | activations [0, 0, 0] round 1 8.580h basis=layout REFUSES | activations [75, 75, 75] HEAD 0.223h basis=ledger passes | activations [75, 0, 0] π‘ ALSO FIXED -3 three artifacts asserting the old exit-0 contract: the `cmd_restore` paragraph, `tmux-restore-observe.sh`'s refusal block (now explains both arms and that ExecMainStatus=0 no longer means cold boot), and `test_tmux_restore_observe.py`'s two. -4 "ConditionPathExists= catches the ordinary shape of that" β deleted, false. -5 the `skew` arm returned a ledger-DISCOUNTED number under layout wording. New `ledger-skew` basis; the vocabulary is six arms and the `why` dict covers all six. -7 the 2026-08-05 fixture sat exactly ON the discount boundary (`worked == gap == 599h`) and asserted no basis, so it passed on both sides of the M8 mutant. Now `pane_age_s=0.0` (overshoots) and asserts the basis. -8 `test_the_refusal_is_checked_BEFORE_...` reached the diagnosis without `_boot_shape`, reading the operator's REAL plan and `/proc/uptime`. -9 the shape-change guard and the boundary now have tests; the sweep is 29 mutants, all killed. -10 "125 pane records against 105" was the LIVE ledger, not the snapshot. The snapshot is 158 pane / 94 non-pane. -11 the remedy said `restore --best`, which is a DEAD END from that state (same refusal, another 30s, same exit). It now says: start tmux. -15 the mtime-not-content caveat is stated. π΄ FOUND BY MY OWN CONTROL, NOT BY THE AUDIT: the first version of the latch tests wrote `alerted-<plan-mtime>` files into the operator's LIVE `/run/user/1000/tmux-session-restore/` β three of them keyed on real plan mtimes, because those tests also read the real plan. Deployed, that is a test run pre-claiming the latch and silencing a real incident's alert. The four leaked files were removed (the deployed script has zero occurrences of `claim_midsession_alert`, so nothing else could have written them) and the fix is structural: an AUTOUSE fixture redirects `XDG_RUNTIME_DIR` for every test in the file, plus a guard pinning that the latch writes nowhere else. β NOT FIXED, DELIBERATELY β π‘-6. The `wall` basis (unreadable `<resurrect-dir>/last`) returns before the cross-check, so the 2026-09-11 shape still refuses there; re-measured here, `(8.600h, 'wall')`. Discounting that number was tried and is unsafe: the wall path has no liveness term to pair with, so `max(gap, live)` is circular, and `min(wall, worked)` alone would wave through a plan whose workspace died weeks ago. There is no signal on that path saying "recover me". The reason is now in the source rather than left as an oversight. RETENTION (192->672) IS SPLIT OUT to #1588, reverted on this branch. Tests: 173 in this file (was 137 at base), 327 across the restore chain. 29/29 mutants killed, green control watched before and after; the one initial survivor (an empty generation matching every record) was a FIXTURE blind spot β every record carried a pid, so the inner comparison rejected anyway β fixed by adding a record with no generation, then re-run and killed. Ambient-independence control: the whole file passes with `HOME` pointed at an empty directory, so no test depends on the operator's plan, ledger or projects. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session-Id: 2fd1ba4f-a498-421f-bda1-30bf1657519e * test(guard): classify this PR's two new kill-mention files β the prose IS the finding CI was red on `test_every_kill_server_call_site_in_the_repo_is_classified`, and it is THIS PR's own red, not the inherited one: measured on the merged tree, the added set is exactly scripts/tmux-session-restore.py scripts/session-analysis/tests/test_tmux_session_restore.py and `removed` is empty. The two-way ledger is working as designed β a new file reaching the wide-kill vocabulary must be classified by a human. All four mentions are PROSE, and the prose is the FINDING rather than decoration: a tmux socket file OUTLIVES its server, so `ConditionPathExists=` keeps passing after the operator's server dies. That is the entire reason the loud arm needed a once-per-incident latch instead of trusting the path trigger to fire once. Measured on a private socket, and on the live host where 7 of 8 runtime sockets were orphans with no server behind them. Deleting either sentence to get this guard green would delete the reason the latch exists, so they are re-justified here rather than reworded to dodge the scanner. Verified NOT call sites, two ways rather than asserted: * an AST walk puts the script's complete spawn argv[0] set at {tmux, grep}, with no `kill-sβ¦` subcommand in any argv list; * the sibling scanner `test_no_tracked_shell_text_writes_a_kill_this_guard_ would_deny` PASSES on both files β i.e. neither carries shell text this guard would deny. Also merges current `origin/main` into the branch (it was behind, and `main` had moved the ledger), so the tree tested here is the tree the merge creates. Both ledger tests pass; 2 passed of that pair, and the restore suites are green on the merged result. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tw7uW1Cypyy6kKVnSRWbga Claude-Session-Id: 2fd1ba4f-a498-421f-bda1-30bf1657519e --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
β¦ONS 192 -> 672) (#1588) SPLIT OUT OF #1586. This shipped there originally as commit 2 of 2; a round-1 audit found two π΄ in the OTHER item (the restore-refusal rework) and nothing in this one. It is the operator's explicit ask, it is independently mergeable, and it should not wait on a rework it has no dependency on. #1586 keeps the refusal fixes and reverts this change on its own branch. β PREMISE, RESTATED CORRECTLY: "store all saves instead of just the most recent" is ALREADY SHIPPED β #1383 introduced generations and they are live on both hosts. This change is PURELY the retention window. Nothing about what gets written changes. WHY 48h WAS NOT ENOUGH. The 192 constant was argued from a worst case of "a Friday-night crash noticed Sunday" (~40h). 2026-09-11 refuted that worst case: the tmux server died TWICE in 24h, and the second recovery leaned on generations from before the first. MEASURED AT THE NEW BOUND, not extrapolated. 672 synthesised generations of the live 53-entry shape (~34 KiB each, plan + cheat-sheet), host at load ~78: n disk richer_generation (warm) prune (no-op) 138 4.6 MiB 5-8 ms 0.2 ms 192 6.5 MiB 7-8 ms 0.3 ms 672 22.6 MiB 25-38 ms 0.6-0.9 ms β Those figures are WARM. The previously recorded 0.26s at 138 was a cold page cache; the measured 138->672 ratio (~3-7x) puts a cold scan at 672 around 1-2s. That is an extrapolation from a measured ratio, not a measured number, and the source says so. π΄ WHAT IT DOES NOT BUY, now stated in the source and the README: the extra retention widens MANUAL recovery (`restore --plan <generation>`) only. `richer_generation` scans at most RECOVERY_ALERT_WINDOW (4) generations behind the pointer, so `--best` reaches exactly as far at 672 as at 192. Those are independent knobs and a new invariant guard fails if anyone couples them. The live generations dir is AT the old cap today β 192 generations, 6.53 MiB, MEASURED 2026-09-12 (an earlier draft of this comment said 146 / 4 MiB, which was already stale when written). Tests: red at 663bc86 (`test_retention_spans_a_week_at_continuums_cadence`, 48.0h vs the 168h asked for), green at HEAD. Pinned as a DERIVED SPAN rather than `== 672`, because the literal is a consequence of the ask and continuum's 15-minute cadence, and `== 672` would go green for a cadence change that silently halved the window. The other two new tests are labelled INVARIANT GUARDS in-source and are green at base by construction β they are the mechanical confirmation that the pruner still behaves at the real bound (every other pruning test runs at a fixture-sized 3) and that the scan window did not move. 5/5 mutants killed, green control watched before and after. Claude-Session-Id: 2fd1ba4f-a498-421f-bda1-30bf1657519e Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The defect, measured
~/.config/initiatives/restore-plan.jsonwas a single mutable file with no backup, andcmd_saveoverwrote it in place. Measured on the workbench 2026-09-06:cmd_saveoverwrote the plan with 10 entries. The cheat-sheet went in the same second. No backup existed.The 47 conversations were recoverable only because tmux-resurrect keeps its saves TIMESTAMPED and each pane line happens to carry a full
claude --resume <id>; 24 were rebuilt from that file.π΄ That asymmetry was the bug: a bad save cost the layout nothing and the bindings everything.
The fix
Mirror what resurrect already does. Each
savewrites an immutable timestamped generation intorestore-plans/restore-plan_<ts>.json(+ the cheat-sheet, which had the identical defect from the identical writer) and repointsrestore-plan.json/restore-cheatsheet.mdat it as relative symlinks.Against each stated constraint
test_a_shrinking_save_is_written_not_refusedpins this, and is labelled an invariant guard (green at base) rather than counted as regression coverage.cmd_restore,plan_staleness_hours,cmd_show,--planandtmux-restore-observe.shall read the same two paths;Path.stat()/Path.exists()follow symlinks. See the second commit β verifying this instead of asserting it found that GNUstat(1)uses lstat, so the observe script needed-L.claude --resume <id>survives.Also fixed β same writer, same shape
_write_atomicβwrite_texttruncates first, so a save killed mid-write left a zero-byte plan. Temp file +os.replace.free_generation_stampβ the stamp has one-second resolution and the incident was a same-second write. A manualsaveracing the 15-min hook would have shared a stamp and clobbered the previous generation from inside the mechanism built to stop that. It must also be strictly newer than everything present, not merely free: a first draft returned the slot pruning had just freed, the new generation sorted OLDEST, and the run reported4 kept (max 3), 0 prunedβ the retention cap silently unenforced. Caught by its own test; a backwards clock reproduces it with no race.adopt_pre_generation_filesβ the deploy of this change must not itself be the bad save. A host still on the old writer has a regular file at the pointer path holding possibly the only good plan; it is copied in as a generation before the pointer moves.prune_generations(protect=β¦)β pruning is the only code here that deletes, so the only code that can recreate the defect. It refuses to unlink the generation the pointer was just aimed at.Red / green matrix β base
c5e425c713 new tests in
test_tmux_session_restore.py(102 in file at HEAD), 1 intest_tmux_restore_observe.py(44 at HEAD). 12 of 13 + the observe test are RED at base; the three core regressions fail on the data loss, not on anAttributeErrorβ the assertions use the test module's own_bound_idsand walk the state dir implementation-blind, because measured, usingtsr.bound_idsmade two of them die on AttributeError before reaching the assertion:a_shrinking_save_does_not_destroy_the_previous_bindings46 of the previous save's 46 bound session ids are no longer recoverablea_shrinking_save_does_not_destroy_the_previous_cheat_sheetno file β¦ still carries 'claude --resume <id>'a_pre_generations_plan_is_preserved_by_the_first_new_save46 bound ids β¦ lost when the pointer was first swappeda_save_dropping_bound_ids_names_the_recovery_commandthe shrink report did not name the 46 dropped bindingsthe_plan_mtime_is_the_PLANS_not_the_symlinksskew was not computed from the plan GENERATION's mtime (1001) against the layout's (1)Green at base by design, labelled as invariant guards (not counted as regression coverage):
a_shrinking_save_is_written_not_refused,a_shrink_that_drops_no_bindings_stays_quiet(the discriminating control β ids, not counts),the_pointer_still_reads_as_the_current_plan.Mutation sweep β 13 mutants
Fresh tree per mutant,
PYTHONDONTWRITEBYTECODE=1, every patch verified to have applied (astr.replacematching nothing scores SURVIVED), each kill confirmed to carry that guard's own assertion text:12 KILLED β no-adoption Β· protect-ignored Β· prune-wrong-end Β· prune-never Β· report-on-count Β· report-silent Β· stamp-free-not-newest Β· raw-second-stamp Β· non-atomic-write Β· empty-plan-spends-a-slot Β· generations-dir-hardcoded-to-the-live-dir Β· full revert to the in-place overwrite.
1 SURVIVED, and it is explained rather than waved away: removing only the generation writes leaves adoption running, and adoption alone then provides a one-deep backup β the mutant fails to break the property, rather than the test failing to see it. Reverting both together (M1b) kills both regression tests with their own messages, and the base-tree measurement above is the definitive form of that same mutant.
Positive control for retention: 5 saves at
keep=3leave exactly the newest 3, with0 prunedasserted below the cap and1 prunedon the run that first exceeds it β a pruner wired to nothing cannot pass.Gate β merged tree, base
c5e425c7scripts/gate.sh --tier both):GATE: RESULT=PASS exit=0; pytestRESULT: PASS (exit=0), nodeRESULT: PASS (exit=0); nodeTOTAL suites=5 files=41 tests=1449 pass=1449 fail=0. session-analysis moved 564 β 577.git archive, one derivation at a time): nodetests PASS (TOTAL suites=5 files=41 tests=1449 pass=1449 fail=0,RESULT: PASS (exit=0)). pytests PASS βPASS scripts/dl-router/tests (collected=1020 passed=1020 skipped=0 floor=942),RESULT: PASS (exit=0)β on the second run; see below.The tier-2 pytests flake, and the control that settled it
The first tier-2 pytests run reported one failure: dl-router's
test_six_writers_with_a_tiny_busy_timeout_still_land_every_rowβ a SQLite writer-contention test with a deliberately tiny busy timeout β failingOperationalError('database is locked'), the exact contention modeCLAUDE.mddocuments. It ran while the box carried ~50 load from eight other sessions' gate runs, and this branch touches zero files underscripts/dl-router.A theory that explains a failure is not evidence for it, so the same derivation was re-run on the same tree:
collected=1020 passed=1019 failed=1RESULT: FAIL (exit=1)collected=1020 passed=1020 failed=0RESULT: PASS (exit=0)π΄ The control's own instrument was validated before its verdict was read. The first attempt used
nix build --rebuild, which errors on a derivation that never built successfully (some outputs β¦ are not valid, so checking is not possible) β the build never ran,nix logserved the stale log from the failed attempt, and the harness printed a confident FAIL it had never measured. The rerun therefore fingerprints the log before and after:600f1632a3b0β634c14b21a77, i.e. a fresh run, not the cached one. Without that check this would have read as a reproduced failure.Both tiers are green on the merged tree at base
c5e425c7.