docs(handoff): two tmux crashes in 24h, both recovered by hand β and autorecovery did nothing for 8.5 hours - #1503
Merged
Conversation
β¦23), both recovered BY HAND and
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Records the 2026-09-11 crashes and the operator's 7-day retention ask.
Two crashes, 00:08 and 02:23. Both recovered by hand, neither attributed: no OOM (0 matches against a positive control of 54), no coredump, nothing logged at teardown. My own audit subagents are ruled out β the last finished 50 minutes before crash 2 (on 2026-09-07 the agent's kill was 72 seconds before).
The second went undetected for 8.5 hours, and that is a design flaw, not a malfunction. Two individually-correct refusals compounded:
02:37:14 no tmux server is running β REFUSING to restore.β fix(tmux-restore): REFUSE when there is no tmux server, instead of manufacturing one systemd then killsΒ #1351's guard, correct (starting one in the unit's cgroup lost 43 conversations on 2026-09-06). But it exits 0 by design, nothing retries, and nothing else on this host starts tmux.11:00:02 restore plan is out of step with the saved layout by 8.6hβ the crash spawned a new server, resurrect saved the degraded layout, and that made the good 53-conversation plan look stale. The guard against restoring a stale plan blocked recovery at exactly the moment it was needed β the same inverted-basis shape fix(tmux-restore): the staleness gate counted POWERED-OFF time against the planΒ #1317 fixed for powered-off time.π΄ #1494 performed both recoveries and is still unmerged, so
maincarries thewindow_statefalse positive and the bare-claudesend. That is rank 1.On the 7-day ask: "store all saves instead of just the most recent" is already shipped (#1383) β generations are live on both hosts. The real delta is 48h β 7d:
KEEP_GENERATIONS 192 β 672, measured at ~34 KiB per generation, so β22 MiB (4 MiB today).Also records the reusable recovery procedure that worked twice, and that the save side protected the plan both times (
no live claude panes found β nothing to snapshot).