Conversation
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Integrate exact no-allocation state sizing and verify measured graph workspace against the admitted host-memory envelope. Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…ste-microsoft-deepseek-v41-memory-admission
Forward SIGHUP through the external watchdog, make unsupported dynamic embedding requests fail the next operation explicitly, and use the bounded ubatch for default model admission. Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…ste-microsoft-deepseek-v41-memory-admission
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Independent review after publication found three remaining admission blockers in head
The correction will preserve explicit |
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Published blocker replacement head |
|
CI note for run |
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Follow-up: inspected Windows job |
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> # Conflicts: # tests/test_strix_memory_watchdog.py
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Final admission replacement published at |
|
Hardware-use freeze: head Written by GPT-5.6 Sol. |
|
HARDWARE FREEZE: head |
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> # Conflicts: # docs/strix-memory-watchdog.md # scripts/strix_memory_watchdog.py # tests/test_strix_memory_watchdog.py
|
Review candidate |
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Corrected review candidate |
Assisted-by: GPT-5.6 Sol Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Overview
Adds the DeepSeek V4.1 Strix Halo host-memory admission and watchdog integration layer for halo-box#48. This PR is stacked on #7 at exact full-graph head
081c549451b92c01fe832ffb0bdd9ad90370e21aand includes the accepted external-watchdog behavior from #1 at exact head778db6f50eae04e6c232c69b9575bdbd0747962b.Before expert-cache or model backend allocation, admission reads integer-byte host use from Linux procfs, rejects any configured swap entry, requires an IGPU-only unified-memory topology, measures full-graph state through the no-allocation memory implementation, accounts the conservative graph workspace and all bounded staging/output categories, and auto-fits complete expert slots under the 116 GiB soft ceiling. It enforces the 118 GiB external-watchdog threshold and strict
<120 GiBhard requirement without changing the GGUF, quantization, requested context, or explicit ubatch.The common CLI and server use ubatch 32 for DeepSeek V4.1 only when
-ubis not specified. The server resolves implicit auto parallelism to one DeepSeek sequence and aligns no-allocation fitting to the admitted envelope; explicit values are preserved and must fit. Direct libllama defaults resolve the admitted DeepSeek ubatch while other architectures retain an effective default of 512. Expert replacement accounting includes the full incoming routed-expert union plus the largest 4096-aligned direct-I/O bounce read. Runtime-context ownership is acquired immediately after parameter validation, before sampler, output, state, backend, or scheduler allocation, with exception rollback.The external watchdog remains a composable process-group wrapper rather than duplicated C++ process control. It monitors host-wide use and swap after startup, forwards
SIGHUP,SIGINT, andSIGTERM, escalates at the configured grace timeout, and caps overrides at repository policy limits. The independently approved canonical replacement publishes a version 2 self-authenticating lease, per-sample atomic heartbeat, guardian fail-closed lifecycle, and persistent JSONL audit that bind the live watchdog, exact thresholds, procfs root, child process group, and exact child command. Independent exact-SHA review closed all nine recorded watchdog findings.Measurements
No performance claim is made. This change is a safety gate and runtime accounting layer; real Strix Halo model validation is currently blocked by active unrelated DeepSeek V4 work, an enabled 32 GiB
/swapfile, and no active watchdog. I did not interrupt workloads or change swap/ROCm configuration.Baseline:
Not applicable: no throughput or latency claim.
After:
Not applicable: no throughput or latency claim.
Correctness:
llama-cli,llama-server, admission/schema/Engram/expert/memory/runtime tests, expert-store, recurrent rollback, save/load state, arg parser, and architecture test targets.ResourceWarningpromoted to errors on macOS: 28 passed and 8 Linux-only tests skipped in each mode. The watchdog script SHA-256 isd2781a25f978dd2bc14fc113079aa2dbf513aa157b44da9d0d51d750daa6c94f, byte-identical to the independently approved artifact.py_compile, flake8 with flake8-no-print, and file-scopedtypassed.test-llama-archs --arch qwen4expand--arch grovemoepassed available backend comparisons; DeepSeek41 compact architecture checks skipped as intended.git diff --checkand the ASCII-only diff scan passed.Additional information
Admission diagnostics report current host use, fixed/dense/state/workspace/staging/output bytes, selected and required expert slots, exact cache bytes, replacement bytes, aligned direct-I/O bounce bytes, safety margin, soft/watchdog/hard thresholds, ignored device-reported bytes, and the rejecting category. Context checkpoints are explicit at 32768, 65536, 98304, and 131072; an unsupported or non-fitting request fails instead of being lowered.
The documented validation command uses NVMe model storage and explicitly binds
ROCR_VISIBLE_DEVICES=0,HIP_VISIBLE_DEVICES=0,HIP_LAUNCH_BLOCKING=1,-dev ROCm0,-ub 32,-np 1, 192 slots, and the exact 72900 MiB expert budget. Its preflight must confirm thatROCm0reportsgfx1151and validate the active watchdog lease. Buffered expert/Engram I/O cannot claim boundedness. Dynamic embedding-output requests are rejected and cause the next encode/decode operation to fail explicitly because they are outside the admitted output profile.Requirements
b26c0e687a556ca501fb056dbbdc4c6294bf90b9andd0a161602f96ae7fb90d5e702f010f198ce71c6awere created during agent-coordinated stack propagation and contain no independent semantic change; they do not carry the required Assisted-by/Copilot trailers. History was not rewritten or force-pushed.