Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
123 commits
Select commit Hold shift + click to select a range
9de5596
kernel parity pass, step 0: the kq kernel microbench, the reference r…
borisbat Sep 1, 2026
a29ce1c
kernel parity, CPU item 2: sign="vec" for iq3s and iq2s - the sign co…
borisbat Sep 1, 2026
c6ce3c1
kernel parity, CPU item 2b: sign="vec" reaches iq3xxs, iq2xs and iq2x…
borisbat Sep 1, 2026
882c2e2
kq_kernel_bench: =name selects one registry row exactly
borisbat Sep 1, 2026
ba586e6
kernel parity, CPU item 3: the iq2 formats' u64 grid pair is one 8-by…
borisbat Sep 1, 2026
c2507a2
kernel parity, CPU item 4: the vector sign column is the grid formats…
borisbat Sep 1, 2026
edaba28
kernel parity, CPU item 5: the column dword read for iq3xxs's grid by…
borisbat Sep 1, 2026
714eb8f
kernel parity: the lab covers all sixteen formats - kq_kernel_bench a…
borisbat Sep 1, 2026
ddac6fb
kernel parity, CPU k-quants: k3 and k2 flush their i16 chains once pe…
borisbat Sep 1, 2026
e45ac5c
kernel parity: the M1 ladder and the post-k3 zen2 rows in the plan
borisbat Sep 1, 2026
e0418cf
the k-quant memo through the Markdown ASCII gate
borisbat Sep 1, 2026
da989ab
followup_general 63: the rollback-and-ban sampler for reasoning trace…
borisbat Sep 1, 2026
c26e70a
kernel parity: the ARM grid memo, its fix on the queue, and the globa…
borisbat Sep 1, 2026
ec2357a
kernel parity: the zen4 ladder - 65/65 TEST, VNNI seats pay on the do…
borisbat Sep 1, 2026
78c193c
the ARM memo through the Markdown ASCII gate
borisbat Sep 1, 2026
b61b88d
llvm_tune: the x86-amx CPU class (AMX-INT8 + AMX-TILE over the 512-bi…
borisbat Sep 1, 2026
220fcec
followup_general 64: Qwen 3.8 thinking control (reasoning_effort low/…
borisbat Sep 1, 2026
4ccd7ea
HOW_TO_GET_SIDECAR.md: the mint, the provenance read, the export and …
borisbat Sep 1, 2026
1c0e29e
the Qwen 3.8 memo through the Markdown ASCII gate
borisbat Sep 1, 2026
17c44b2
kernel parity: the zen4 model-level stamp and the repack caveat on th…
borisbat Sep 1, 2026
f6e6b12
kq_kernel_bench: 64-byte-aligned planes (Intel splits a misaligned 64…
borisbat Sep 1, 2026
6c9c07f
the x86-amx defaults profile, minted on a Granite Rapids c8i.4xlarge;…
borisbat Sep 1, 2026
612d7ac
HOW_TO_GET_SIDECAR.md: the new-class walk as run on the Intel box, an…
borisbat Sep 1, 2026
d99bce9
kernel parity: the Intel ladder (Granite Rapids, x86-amx) with its tw…
borisbat Sep 1, 2026
cd2f427
kernel parity: the zen2 ladder re-run with the fixed bench
borisbat Sep 1, 2026
4e81727
kq_kernel_bench: one arena for every plane at fixed staggered offsets…
borisbat Sep 1, 2026
37bdeaa
kernel parity, ARM grid decode: row pairs under the sdot lattice - tw…
borisbat Sep 1, 2026
7ddb256
kernel parity, ARM grid decode: the ksigns formats sign through a +-1…
borisbat Sep 1, 2026
9e61caf
kernel parity ledgers: the ARM grid landing under followup 61 (the du…
borisbat Sep 1, 2026
6641a2e
kernel parity, grid decode: one row-group emitter at any width - and …
borisbat Sep 1, 2026
0f07299
kq_kernel_bench: --base-align and --base-offset put the arena at a ch…
borisbat Sep 1, 2026
50650f2
kq_kernel_bench --team N: the GEMV dispatched the engine's way - N ro…
borisbat Sep 1, 2026
4af2d6e
k6 planes, CPU flavor: the qh half packed per sub-block - one byte ca…
borisbat Sep 1, 2026
7211c6f
k3 planes, CPU flavor: qs packed per sub-block at 2j and hmask one co…
borisbat Sep 1, 2026
02af2f3
kernel parity plan: the k3 transpose landing and the k5 shape
borisbat Sep 1, 2026
a166a65
kernel parity plan: the k5 transpose measured and killed - x86 byte-v…
borisbat Sep 1, 2026
1605d3e
kq_kernel_bench --team: the engine's own splitter and dispatcher, not…
borisbat Sep 1, 2026
3b67227
kernel ladder BIG=1: the DRAM-bound decode row (d=32768) - the many-l…
borisbat Sep 1, 2026
04d54f5
k6/k3 decode, x86: the per-16 sub-scale rides the chain flush - pmadd…
borisbat Sep 1, 2026
1880c43
k4/k5 decode, x86: the sub-scale rides the chain flush too - one 6-bi…
borisbat Sep 1, 2026
e538ccf
kernel ladder: team + DRAM-bound decode rows are the default, SOLO=1 …
borisbat Sep 1, 2026
fedd15b
kernel parity plan: zen2 closing tables (v4 one-thread + engine shape…
borisbat Sep 1, 2026
6cb455c
kernel parity plan: the M1 closing tables (one thread at the engine p…
borisbat Sep 1, 2026
0455b11
grid decode: DASLLAMA_GRID_ROWS_X86=1 re-arms the x86 row-group compo…
borisbat Sep 1, 2026
6a6ee90
grid gather panel and the tile's byte-expand scratch: a cache line's …
borisbat Sep 1, 2026
f548521
grid decode, x86-vnni512: iq2s and iq2xxs take the row form by class;…
borisbat Sep 1, 2026
df4409f
grid_rows_path: the comment at the three-line cap
borisbat Sep 1, 2026
04a122f
kernel parity plan: Intel v2 round one (engine shape), M4 v1
borisbat Sep 1, 2026
1224367
grid row form, x86: the sign column's mask for every format - the ksi…
borisbat Sep 1, 2026
0ae7686
kernel parity plan: the zen4 row-form results, the class gate, the iq…
borisbat Sep 1, 2026
4592153
kernel parity plan: zen4 round four (column-sign row form)
borisbat Sep 1, 2026
e7249de
kernel parity plan: Intel round two one-thread table
borisbat Sep 1, 2026
b5925fb
kernel parity plan: the M4 minted-profile table
borisbat Sep 1, 2026
f2f8908
kernel parity plan: Intel seat races (k6 at parity at 16 lanes), the …
borisbat Sep 1, 2026
32911d4
kernel parity plan: the M4 engine-shape table on the minted profile
borisbat Sep 1, 2026
111becf
kernel parity plan: the M4 k5 deposit result - the M4 closed
borisbat Sep 1, 2026
1e94e38
grid decode: the x86-amx class takes the row form for iq2xxs and iq3x…
borisbat Sep 1, 2026
3bd22a1
kernel parity plan: Intel round three and the amx gate verification
borisbat Sep 1, 2026
eec815e
kernel parity plan: the Intel k6 mode probe (the AVX-512 license shap…
borisbat Sep 1, 2026
f51bb43
generated kernels: a 64-byte-aligned entry alloca realigns the frame …
borisbat Sep 1, 2026
293180c
kernel parity plan: the frame realignment, the first-format bench art…
borisbat Sep 1, 2026
4a611cd
grid decode: iq2s leaves the zen4 row form (3257 vs the panel's 2525 …
borisbat Sep 1, 2026
1e9a1a5
kernel parity plan: the burn-in verification and the final class gate
borisbat Sep 1, 2026
21bb51c
kernel parity plan: the next session's rulings - the gemv races the t…
borisbat Sep 1, 2026
a8ba469
tune: the gemv companion gets its own seat - raced among the tile win…
borisbat Sep 1, 2026
11eba92
kernel parity plan + sidecar how-to: the gemv's own seat landed, the …
borisbat Sep 1, 2026
7eb39c0
research memo: llama.cpp's grid kernels on zen4 - no AVX-512 path, th…
borisbat Sep 1, 2026
9a9aa55
kernel parity plan: the zen2 mint smoke, the zen4 research memo in th…
borisbat Sep 1, 2026
d35f24b
grid row form, x86: DASLLAMA_GRID_QPANEL=1 builds the row group's vec…
borisbat Sep 1, 2026
5869d5b
tune profiles x86-vnni512 and x86-amx: re-minted on the emitter they …
borisbat Sep 1, 2026
9ace1f8
gemm emitter: the ymm vreg budget counts 32 registers on EVEX targets
borisbat Sep 1, 2026
3a8ec2e
kernel parity plan: the re-mint round, the declined width-256 row, th…
borisbat Sep 1, 2026
c8c1ea6
kernel parity plan: zen4 v3 and Intel v3 ladders on the re-minted pro…
borisbat Sep 1, 2026
611bd35
kq_kernel_bench: --each appends every round's sample - the per-round …
borisbat Sep 1, 2026
6bd0919
grid row form: the qpanel spelling is gone - a loss or a wash on zen2…
borisbat Sep 1, 2026
3373bea
grid gemv: the VBMI symbol lattice as a perm row for iq2xxs and iq3xx…
borisbat Sep 1, 2026
e52fd31
vbmi lattice: the iq3 planes pack two weights each (2 x 3 bits), not …
borisbat Sep 1, 2026
2a6af04
vbmi lattice: iq3s, iq2xs and iq2s join (512 / 512 / 1024-entry grids…
borisbat Sep 1, 2026
9ec45a6
kq_kernel_bench: --chunks-per-lane sets the engine splitter's gemv gr…
borisbat Sep 1, 2026
03864ea
kernel parity plan: the mr16 race verdict on both boxes; the Intel k6…
borisbat Sep 1, 2026
cb3238e
kq_kernel_bench: three unmeasured warm rounds in the SOLO arm - the I…
borisbat Sep 1, 2026
a8748f2
kernel parity plan: the lattice's Intel verdict; the k6 gap is mode-s…
borisbat Sep 1, 2026
7544ae3
kernel parity plan: the lattice's full Intel race - every grid row cl…
borisbat Sep 1, 2026
8535293
kq_kernel_bench: a 'timed rounds begin' log line before each timed lo…
borisbat Sep 1, 2026
e206ae4
kernel parity plan: the lattice sweeps zen4; the Intel gap is memory-…
borisbat Sep 1, 2026
3bf5fef
vbmi lattice: compile-time hygiene - grid tables held as fixed-array …
borisbat Sep 1, 2026
e42c386
kernel parity plan: the cold-start crash bisected - vbmi perm-row pre…
borisbat Sep 1, 2026
ae2a881
vbmi lattice: one grid table per helper frame - five by-value tables …
borisbat Sep 1, 2026
c31d158
llvm_boost: LLVMBuildGEP2's inbounds goes through LLVMBuildInBoundsGE…
borisbat Sep 1, 2026
bde35df
kernel parity plan: the cold-start crash found and fixed - the GEP wr…
borisbat Sep 1, 2026
d4a5876
kernel_ladder.sh: KL_MODULE_CACHE passes -module-cache to the bench s…
borisbat Sep 1, 2026
b4d4918
tune profile x86-amx: re-minted on the lattice emitter (Granite Rapid…
borisbat Sep 1, 2026
369d552
tune profile x86-vnni512: re-minted on the lattice emitter (zen4 c7a)…
borisbat Sep 1, 2026
fe439d0
kq_kernel_bench: a '# ysum' line per gemv row under --tsv - modes and…
borisbat Sep 1, 2026
0d6cf5d
kernel parity plan: the Intel k6 mode gap vanished with the rebase; z…
borisbat Sep 1, 2026
8e71b52
tune harness: the gemv seat races at the engine's decode shape; gemv …
borisbat Sep 1, 2026
7db15e6
gemm emitter: a pf perm knob - software prefetch pf bytes ahead on th…
borisbat Sep 1, 2026
8187672
tune harness: the engine-shape seat race tiles the streamed fixture (…
borisbat Sep 1, 2026
972d08e
kernel parity plan: x86 hybrid core detection follow-up (Apple-only t…
borisbat Sep 1, 2026
14f72ca
followup_general 66: x86 hybrid core detection for the lane policy (A…
borisbat Sep 1, 2026
589139b
tune harness: the engine-shape seat fixture is built at the ffn width…
borisbat Sep 1, 2026
d98323e
kq_kernel_bench: a q8s16 arm - q8 over binary16 group scales, the wsc…
borisbat Sep 1, 2026
05ec4c3
kernel parity plan: q8s16 like-for-like row, the prefetch null, the p…
borisbat Sep 1, 2026
179043e
followup_general 67: per-format decode lane cap on SMT x86 - zen4's s…
borisbat Sep 1, 2026
317a9a1
tune harness: the engine-shape seat fixture beats any L3 (320 tiles, …
borisbat Sep 1, 2026
33ed27a
tune profile x86-amx: re-minted with the L3-proof engine-shape seat r…
borisbat Sep 1, 2026
12e0dc6
tune profile x86-vnni512: re-minted with the L3-proof engine-shape se…
borisbat Sep 2, 2026
8c07d77
dasLLVM harvest: the x64 tier-gate inventory (ARCHITECTURE.md#x64-tie…
borisbat Sep 2, 2026
9d55c0c
dasLLAMA harvest: three anchored mechanisms (the sub-block k3/k6 plan…
borisbat Sep 2, 2026
b94f2ef
tune profile arm-i8mm: re-minted on the rebased branch (M4 Pro; 49 en…
borisbat Sep 2, 2026
44cbfb2
followup_general 68: iq2xs on arm-i8mm regressed with the shared row-…
borisbat Sep 2, 2026
a467ee5
followup_general 68 corrected: M4 iq2xs 0.90/0.95 is the format's tru…
borisbat Sep 2, 2026
e20600b
dasLLVM review-audit fixes: the tier-gates section's two factual bugs…
borisbat Sep 2, 2026
bdeacff
dasLLAMA review-audit fixes: the bare model-scaled resize, the grain …
borisbat Sep 2, 2026
7b951e3
style-hygiene round: dead gather names, family docs onto their owners…
borisbat Sep 2, 2026
a931270
TDD + woodpecker round: the lattice tables and the GPU gather get sil…
borisbat Sep 2, 2026
a3cfd7d
test fixes from the first run: the GEP scanner self-excludes and gets…
borisbat Sep 2, 2026
73721d0
popen_argv respells argv[0] natively on Windows: CreateProcess reject…
borisbat Sep 2, 2026
eafdd0f
dasLLAMA: the zen2 board row for gemma-4-12B Q4_K_M re-minted on the …
borisbat Sep 2, 2026
09a1702
spawn fix, style round: the three caller-side backslash normalizes go…
borisbat Sep 2, 2026
fb76369
preflight round: the lattice test's grid words come from one helper, …
borisbat Sep 2, 2026
2d7a6a3
the GEP scanner matches the call shape and ignores comments: a mentio…
borisbat Sep 2, 2026
23cc1a3
kq_kernel_bench reads the tune mode off g_env_tune, not the environme…
borisbat Sep 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 1 addition & 2 deletions daslib/fio.das
Original file line number Diff line number Diff line change
Expand Up @@ -692,8 +692,7 @@ def rmdir_rec_result(path : string) : fs_result_bool {

def run_and_capture(args : array<string>; var output : string&; timeout_sec : float = 0.0) : int {
//! Run an external command and capture its stdout+stderr (merged into one pipe by the underlying ``popen_argv``). Returns the process exit code;
//! -1 means the spawn itself failed. No shell is involved, but on Windows the program path (``args[0]``) must use backslashes —
//! CreateProcess does not resolve ``bin/daslang``-style forward-slash relative paths (probe-verified); ``replace(exe, "/", "\\")`` first.
//! -1 means the spawn itself failed. No shell is involved, and a forward-slash ``args[0]`` spawns on every platform (``popen_argv`` hands Windows the backslash spelling).
var captured : string
let exit_code = unsafe(popen_argv(args, timeout_sec, $(f) {
if (f != null) {
Expand Down
7 changes: 5 additions & 2 deletions modules/dasLLAMA/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,8 +55,11 @@ re-transcoding `$LCPP/src/unicode-data.cpp`).
MoE region split.
- `ARCHITECTURE_MEDIA.md` - sec.2.13-2.16: the padded tower GEMM widths, the family GPU hooks,
the tower weight lane, and the plain-Model ASR decoders.
- `ARCHITECTURE_MEASUREMENT.md` - sec.2.5, 2.10, 2.20: the benchmark rig, the tune gate, and the
sanctioned instrumentation rails.
- `ARCHITECTURE_MEASUREMENT.md` - sec.2.5, 2.10, 2.20, 2.21, 2.26-2.27: the benchmark rig, the tune
gate, the sanctioned instrumentation rails, kernel-race fidelity, the gemv's own tune seat, and
the CPU kernel bench's fixture conditions.
- `ARCHITECTURE_CPU_KERNELS.md` - sec.2.22-2.24: the sub-block-packed k3/k6 planes, the grid
formats' panel and row-group decodes, and the VBMI symbol lattice.

## 3. Inherited invariants

Expand Down
45 changes: 45 additions & 0 deletions modules/dasLLAMA/ARCHITECTURE_CPU_KERNELS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
# dasLLAMA architecture - the CPU kernel planes and decodes

Companion of `ARCHITECTURE.md` (contract: `../../ARCHITECTURE_COMMON.md`). Section 2 here continues
the mechanism numbering; each section is cited by the code that embodies it.

## 2. Mechanisms

### 2.22 The k3 and k6 planes are packed per sub-block {#kq-subblock-planes}

A k6 grp<mr> plane's qh columns `2blk` and `2blk + 1` carry one sub-block's four `j` sites, each as
a 2-bit field at bit `2j`; the disk byte `h*32 + half*16 + j*4 + t` feeds sub-blocks `4h..` at bit
`2b`. k3 packs the same way - qs columns `2blk + half` carry the sub-block's four `j` sites at `2j`,
and the hmask column `blk` carries its eight sites at bit `s` (lo at `j`, hi at `4 + j`). One
sub-block's decode then costs two loads for k6 and three for k3, and nothing loaded lives past it;
the row-interleaved disk order cost one load per `j`. The layout is CPU-flavor: the `.dlim` a box
bakes is for the hardware that runs it, so a CPU plane owes nothing to the GPU tiers' shapes.
`IMAGE_VERSION` 28 is this layout.

### 2.23 A grid format's CPU gemv decodes as a panel or as row groups {#grid-decode-forms}

Five formats (iq3s, iq3xxs, iq2s, iq2xs, iq2xxs) have two gemv decode forms. The PANEL form gathers
a superblock into an alloca panel first and reads packed positions through one dword load per
4-byte column, the four positions of a column sharing the load. The ROW-GROUP form composes a
weight-width vector straight from the grid words - width/64 rows x 8 weights, one u64 grid entry
per iq2 row - and reverts to byte loads, because the column dword read only pays inside the panel.
The sdot lattice always takes row groups; on x86 the panel's latency chain does not scale with the
core, so `x86-vnni512` takes row groups for iq2xxs and `x86-amx` for iq2xxs and iq3xxs, everything
else the panel. A VBMI seat (sec.2.24) takes row groups unconditionally.

### 2.24 The VBMI symbol lattice {#vbmi-lattice}

On a zmm VBMI target a grid block decodes as row groups through a symbol lattice. Every grid byte
comes from a tiny alphabet (three symbols for the iq2 family, eight for iq3), so the grid is baked
as two compact code planes - entry `e`'s low and high half, four 2-bit symbols each for iq2, two
3-bit for iq3 - plus the alphabet as a per-lane `vpshufb` table and `ksigns_iq2xs` whole (128 bytes:
exactly the two registers one `VPERMI2B` indexes). Per format the block's index bytes gather into one
64-lane vector per column (lane `r*4 + position`) and look up in the code planes - `VPERMI2B` per 128
entries, index bit 7 blends the pairs, the 9th and 10th index bits arrive as lane masks (iq3s and iq2s
from the row's qh byte, iq2xs from bit 0 of its u16 word's high byte). The row's sign bytes land in
the same lane layout: the plane's own column for iq3s and iq2s, one ksigns `VPERMI2B` over the 7-bit
codes for the rest. Per row group and weight octet a constant two-source shuffle places each row's
code bytes in its qword, `VPMULTISHIFTQB` spreads the symbols into bytes, one `vpshufb` maps them to
magnitudes, and the signs ride the activation copy as a mask `(x ^ m) - m`. The lattice row shares
its tile body and planes with the 512/mr16 row, so only the gemv differs - what the gemv's own seat
(`ARCHITECTURE_MEASUREMENT.md` sec.2.26) races.
29 changes: 29 additions & 0 deletions modules/dasLLAMA/ARCHITECTURE_MEASUREMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,3 +160,32 @@ no model, no tuner, every arm gated against a CPU plane-dequant oracle before it
chain every dispatch through ONE shared output buffer on purpose - the serialized regime is
the instrument's probe shape, imitating the reference tool it is compared against - and its
numbers reach the engine only through a human porting decision, never a minted crown.

### 2.26 The gemv takes its own tune seat {#gemv-seat}

A kq family's manifest entry is its tile-best row, and the gemv gets a SECOND entry when a different
row serves the streamed decode better. Only same-mr rows can differ, because the layout companion
pins the plane's interleave; of those the two best by tile time race, the winner takes the gemv only
by the margin over the tile winner's own gemv, and the incumbent keeps a tie. Every family's perm grid
therefore carries a 256-wide `mr = 16` alternate beside its 512-wide tile crown. The seat is decided
at the engine's decode shape - a DRAM-bound plane streamed by every lane through the engine's own
splitter - because the engine's row length moves the answer (k3 on Granite Rapids: the 256 seat wins
at n=2048 and loses at 14336 - `benchmarks/matmul/kq_kernel_bench.das`, tune mode, seats pinned, d=32768). The seat fixture is a 512-row build at the ffn width tiled 320 times,
past the largest L3 a socket lends a slice of, and the seat takes the MEDIAN of seven rounds: a round
that finds the plane in L3 must not crown it. In normal mode `llvm_tune` stamps a companion from its
own manifest entry when one exists and is a perm this box can run, else from the tile's.

### 2.27 The CPU kernel bench's fixture conditions {#cpu-kernel-bench-fixture}

`benchmarks/matmul/kq_kernel_bench.das` times raw kernels on synthetic planes, and three fixture
properties decide whether its numbers mean anything. Every plane of one format lives in ONE arena at
fixed offsets, staggered so no two starts share their low 12 address bits: the heap places separate
arrays at run-dependent relative addresses, and planes that alias in the L1/L2 set logic make a run's
time depend on where the heap put them. Scale planes are filled with a byte that is a normal number in
every scale form, never random bytes, because denormal math runs orders of magnitude slower. Each row
is warmed before it is timed - three unmeasured rounds solo, six dispatches per row on the team arm -
because a core ramps over several rounds and one warm call is not enough. The q8 row exists in two
flavors: f32 group scales (the engine's own quantization) and `q8s16` over binary16 scales - the
wscale_f16 rail a GGUF q8_0 tensor runs, and the like-for-like row against the reference's q8_0.
Provenance for every figure in this section: `benchmarks/matmul/kq_kernel_bench.das` under
`DAS_TUNE_MODE=tune`, one thread, its default `--fmt` / `-n` / `-d` shape.
22 changes: 22 additions & 0 deletions modules/dasLLAMA/HOW_TO_ADD_A_FORMAT.md
Original file line number Diff line number Diff line change
Expand Up @@ -260,6 +260,19 @@ loses cross-module inlining); the first run after a cache write pays one cold co
deser re-key); QUIRK 21 still applies to emitter edits. Numbers, caveats and the
invalidation ledger: `plans/jit_compile_time.md`.

**Kernel-loop invocation (adopted 2026-09-01):** a kernel spelling is raced without a model in
`benchmarks/matmul/kq_kernel_bench.das` - `DAS_TUNE_MODE=tune bin/Release/daslang.exe -jit
modules/dasLLAMA/benchmarks/matmul/kq_kernel_bench.das -- --fmt <fmt> --perm <substr>` times every
row of the tile's and gemv's `_variants()` registries at one thread on synthetic planes (seconds
per try), and the reference row at the same shape is the reference exe's `test-backend-ops perf
-o MUL_MAT -p "type_a=<type>,type_b=f32,m=4096,n=1,"` under `GGML_BENCH_THREADS=1`;
`harness/kernel_ladder.sh` runs both sides for every format and prints the box's ratio table. The
app run comes only after a spelling wins there. Procedure, fact base and work queue:
`plans/kernel_parity_pass.md`.
2026-09-01: the pass closed the M1 CPU at kernel level - all 32 ladder rows at or above the
reference (the grid decodes 1.16-1.51x, from 0.51-0.93x) - so the per-format M1 CPU tg tails stamped
below (0.51x-0.73x) predate the ARM row-group decode; re-stamp the vehicles before quoting them.

A real file whose every tensor type is now loadable (the header census script in the session
scratchpad, or `harness/gguf_dump.das`), through `examples/dasLLAMA/run.das` against
`simple_ids.exe` from the llama.cpp reference build for the same prompt; then `test_model_image`
Expand Down Expand Up @@ -554,6 +567,9 @@ only the step-3 0.0267 top-2 flip vs llama.cpp. M1 benches: CPU das 897.4/56.3 v
is CLOSED - and with it THE FORMAT LADDER: four-tier table zen2 2.86x/0.70x, vk 0.78x/0.70x,
M1 CPU 6.41x/0.57x, Metal 0.93x/0.78x.

2026-09-01 (kernel parity pass): zen2 16t (`benchmarks/lcpp_bench.das --for-debug-purposes`, debug-jit) vs the clean-cpu `llama-bench` on the IQ2_XXS-local 1B: pp512 506.9 vs
181.2 (2.80x), tg128 88.6 vs 86.2 (1.03x, was 0.70x) - the sign column + u64 pair gemv.

### IQ2_XS Phase A (CPU, 2026-08-31) - the ksigns u64 tier

Shape: 256-superblock grid format - each of the 32 u16 qs words carries a 9-bit index into
Expand Down Expand Up @@ -621,6 +637,9 @@ tiers. M1 16GB benches: CPU das 746.6/51.6 vs llama.cpp 144.8/101.1 (5.16x/0.51x
(0.93x/0.86x - the iq2s pp class). The format is CLOSED on all four tiers; four-tier table:
zen2 2.78x/0.70x, vk 0.77x/0.54x, M1 CPU 5.16x/0.51x, Metal 0.93x/0.86x.

2026-09-01 (kernel parity pass): zen2 16t (`benchmarks/lcpp_bench.das --for-debug-purposes`, debug-jit) vs the clean-cpu `llama-bench` on the IQ2_XS-local 1B: pp512 480.7 vs
174.6 (2.75x), tg128 93.8 vs 85.9 (1.09x, was 0.70x) - sign column + u64 pair + column dword read.

### IQ2_S Phase A (CPU, 2026-08-31) - the u64-grid tier

Shape: 256-superblock grid format, the first with a u64 grid - a 10-bit index (qs byte |
Expand Down Expand Up @@ -699,6 +718,7 @@ the mradermacher i1 vehicle - IQ2_S attn x32 + IQ3_XXS/IQ3_S/Q4_K/Q5_K):
| tier | pp512 das / llama.cpp | tg128 das / llama.cpp |
|---|---|---|
| zen2 CPU | 501.5 / 138.5 (3.62x) | 55.8 / 73.5 (0.76x) |
| zen2 CPU, 2026-09-01 kernel parity pass | 494.4 / 138.1 (3.58x) | 71.9 / 74.5 (0.97x, +-4.2) - sign column + u64 pair |
| 5060 Ti Vulkan | 12099.3 / 17377.5 (0.70x - the tier class) | 292.4 / 362.5 (0.81x) |
| M1 CPU | 883.3 / 413.4 (2.14x) | 53.7 / 73.9 (0.73x) |
| M1 Metal | 3170.5 / 3427.2 (0.93x) | 205.8 / 220.7 (0.93x) |
Expand Down Expand Up @@ -915,6 +935,7 @@ the local --tensor-type requant, iq3_xxs on attn_k/q + all ffn):
| tier | pp512 das / llama.cpp | tg128 das / llama.cpp |
|---|---|---|
| zen2 CPU | 507.2 / 136.0 (3.73x) | 56.7 / 72.6 (0.78x) |
| zen2 CPU, 2026-09-01 kernel parity pass | 501.0 / 135.8 (3.69x) | 74.9 / 74.3 (1.01x) - sign column + column dword read |
| 5060 Ti Vulkan | 12225.7 / 17807.7 (0.69x - the tier class) | 372.1 / 389.9 (0.95x) |
| M1 CPU | 906.0 / 410.5 (2.21x) | 53.5 / 74.0 (0.72x) |
| M1 Metal | 3224.0 / 3429.9 (0.94x) | 213.5 / 227.3 (0.94x) |
Expand Down Expand Up @@ -1017,6 +1038,7 @@ embedding head is Q6_K - three formats share every decode step):
| tier | pp512 das / llama.cpp | tg128 das / llama.cpp |
|---|---|---|
| zen2 CPU | 516.9 / 104.9 (4.93x) | 52.4 / 57.0 (0.92x) |
| zen2 CPU, 2026-09-01 kernel parity pass | 526.1 / 104.7 (5.03x) | 69.9 / 56.5 (1.24x) - the vector sign column gemv |
| 5060 Ti Vulkan | 12539.6 / 17865 (0.70x) | 288.1 / 324.2 (0.89x) |
| M1 CPU | 886.2 / 433.6 (2.04x) | 57.4 / 66.6 (0.86x) |
| M1 Metal | 3237.6 / 3344.3 (0.97x) | 199.4 / 209.0 (0.95x) |
Expand Down
Loading
Loading