ggml-cpu: 2x2 register-blocked kernel for ggml_vec_dot_f32 (F32 matmul) - #34
Open
codspeed-hq[bot] wants to merge 1 commit into
CodSpeed HQ / CodSpeed Performance Analysis
succeeded
Jul 29, 2026 in 0s
Performance Gate Passed
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 3 improved benchmarks
✅ 25 untouched benchmarks
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | Simulation | ffn_swiglu |
1,067.7 ms | 858.1 ms | +24.42% |
| ⚡ | Simulation | mul_mat[f32] |
346.1 ms | 282.1 ms | +22.7% |
| ⚡ | WallTime | prompt_layer[f32] |
247.1 ms | 213.6 ms | +15.69% |
Tip
Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.
Comparing codspeed-optim-2-2-register-blocked-kernel-for-ggml-vec-dot-f32-f-1785295416304 (5a7c570) with master (46819c9)
Loading