Skip to content

ggml-cpu: vectorize sum-of-squares reduction in RMS-norm F32 kernel - #33

Open
codspeed-hq[bot] wants to merge 1 commit into
masterfrom
codspeed-optim-vectorize-the-sum-of-squares-reduction-in-the-fuse-1785230901711
Open

ggml-cpu: vectorize sum-of-squares reduction in RMS-norm F32 kernel#33
codspeed-hq[bot] wants to merge 1 commit into
masterfrom
codspeed-optim-vectorize-the-sum-of-squares-reduction-in-the-fuse-1785230901711

ggml-cpu: vectorize sum-of-squares reduction in RMS-norm F32 kernel

d8d6aa6
Select commit
Loading
Failed to load commit list.
CodSpeed HQ / CodSpeed Performance Analysis succeeded Jul 28, 2026 in 0s

Performance Gate Passed

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
✅ 27 untouched benchmarks

Performance Changes

Mode Benchmark BASE HEAD Efficiency
Simulation rms_norm 561.2 µs 464.4 µs +20.85%

Tip

Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.


Comparing codspeed-optim-vectorize-the-sum-of-squares-reduction-in-the-fuse-1785230901711 (d8d6aa6) with master (46819c9)

Open in CodSpeed