et-backend: Q4_0 matrix-engine GEMM rewrite for prefill#24
Open
RehanQasim-dev wants to merge 1 commit into
Open
et-backend: Q4_0 matrix-engine GEMM rewrite for prefill#24RehanQasim-dev wants to merge 1 commit into
RehanQasim-dev wants to merge 1 commit into
Conversation
RehanQasim-dev
marked this pull request as draft
July 23, 2026 11:13
RehanQasim-dev
force-pushed
the
upstream-q4_0-tensor-gemm
branch
from
July 23, 2026 11:16
55e06a7 to
082deb4
Compare
Static-reuse dequant across N-tiles with K-splitting and a materialized register-resident path for deep reuse, plus software-pipelined activation prefetch and hart-sync fixes. Adds flush_to_l2_multi to work around the 16-line cap of a single FlushVA. Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
RehanQasim-dev
force-pushed
the
upstream-q4_0-tensor-gemm
branch
from
July 23, 2026 11:21
082deb4 to
59ee193
Compare
RehanQasim-dev
marked this pull request as ready for review
July 23, 2026 11:59
marty1885
approved these changes
Jul 24, 2026
marty1885
left a comment
There was a problem hiding this comment.
Tested and I can replicate the performance. Decode is still coherent. Looks good!
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Rewrites the Q4_0 matrix-engine mul_mat kernel with static-reuse dequant across N-tiles, K-splitting and a materialized register-resident path for deep reuse, plus software-pipelined activation prefetch and hart-sync fixes. Adds flush_to_l2_multi to work around the 16-line cap of a single FlushVA.
Prefill improves roughly 3x on Llama-3.2-1B. Decode is unchanged since decode (N=1) always routes to the vecdot kernel, not the matrix-engine path this PR touches.
Additional information
Performance (Llama-3.2-1B-Instruct, ET-SoC-1):
Prefill t/s
Verified with llama-cli on ET-SoC-1 hardware (Llama-3.2-1B-Instruct).
Requirements