Skip to content

DRAFT: ggml: optimize Flash Next inference paths - #39

Draft
gaetan-puleo wants to merge 1 commit into
masterfrom
exp/qwen-38-next-iq4xs-flash-rocm-vulkan-max-kernel-improvements
Draft

gaetan-puleo wants to merge 1 commit into
masterfrom
exp/qwen-38-next-iq4xs-flash-rocm-vulkan-max-kernel-improvements

Conversation

@gaetan-puleo

Copy link
Copy Markdown
Collaborator

Overview

Measurements

Device:     Ryzen AI Max+ 395 / <mini-PC or laptop model>
Memory:     128 GB LPDDR5X-8000
Power:      <sustained TDP / platform profile / power governor>
BIOS:       UMA split <n> GB
Kernel:     <uname -r>, amdgpu params: <gttsize / ttm.pages_limit, or none>
Backend:    Vulkan RADV, Mesa <version>   (or: ROCm <version>, HIP_LAUNCH_BLOCKING=<0|1>)
Build:      <cmake flags>
Baseline:   <merge-base commit sha, built and run in this same session>
Change:     <your branch commit sha>
Model:      <HF repo / file>, <quant>

Baseline:

After:

Correctness:

Additional information

Requirements

  • I have read and agree with the contributing guidelines
  • This change is Strix Halo specific, or justified by measurements on Strix Halo. General llama.cpp improvements
    belong in halo-box/llama.cpp instead
  • AI usage disclosure:
  • What was NOT verified:

@gaetan-puleo gaetan-puleo changed the title ggml: optimize Flash Next inference paths DRAFT: ggml: optimize Flash Next inference paths Sep 10, 2026
@gaetan-puleo
gaetan-puleo marked this pull request as draft September 10, 2026 11:14

@dzannotti dzannotti left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validation: static whitespace checks pass. I did not run the implementation because the PR is draft, based on an old revision, is marked dirty/merge-conflicted, and currently has a failed Windows check. Those are prerequisites to meaningful validation.

Scope: the work is broad generic llama.cpp functionality, not a Strix-only patch. Rebase onto current upstream/staging, resolve the merge state, split independent concerns, and provide CPU/CUDA/Vulkan correctness coverage before requesting a new review.

comment generated by my clanker Codex

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants