Skip to content

Update llama.cpp models and Gemma 4 chat templates - #36

Merged
addianto merged 10 commits into
masterfrom
feature/update-llamacpp-models
Jul 16, 2026
Merged

Update llama.cpp models and Gemma 4 chat templates#36
addianto merged 10 commits into
masterfrom
feature/update-llamacpp-models

Conversation

@addianto

Copy link
Copy Markdown
Owner

No description provided.

@addianto addianto self-assigned this Jul 16, 2026
Copilot AI review requested due to automatic review settings July 16, 2026 13:59

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates local llama.cpp presets to newer Gemma 4 / Qwen model entries and refreshes the vendored Gemma 4 chat templates, plus adds an idle-sleep flag to the llama-server start scripts.

Changes:

  • Update model preset .ini files (Gemma 4 repo entries, speculative drafting settings, and embedding/reranker model selections).
  • Vendor updated “canonical” Gemma 4 chat templates with revised tool-calling / turn-closure logic.
  • Add --sleep-idle-seconds to both the PowerShell and POSIX launcher scripts to reduce idle resource usage.

Reviewed changes

Copilot reviewed 5 out of 7 changed files in this pull request and generated 6 comments.

Show a summary per file
File Description
config/llama.cpp/rinze.ini Updates Gemma 4 E4B preset, adds speculative drafting config, and refreshes Qwen/embedding/reranker entries.
config/llama.cpp/alice.ini Updates Gemma 4 presets, increases context size for 26B preset, adds speculative drafting config, and refreshes Qwen/embedding/reranker entries.
config/llama.cpp/gemma-4-e4b-chat_template.jinja Updates Gemma 4 E4B chat template logic around tool-calls, thinking gating, and turn continuation.
config/llama.cpp/gemma-4-e2b-chat_template.jinja Updates Gemma 4 E2B chat template logic around tool-calls, thinking gating, and turn continuation.
config/llama.cpp/gemma-4-26b-a4b-chat_template.jinja Updates Gemma 4 26B A4B chat template logic around tool-calls, thinking gating, and turn continuation.
bin/Start-LlamaServer.ps1 Adds an idle sleep duration and passes --sleep-idle-seconds to llama-server.
bin/start-llama-server.sh Adds an idle sleep duration and passes --sleep-idle-seconds to llama-server.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread config/llama.cpp/gemma-4-e4b-chat_template.jinja
Comment thread config/llama.cpp/gemma-4-e4b-chat_template.jinja
Comment thread config/llama.cpp/gemma-4-e2b-chat_template.jinja
Comment thread config/llama.cpp/gemma-4-e2b-chat_template.jinja
Comment thread config/llama.cpp/gemma-4-26b-a4b-chat_template.jinja
Comment thread config/llama.cpp/gemma-4-26b-a4b-chat_template.jinja
@addianto
addianto merged commit ff23897 into master Jul 16, 2026
1 check passed
@addianto
addianto deleted the feature/update-llamacpp-models branch July 16, 2026 14:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants