Sync master with upstream release b8149 by jan-service-account · Pull Request #435 · janhq/llama.cpp

jan-service-account · 2026-02-25T08:40:55Z

Updates dev branch with latest release (b8149) from ggml-org/llama.cpp

* ci : add metal server workflows * cont : try fix python init * cont : move to a separate workflow that runs only on master * cont : fix num jobs Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

* spec: remove parameter spec-ngram-check-rate * spec : renamed statistics vars * spec : add n_call_begin, n_call_accept * spec : don't enable key-map-stats

…-org#19457) * Log converting requests * Print as debug instead of info [no ci] --------- Co-authored-by: openingnow <>

* chat: fix case where template accepts type content only * rm stray log * reuse render_message_to_json

* cuda : extend GGML_OP_PAD to work with non-cont src0 * tests : add permuted pad

Implement ggml_cann_mul_mat_id_quant function to support quantized matrix multiplication for Mixture of Experts (MoE) architectures on CANN backend. Key features: - Support Q4_0 and Q8_0 quantized weight formats - Use IndexSelect to dynamically route expert-specific weights based on indices - Leverage WeightQuantBatchMatmulV2 for efficient quantized computation - Handle automatic F16 type conversion for hardware compatibility - Support both per-expert and broadcast input modes Implementation details: - Extract expert weights and scales using CANN IndexSelect operation - Process each batch and expert combination independently - Create proper tensor views with correct stride for matmul operations - Automatic input/output type casting to/from F16 as needed Testing: All test cases passed for supported types (F32, F16, Q4_0, Q8_0).

…-org#18968)

…xtModel (ggml-org#19445) * Add special case for Qwen3VLMoe * Fix down path, remove arrows and checkmarks * ws * Moved to Qwen3VL * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

…ion (ggml-org#19452) using noexcept std::filesystem::directory_entry::is_regular_file overload prevents abnormal termination upon throwing an error (as caused by symlinks to non-existent folders on linux) Resolves: ggml-org#18560

…ons (dotprod) (ggml-org#19360) * First working version of GEMM and GEMV * interleave loads and compute * Clang-format * Added missing fallback. Removed tested TODO. * Swap M and N to be consistent with the repack template convention

* support qwen3.5 series * remove deepstack for now, and some code clean * code clean * add FULL_ATTENTION_INTERVAL metadata * code clean * reorder v heads for linear attention to avoid expensive interleaved repeat

…9315) * Fix memory leaks in shader lib, backend, backend_context, buffer_context, and webgpu_buf_pool * Free pools * Cleanup * More cleanup * Run clang-format * Fix arg-parser and tokenizer test errors that free an unallocated buffer * Fix device lost callback to not print on device teardown * Fix include and run clang-format * remove unused unused * Update binary ops --------- Co-authored-by: Reese Levine <reeselevine1@gmail.com>

CCCL 3.2 has been released since it was added to llama.cpp as part of the backend-sampling PR, and it makes sense to update from RC to final released version. https://github.com/NVIDIA/cccl/releases/tag/v3.2.0

…19368) * llama : refactor sampling_info to use buffer_view template This commit updates the sampling_info struct in llama-context to use a buffer_view template for the logits, probs, sampled tokens, and candidates buffers. The motivation for this is to simplify the code, improve type safety and readability.

* tests : extend bin bcast for permuted src1 * cont : extend bin support * cont : s0 is always 1 * tests : simplify

Co-authored-by: thecaptain789 <thecaptain789@users.noreply.github.com>

* hexagon: add ARGSORT op Co-authored-by: Yarden Tal <yardent@qti.qualcomm.com> * hexagon: argsort reject tensors with huge rows for now * Adding support for DIV,SQR,SQRT,SUM_ROWS ops in hexagon backend * hexagon : Add GEGLU op * hexagon: fix editor config check * hexagon: rewrite and optimize binary ops ADD/SUB/MUL/DIV/ADD_ID to use DMA --------- Co-authored-by: Yarden Tal <yardent@qti.qualcomm.com> Co-authored-by: Manohara Hosakoppa Krishnamurthy <mhosakop@qti.qualcomm.com>

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

This commit updates an incorrect dSYMs where the the 's' was uppercase by mistake. The motivation for fixing this is that this can cause issues on case sensitive operating systems. Refs: ggml-org/whisper.cpp#3630

* Move dequant_model to after the text_config merge Add new kimi-k2.5 keys to mtmd convert Update V_MMPROJ tensor mapping for new mm_projector.proj keys Update V_M_IMP_NORM for new mm_projector.pre_norm key * Fix a couple of oversights * Add image support for Kimi-K2.5 * Revert changes to KimiVLForConditionalGeneration * Fix an assert crash * Fix permute swapping w / h on accident * Kimi-K2.5: Use merged QKV for vision * Kimi-K2.5: pre-convert vision QK to use build_rope_2d * Kimi-K2.5: support non-interleaved rope for vision * Kimi-K2.5: fix min / max pixel * Kimi-K2.5: remove v/o permutes, unnecessary * Kimi-K2.5: update permute name to match * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Kimi-K2.5: replace build_rope_2d ggml_cont with ggml_view_3d pointers --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

This commit removes two unused functions `common_lcp` and `common_lcs`. The last usage of these functions was removed in Commit 33eff40 ("server : vision support via libmtmd") and are no longer used anywhere in the codebase.

…g#19511) * ggml : unary ops support non-cont src0 * metal : support F16 unary ops + fix ELU

* opencl: add general q6_k mm * opencl: refine condition for q6_K mm * opencl: add general q4_K mv * opencl: fix whitespace

Signed-off-by: Adrien Gallouët <adrien@gallouet.fr>

…g#19433) This builds the following targets: * gfx1151 * gfx1150 * gfx1200 * gfx1201 * gfx1100 * gfx1101 * gfx1030 * gfx908 * gfx90a * gfx942

Also update architectures

…nt message (ggml-org#19773) * server : merge contiguous input items into a single assistant message * cont : simplify tool call msg * cont : reduce and combine content * cont : fix merging content items

* model: Add Kanana-2 model support * lint: adjust spacing

…l-org#19805) Co-authored-by: Jules LEIDELINGER <11395311+julio75012@users.noreply.github.com>

…8862) * llama : remove write/read of output ids/logits/embeddings This commit removes the write/read of output ids, logits and embeddings from the llama context state. Refs: ggml-org#18862 (comment) * completion : add replying of session state This commit updates the session handing in the completion tool to handle the that logits are no longer stored in the session file. Instead, we need to replay the last token to get the logits for sampling. * common : add common_prompt_batch_decode function This commit adds a new function which is responsible for decoding prompt and optionally handle the saving for session data. * update save-state.cpp to use llama_state_load_file This commit updates the save-load-state example to utilize the new llama_state_load_file function for loading the model state from a file. And it also replays the last token after loading since this state is now stored before the last token is processed. * examples : set n_seq_max = 2 for ctx3 This commit updates the save-load-state example to set the n_seq_max parameter to 2 when initializing the ctx3 context. The motivation for this change is that using 1 as n_parallel/n_seq_max the context only supports one sequence, but the test laster tries to use a second sequence which results in the following error: ```console main : loaded state with 4 tokens main : seq 0 copied, 225760 bytes main : kv cache cleared find_slot: seq_id=1 >= n_seq_max=1 Try using a bigger --parallel value state_read_meta: failed to find available cells in kv cache ``` This seems to only happen for recurrent/hybrid models.

…ons (dotprod) (ggml-org#19356) * Generic GEMV and boilerplate for q5_K dotprod * Generic GEMM and boilerplate for q5_K dotprod * ARM64 q5_K dotprod GEMM * ARM64 q5_K dotprod GEMV

…ml-org#19823) This commit replaces/merges the inspect-org-model.py script with the contents tensor-info.py script. The merged script has also been updated to also print tensor sizes which was the only thing that was not done before (by tensor-info.py that is). The motivation for this is that tensor-info.py does not load the tensor weights which can be time consuming for larger models. And also now that both are doing almost the same thing it makes sense to just have one and not two scripts to maintain.

…gml-org#19829)

…rg#19824) * tests : fix typos in comments in test-backend-sampler [no ci]

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

…gml-org#19835)

* hexagon: refactor set/get/sum-rows ops to use local context * hexagon: refactor ROPE and Softmax Ops to use local context Improves performance a bit by precomputing things and saving in the context. * hexagon: refactor activation ops to use local context struct * hexagon: refactor unary ops to use local context struct and DMA/VTCM * hexagon: use aligned hvx_scale function * hexagon: remove unused fields from op_context * hexagon: rewrite ROPE to use DMA and VTCM scratchpad * hex-rope: keep N rows in scratchpad (instead of just two) * hex-rope: introduce rowidx cache * hex-rope: remove unused fields * hex-rope: rewrite dma prefetch logic to allow for multi-row fetch/compute also removes the need for fastdiv. * hex-rope: minor formatting * hex-rope: use indices and unroll the loops * hex-rope: more updates to cleanup rope-block handling * hexagon: cleanup supported type/dims checks * hexagon: all reduce funcs replicated across lanes There is no need to explicitly replicate the first value. * snapdragon: update adb and windows scripts to use ubatch-size 256 Updated Ops support handles larger ubatches.

* vulkan: allow using fp16 in scalar flash attention shader * split rows inside of subgroups for faster synchronization * use row_split when Br >= 4, change reductions to use shared memory if row_split == 1 * use f32 scalar FA if f16 is not supported by device * fix amd workgroup size issue * optimize masksh use * add medium rows FA shader Br size * fixes * add padding to mask shmem buffer * cache q values into registers for KQ * fuse lf accumulation, pf and v accumulation into a loop * stage K loads through shmem * stage V loads through shmem * only stage through shmem on Nvidia * default to Bc 32 * also stage V through shmem when this is done for K * dynamic subgroups for intel * use vectorized stores * use float_type for dequantize4 functions * use smaller scalar rows size for smaller rows count * relax flash attention split_k condition to allow non-gqa use * use minimal subgroup size on Intel * fix shmem support function * fix rebase issues * fixes * Bc 4 for scalar FA is not a valid configuration * Use wave32 on AMD RDNA for scalar FA * add Intel shader core count lookup-table * fix regressions * device tuning * tmpsh size fix * fix editorconfig * refactor fa tuning logic into a single place * fix gqa opt logic * fix block_rows with small n_rows * amd tuning * fix hsk=72/80 issue * tuning * allow condition skipping for column check * use float16 for Of if available * address feedback * fix bad RDNA performance on head size <= 128 by limiting occupancy * allow printing pipeline stats * cleanup and fixes * limit occupancy for GCN for small batch FA with large HSK * disable f16 FA for GCN AMD GPUs on the proprietary driver

"max_tokens" is deprectated in favor of "max_completion_tokens" which sets the upper bound for reasoning+output token. Closes: ggml-org#13700

* model : Update label for LFM2-24B-A2B ``` ❯ build/bin/llama-bench -m /data/playground/checkpoints/LFM2-24B-A2B-Preview-Q4_0.gguf,/data/playground/checkpoints/LFM2-8B-A1B-Q4_0.gguf -p 1 -n 0 | model | size | params | backend | threads | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: | | lfm2moe 24B.A2B Q4_0 | 12.54 GiB | 23.84 B | CPU | 10 | pp1 | 30.35 ± 2.49 | | lfm2moe 8B.A1B Q4_0 | 4.41 GiB | 8.34 B | CPU | 10 | pp1 | 49.24 ± 1.93 | ``` * Remove extra line

* gguf : prevent integer overflow for ggml_context mem size * ggml : fix int overflows in ggml_new_object() * gguf : prevent string exhaustion * gguf : prevent array elements exhaustion * ggml : fix negative tensor type oob * py : assert that alignment is non-zero power of 2 * ggml : check int overflow in ggml_new_tensor_impl and ggml_new_object * gguf-py : error on duplicate keys when reading * py : restore tensor_fields * enforce proper alignment in add_custom_alignment * gguf : better name * gguf : fix ctx size for no_alloc == true * gguf : minor print fix * ggml : print values when overflow * ggml : remove deprecated ggml_type_sizef() * ggml : relax ggml_type asserts to debug-only * gguf : add mem_size overflow test * gguf : add file size check for arrays * ggml : relax asseerts for ggml_get_type_traits() * flake8 fix --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

…outer mode (ggml-org#19854) * server: fix query params lost when proxying requests in multi-model router mode * server: re-encode query params using httplib::encode_query_component in proxy

ggerganov and others added 30 commits February 9, 2026 15:09

spec : remove check rate (ggml-org#19377)

292f690

* spec: remove parameter spec-ngram-check-rate * spec : renamed statistics vars * spec : add n_call_begin, n_call_accept * spec : don't enable key-map-stats

Server: log when converting requests to chat completions format (ggml…

820ebfa

…-org#19457) * Log converting requests * Print as debug instead of info [no ci] --------- Co-authored-by: openingnow <>

mtmd: Implement tiling for LFM2-VL (ggml-org#19454)

262364e

chat: fix case where template accepts type content only (ggml-org#19419)

98e57ca

* chat: fix case where template accepts type content only * rm stray log * reuse render_message_to_json

cuda : extend GGML_OP_PAD to work with non-cont src0 (ggml-org#19429)

a0d5855

* cuda : extend GGML_OP_PAD to work with non-cont src0 * tests : add permuted pad

CANN: Remove unnecessary wrapper for gml_backend_buft_is_cann (ggml…

f0bfe54

…-org#18968)

tts : fix typos in README.md [no ci] (ggml-org#19463)

66d403c

test: fix IMROPE perf test case (ggml-org#19465)

9a96352

models : support qwen3.5 series (ggml-org#19468)

fc0fe40

* support qwen3.5 series * remove deepstack for now, and some code clean * code clean * add FULL_ATTENTION_INTERVAL metadata * code clean * reorder v heads for linear attention to avoid expensive interleaved repeat

CUDA : Update CCCL-tag for 3.2 to final release from RC (ggml-org#19486)

612db61

CCCL 3.2 has been released since it was added to llama.cpp as part of the backend-sampling PR, and it makes sense to update from RC to final released version. https://github.com/NVIDIA/cccl/releases/tag/v3.2.0

metal : consolidate unary ops (ggml-org#19490)

ceaa89b

ggml : extend bin bcast for permuted src1 (ggml-org#19484)

89181c0

* tests : extend bin bcast for permuted src1 * cont : extend bin support * cont : s0 is always 1 * tests : simplify

model : fix wavtokenizer embedding notions (ggml-org#19479)

6d95707

llama : correct typos 'occured' and 'occurences' (ggml-org#19414)

8ee538c

Co-authored-by: thecaptain789 <thecaptain789@users.noreply.github.com>

common : improve download error reporting (ggml-org#19491)

0c1f39a

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

docs: ban AI for issues and discussions [no CI] (ggml-org#19512)

ada90bf

metal : extend l2_norm support for non-cont src0 (ggml-org#19502)

9ab072e

build : fix case in dSYMs path for build-macos [no ci] (ggml-org#19515)

53de59f

This commit updates an incorrect dSYMs where the the 's' was uppercase by mistake. The motivation for fixing this is that this can cause issues on case sensitive operating systems. Refs: ggml-org/whisper.cpp#3630

ggml : unary ops support non-cont src0 + metal F16 unary ops (ggml-or…

914dde7

…g#19511) * ggml : unary ops support non-cont src0 * metal : support F16 unary ops + fix ELU

opencl: add general Q6_K mm and Q4_K mv (ggml-org#19347)

4d3daf8

* opencl: add general q6_k mm * opencl: refine condition for q6_K mm * opencl: add general q4_K mv * opencl: fix whitespace

angt and others added 28 commits February 21, 2026 19:12

vendor : update cpp-httplib to 0.33.1 (ggml-org#19778)

99156f3

Signed-off-by: Adrien Gallouët <adrien@gallouet.fr>

Add a build target to generate ROCm artifacts using ROCm 7.2 (ggml-or…

f75c4e8

…g#19433) This builds the following targets: * gfx1151 * gfx1150 * gfx1200 * gfx1201 * gfx1100 * gfx1101 * gfx1030 * gfx908 * gfx90a * gfx942

Update ROCm docker container to 7.2 release (ggml-org#19418)

3571565

Also update architectures

ci : fix rocm release path [no ci] (ggml-org#19784)

e877ad8

server : merge contiguous Responses input items into a single assista…

34ec1c3

…nt message (ggml-org#19773) * server : merge contiguous input items into a single assistant message * cont : simplify tool call msg * cont : reduce and combine content * cont : fix merging content items

ci : fix rocm archive name [no ci] (ggml-org#19808)

9f0684f

model : add Kanana-2 model support (ggml-org#19803)

ae2368e

* model: Add Kanana-2 model support * lint: adjust spacing

Fix wrong cli-argument in documentation (ggml-org#19804)

cacc371

common : fix improper trimming in XML parser on complete message (ggm…

ed48378

…l-org#19805) Co-authored-by: Jules LEIDELINGER <11395311+julio75012@users.noreply.github.com>

jinja: correct stats for tojson and string filters (ggml-org#19785)

5452d73

cli : provide model with text filename (ggml-org#19783)

e8e2616

ggml-cpu: arm64: q5_K repack gemm and gemv (and generic) implementati…

bc160d3

…ons (dotprod) (ggml-org#19356) * Generic GEMV and boilerplate for q5_K dotprod * Generic GEMM and boilerplate for q5_K dotprod * ARM64 q5_K dotprod GEMM * ARM64 q5_K dotprod GEMV

webui: Add setting to have full height Code Blocks in Chat Messages (g…

9051663

…gml-org#19829)

tests : fix typos in comments in test-backend-sampler [no ci] (ggml-o…

d8aeb65

…rg#19824) * tests : fix typos in comments in test-backend-sampler [no ci]

vendor : update cpp-httplib to 0.34.0 (ggml-org#19830)

b68a83e

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

feat: Add code blocks full height setting to parameter sync service (g…

5eb0ea3

…gml-org#19835)

vulkan: fix data race in mul_mat_id shader (ggml-org#19790)

3ea5360

vulkan: fix coopmat1 without bf16 support (ggml-org#19793)

8c2c010

server : support max_completion_tokens request property (ggml-org#19831)

c830f99

"max_tokens" is deprectated in favor of "max_completion_tokens" which sets the upper bound for reasoning+output token. Closes: ggml-org#13700

server: fix query params lost when proxying requests in multi-model r…

47eb12b

…outer mode (ggml-org#19854) * server: fix query params lost when proxying requests in multi-model router mode * server: re-encode query params using httplib::encode_query_component in proxy

models : fix graph splits (ggml-org#19866)

2446419

gguf : fix ftell/fseek for Windows (ggml-org#19870)

a96a112

jan-service-account merged commit f765a4a into dev Feb 25, 2026
3 checks passed

jan-service-account deleted the update-dev-from-master-2026-02-25-08-40 branch February 25, 2026 08:43

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Sync master with upstream release b8149#435

Sync master with upstream release b8149#435
jan-service-account merged 173 commits intodevfrom
update-dev-from-master-2026-02-25-08-40

jan-service-account commented Feb 25, 2026

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

20 participants

Conversation

jan-service-account commented Feb 25, 2026

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

20 participants