Sync master with upstream release b8067 by jan-service-account · Pull Request #425 · janhq/llama.cpp

jan-service-account · 2026-02-16T00:48:21Z

Updates dev branch with latest release (b8067) from ggml-org/llama.cpp

* ci : add metal server workflows * cont : try fix python init * cont : move to a separate workflow that runs only on master * cont : fix num jobs Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

* spec: remove parameter spec-ngram-check-rate * spec : renamed statistics vars * spec : add n_call_begin, n_call_accept * spec : don't enable key-map-stats

…-org#19457) * Log converting requests * Print as debug instead of info [no ci] --------- Co-authored-by: openingnow <>

* chat: fix case where template accepts type content only * rm stray log * reuse render_message_to_json

* cuda : extend GGML_OP_PAD to work with non-cont src0 * tests : add permuted pad

Implement ggml_cann_mul_mat_id_quant function to support quantized matrix multiplication for Mixture of Experts (MoE) architectures on CANN backend. Key features: - Support Q4_0 and Q8_0 quantized weight formats - Use IndexSelect to dynamically route expert-specific weights based on indices - Leverage WeightQuantBatchMatmulV2 for efficient quantized computation - Handle automatic F16 type conversion for hardware compatibility - Support both per-expert and broadcast input modes Implementation details: - Extract expert weights and scales using CANN IndexSelect operation - Process each batch and expert combination independently - Create proper tensor views with correct stride for matmul operations - Automatic input/output type casting to/from F16 as needed Testing: All test cases passed for supported types (F32, F16, Q4_0, Q8_0).

…-org#18968)

…xtModel (ggml-org#19445) * Add special case for Qwen3VLMoe * Fix down path, remove arrows and checkmarks * ws * Moved to Qwen3VL * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

…ion (ggml-org#19452) using noexcept std::filesystem::directory_entry::is_regular_file overload prevents abnormal termination upon throwing an error (as caused by symlinks to non-existent folders on linux) Resolves: ggml-org#18560

…ons (dotprod) (ggml-org#19360) * First working version of GEMM and GEMV * interleave loads and compute * Clang-format * Added missing fallback. Removed tested TODO. * Swap M and N to be consistent with the repack template convention

* support qwen3.5 series * remove deepstack for now, and some code clean * code clean * add FULL_ATTENTION_INTERVAL metadata * code clean * reorder v heads for linear attention to avoid expensive interleaved repeat

…9315) * Fix memory leaks in shader lib, backend, backend_context, buffer_context, and webgpu_buf_pool * Free pools * Cleanup * More cleanup * Run clang-format * Fix arg-parser and tokenizer test errors that free an unallocated buffer * Fix device lost callback to not print on device teardown * Fix include and run clang-format * remove unused unused * Update binary ops --------- Co-authored-by: Reese Levine <reeselevine1@gmail.com>

CCCL 3.2 has been released since it was added to llama.cpp as part of the backend-sampling PR, and it makes sense to update from RC to final released version. https://github.com/NVIDIA/cccl/releases/tag/v3.2.0

…19368) * llama : refactor sampling_info to use buffer_view template This commit updates the sampling_info struct in llama-context to use a buffer_view template for the logits, probs, sampled tokens, and candidates buffers. The motivation for this is to simplify the code, improve type safety and readability.

* tests : extend bin bcast for permuted src1 * cont : extend bin support * cont : s0 is always 1 * tests : simplify

Co-authored-by: thecaptain789 <thecaptain789@users.noreply.github.com>

* hexagon: add ARGSORT op Co-authored-by: Yarden Tal <yardent@qti.qualcomm.com> * hexagon: argsort reject tensors with huge rows for now * Adding support for DIV,SQR,SQRT,SUM_ROWS ops in hexagon backend * hexagon : Add GEGLU op * hexagon: fix editor config check * hexagon: rewrite and optimize binary ops ADD/SUB/MUL/DIV/ADD_ID to use DMA --------- Co-authored-by: Yarden Tal <yardent@qti.qualcomm.com> Co-authored-by: Manohara Hosakoppa Krishnamurthy <mhosakop@qti.qualcomm.com>

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

This commit updates an incorrect dSYMs where the the 's' was uppercase by mistake. The motivation for fixing this is that this can cause issues on case sensitive operating systems. Refs: ggml-org/whisper.cpp#3630

* Move dequant_model to after the text_config merge Add new kimi-k2.5 keys to mtmd convert Update V_MMPROJ tensor mapping for new mm_projector.proj keys Update V_M_IMP_NORM for new mm_projector.pre_norm key * Fix a couple of oversights * Add image support for Kimi-K2.5 * Revert changes to KimiVLForConditionalGeneration * Fix an assert crash * Fix permute swapping w / h on accident * Kimi-K2.5: Use merged QKV for vision * Kimi-K2.5: pre-convert vision QK to use build_rope_2d * Kimi-K2.5: support non-interleaved rope for vision * Kimi-K2.5: fix min / max pixel * Kimi-K2.5: remove v/o permutes, unnecessary * Kimi-K2.5: update permute name to match * Update convert_hf_to_gguf.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Kimi-K2.5: replace build_rope_2d ggml_cont with ggml_view_3d pointers --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

This commit removes two unused functions `common_lcp` and `common_lcs`. The last usage of these functions was removed in Commit 33eff40 ("server : vision support via libmtmd") and are no longer used anywhere in the codebase.

…g#19511) * ggml : unary ops support non-cont src0 * metal : support F16 unary ops + fix ELU

* opencl: add general q6_k mm * opencl: refine condition for q6_K mm * opencl: add general q4_K mv * opencl: fix whitespace

…gml-org#19583) * ggml-hexagon: fa improvements ggml-hexagon: optimize flash attention calculations with improved variable handling ggml-hexagon: streamline flash attention operations by removing redundant checks for FP32 ggml-hexagon: optimize hvx_dot_f16_f16_aa_rx2 by simplifying variable handling for unused elements ggml-hexagon: optimize flash attention by changing slope vector type to F16 * hexfa: fixed test-backend-ops failurs due to leftover element handling * hexagon: refactor and optimize fa to use local context struct * ggml-hexagon: optimize flash-attention using hvx_vec_expf Use HVX for online softmax. --------- Co-authored-by: chraac <chraac@gmail.com>

This commit allows Qualcomm native vulkan driver to be used on Windows instead of Mesa Dozen.

Run libtool via xcrun like strip and dsymutil, to have proper tool resolution. Signed-off-by: Adrien Gallouët <angt@huggingface.co>

* scripts : use official split.py for cpp-httplib Using the official script is safer and ensures the generated code aligns with the library's standards. Signed-off-by: Adrien Gallouët <angt@huggingface.co> * Catch generic errors Signed-off-by: Adrien Gallouët <angt@huggingface.co> * Allow print() Signed-off-by: Adrien Gallouët <angt@huggingface.co> * Ensure robust cleanup Signed-off-by: Adrien Gallouët <angt@huggingface.co> --------- Signed-off-by: Adrien Gallouët <angt@huggingface.co>

* ggml: added cleanups in ggml_quantize_free Add missing cleanup calls for IQ2_S, IQ1_M quantization types and IQ3XS with 512 blocks during quantization cleanup. * mmap: Fix Windows handle lifetime Move hMapping from local variable to member variable so it stays alive for the entire lifetime of the mapping. The file mapping handle must remain valid until UnmapViewOfFile is called. Fixes cleanup order in destructor. * Update llama-mmap.cpp * Update llama-mmap.cpp Remove trailing whitespace from line 567

* Refactoring to use new llama_put_adapter_loras * cont : alternative lora API --------- Co-authored-by: Jake Chavis <jakechavis6@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

last_graph is only available without OpenMP, but ggml_graph_compute_thread() is called in both cases. Signed-off-by: Adrien Gallouët <angt@huggingface.co>

* models : optimizing qwen3next graph * cont * wip * wip * wip * wip * wip * wip * wip * wip * wip * wip * cont : remove redundant q, g chunking * minor * minor * avoid passing masks around * avoid concats during chunking * naming + shapes * update names and use prefix to disable CUDA graphs

* nemotron nano v2 vlm support added * simplified code; addressed reviews * pre-downsample position embeddings during GGUF conversion for fixed input size

* ensure all models handle new experts count * revert removal for PhiMoeModel, does not inherit from base

…ml-org#19581) * cmake: fix KleidiAI install target failure with EXCLUDE_FROM_ALL Fix for the bug ggml-org#19501 by adding EXCLUDE_FROM_ALL to FetchContent_Declare. This properly excludes KleidiAI from both build and install targets, preventing install failures when GGML_CPU_KLEIDIAI=ON is used. The KleidiAI source files are still compiled into libggml-cpu.so, preserving all functionality. * addressed code review comments

* ggml-cpu: FA add GEMM microkernel * add guard for sizeless vector types * fix case where DV % GGML_F32_EPR !=0 * move memset out of the loop * move another memset out of the loop * use RM=4 for arm * simd_gemm: convert everything to int * convert everything to size_t to avoid warnings * fixup * add pragma for ignoring aggressive loop optimizations

This commit addresses a build issue with the KleidiAI backend when building multiple cpu backends. Commmit 3a00c98 ("cmake : fix KleidiAI install target failure with EXCLUDE_FROM_ALL") introduced a change where FetchContent_Populate is called instead of FetchContent_MakeAvailable, where the latter does handle this case (it is idempotent but FetchContent_Populate is not). I missed this during my review and I should not have commited without verifying the CI failure, sorry about that.

This option was introduced as a workaround because cpp-httplib could not build on visionOS. Since it has been fixed and now compiles on all platforms, we can remove it and simplify many things. Signed-off-by: Adrien Gallouët <angt@huggingface.co>

* cuda: optimize iq2xxs/iq2xs/iq3xxs dequantization - load all 8 int8 for a grid position in one load - calculate signs via popcnt instead of fetching from ksigns table - broadcast signs to drop individual shift/mask * cuda: iq2xxs: simplify sum scaling express `(sum * scale + sum / 2) / 4` as `(sum * (scale * 2 + 1)) / 8` express `((aux32 >> 28) * 2 + 1)` as `(aux32 >> 27 | 1)` saves 3 registers for mul_mat_vec_q (152 -> 149) according to nsight AFAICT no overflow can occur here as iq2xxs values are far too small * uint -> uint32_t error: identifier "uint" is undefined

ggerganov and others added 30 commits February 9, 2026 15:09

spec : remove check rate (ggml-org#19377)

292f690

* spec: remove parameter spec-ngram-check-rate * spec : renamed statistics vars * spec : add n_call_begin, n_call_accept * spec : don't enable key-map-stats

Server: log when converting requests to chat completions format (ggml…

820ebfa

…-org#19457) * Log converting requests * Print as debug instead of info [no ci] --------- Co-authored-by: openingnow <>

mtmd: Implement tiling for LFM2-VL (ggml-org#19454)

262364e

chat: fix case where template accepts type content only (ggml-org#19419)

98e57ca

* chat: fix case where template accepts type content only * rm stray log * reuse render_message_to_json

cuda : extend GGML_OP_PAD to work with non-cont src0 (ggml-org#19429)

a0d5855

* cuda : extend GGML_OP_PAD to work with non-cont src0 * tests : add permuted pad

CANN: Remove unnecessary wrapper for gml_backend_buft_is_cann (ggml…

f0bfe54

…-org#18968)

tts : fix typos in README.md [no ci] (ggml-org#19463)

66d403c

test: fix IMROPE perf test case (ggml-org#19465)

9a96352

models : support qwen3.5 series (ggml-org#19468)

fc0fe40

* support qwen3.5 series * remove deepstack for now, and some code clean * code clean * add FULL_ATTENTION_INTERVAL metadata * code clean * reorder v heads for linear attention to avoid expensive interleaved repeat

CUDA : Update CCCL-tag for 3.2 to final release from RC (ggml-org#19486)

612db61

CCCL 3.2 has been released since it was added to llama.cpp as part of the backend-sampling PR, and it makes sense to update from RC to final released version. https://github.com/NVIDIA/cccl/releases/tag/v3.2.0

metal : consolidate unary ops (ggml-org#19490)

ceaa89b

ggml : extend bin bcast for permuted src1 (ggml-org#19484)

89181c0

* tests : extend bin bcast for permuted src1 * cont : extend bin support * cont : s0 is always 1 * tests : simplify

model : fix wavtokenizer embedding notions (ggml-org#19479)

6d95707

llama : correct typos 'occured' and 'occurences' (ggml-org#19414)

8ee538c

Co-authored-by: thecaptain789 <thecaptain789@users.noreply.github.com>

common : improve download error reporting (ggml-org#19491)

0c1f39a

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

docs: ban AI for issues and discussions [no CI] (ggml-org#19512)

ada90bf

metal : extend l2_norm support for non-cont src0 (ggml-org#19502)

9ab072e

build : fix case in dSYMs path for build-macos [no ci] (ggml-org#19515)

53de59f

This commit updates an incorrect dSYMs where the the 's' was uppercase by mistake. The motivation for fixing this is that this can cause issues on case sensitive operating systems. Refs: ggml-org/whisper.cpp#3630

ggml : unary ops support non-cont src0 + metal F16 unary ops (ggml-or…

914dde7

…g#19511) * ggml : unary ops support non-cont src0 * metal : support F16 unary ops + fix ELU

opencl: add general Q6_K mm and Q4_K mv (ggml-org#19347)

4d3daf8

* opencl: add general q6_k mm * opencl: refine condition for q6_K mm * opencl: add general q4_K mv * opencl: fix whitespace

max-krasnyansky and others added 28 commits February 13, 2026 16:27

vulkan: Add vendor id for Qualcomm drivers (ggml-org#19569)

2dec548

This commit allows Qualcomm native vulkan driver to be used on Windows instead of Mesa Dozen.

vulkan: support GGML_OP_SET (ggml-org#19584)

53aef25

vulkan: support L2_NORM with contiguous rows (ggml-org#19604)

dbb0233

build : fix libtool call in build-xcframework.sh (ggml-org#19605)

91ea5d6

Run libtool via xcrun like strip and dsymutil, to have proper tool resolution. Signed-off-by: Adrien Gallouët <angt@huggingface.co>

convert : store ffn_gate_inp_shexp as F32 (ggml-org#19606)

0d00ef6

metal : fix ACC op (ggml-org#19427)

6e473fb

llama : update LoRA API. + fix excessive graph reserves (ggml-org#19280)

2d8015e

* Refactoring to use new llama_put_adapter_loras * cont : alternative lora API --------- Co-authored-by: Jake Chavis <jakechavis6@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

webui: Architecture and UI improvements (ggml-org#19596)

baa12f3

NetBSD build support (ggml-org#19589)

badba89

ggml : fix GGML_DEBUG with OpenMP (ggml-org#19599)

b7742cf

last_graph is only available without OpenMP, but ggml_graph_compute_thread() is called in both cases. Signed-off-by: Adrien Gallouët <angt@huggingface.co>

mtmd : Add Nemotron Nano 12B v2 VL support (ggml-org#19547)

01d8eaa

* nemotron nano v2 vlm support added * simplified code; addressed reviews * pre-downsample position embeddings during GGUF conversion for fixed input size

convert : ensure all models handle new experts count (ggml-org#19621)

079feab

* ensure all models handle new experts count * revert removal for PhiMoeModel, does not inherit from base

ggml-cpu: optimize ggml_vec_dot_bf16 for s390x (ggml-org#19399)

184c694

ggml : avoid UB in gemm ukernel (ggml-org#19642)

08e6d91

context : fix output reorder with backend sampling (ggml-org#19638)

341bc7d

docs: update s390x build docs (ggml-org#19643)

6e67fd2

ggml : bump version to 0.9.6 (ggml/1423)

1a8c700

ggml : bump version to 0.9.7 (ggml/1425)

55d5859

sync : ggml

ff4affb

jan-service-account merged commit ff4affb into dev Feb 25, 2026
5 of 9 checks passed

jan-service-account deleted the update-dev-from-master-2026-02-16-00-48 branch February 25, 2026 08:43

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Sync master with upstream release b8067#425

Sync master with upstream release b8067#425
jan-service-account merged 91 commits intodevfrom
update-dev-from-master-2026-02-16-00-48

jan-service-account commented Feb 16, 2026

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

20 participants

Conversation

jan-service-account commented Feb 16, 2026

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

20 participants