b11399: CUDA: refactor swizzling code (#29612)
CUDA: refactor swizzling code fix templates/loop bounds
CUDA: refactor swizzling code fix templates/loop bounds
ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (#29806) ggml-cpu: vectorize BF16 K tails in tinyBLAS tests: Skip tinyBLAS when use_ref is enabled so CPU tests compare against the vec_dot path. ggml-cpu: vectorize tinyBLAS F16/F32 tails Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52638584 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple…
cuda : move neu_padded to where it is used (#29940) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52636582 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu…
ci : windows llvm build requires ninja multi-config (#29959) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52619546 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu…
add windows vulkan arm64 release add link
chat-peg-parser : clear current_tool when pending_tool_call is reset (#29942) A TOOL_ID node that arrives after TOOL_CLOSE wrote through current_tool, which still pointed into the just-destroyed pending_tool_call optional (use-after-free, then a second free of the id buffer). Clear the pointer on reset. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/526…
ci : set default permissions (#29945) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52595372 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 1…
cuda : move blocks_per_col to where it is used (#29939) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52591217 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ub…
CUDA: fix MMQ memory fault if n_expert >> n_ubatch (#29941) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52581427 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu…
vulkan: fix rdna4 mat_vec tuning (#29934) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52579133 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CU…
imatrix: calculate activation-based statistics for new format (GGUF) imatrices (#14891) Use activations to calculate the stats Determine calculation mode Compute entropy for activations Compute cosine similarity based on activations Compute l2 norm Add compute_layer_statistics() function Update aggregated statistic report layout Fix printing l2 norm when calc_mode = 1 Refactor variable name Comput…
spec : fix n-gram drafts rejected at temp > 0 after truncation (#29924) Co-authored-by: Pranesh Gonegandla pgonegandla@nvidia.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52563003 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU…
common : prepare load_from_models_dir() for path conversion (#29674) This is part of the fs::path modernization series. That was also the opportunity to remove fs_list(). Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52560765 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI e…
server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (#29938) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52557999 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ub…
ci : pushing tag needs deploy key (#29937) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52545484 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - C…
webgpu: add f16 support to fill/set_rows (#29897) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52497832 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA…
mtmd : fix deprecated strdup warning on Windows (#29863) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52479259 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) U…
vendor : update cpp-httplib to 0.59.0 (#29886) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52476117 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12)…
server : fix laya abort by limiting n_batch to n_ubatch (#29903) server : fix laya abort by limiting n_batch to n_ubatch Fixes #29902 Assisted-by: Claude fix(review) : rm tests, embeddings cond Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52435853 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…
common : add common_is_tty() helper and fix deprecated warnings on Windows (#29860) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52429962 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64…
chat : honor json_schema in Ling 3.0 parser (#29813) chat : honor json_schema in Ling 3.0 parser Ling 3.0 only built a grammar for tool calls and did not handle inputs.json_schema, so response_format requests were left unconstrained. Add an eager response-format grammar path with precedence over tools, following the existing parser patterns. Require before JSON when thinking is enabled and do not…
ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (#29904) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52418791 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan)…
graph: gather the recurrent states once so the reserve covers every split (#29856) build_rs gathered the extra states (n_rs - n_seqs rows) with their own get_rows. The worst-case reserve has n_rs == n_seqs, so that node was sized at zero rows, and any ubatch whose cells are not contiguous forced a graph reallocation at an unchanged node count, which aborts under GGML_SCHED_NO_REALLOC. A single get…
ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (#29852) ggml-openvino : Qwen3.5 MoE perf (#312) Squash of ravi9#312: ggml-openvino: add detailed inference profiling (Yu, Zijun) ggml-openvino: use remote output tensors by default (Yu, Zijun) ggml-openvino: optimize single-sequence recurrent state (Yu, Zijun) opt1: remove recurrent reset for single seque…
qwen4exp : halve the indexer score memory (#29825) qwen4exp : halve the indexer score memory The indexer scored all heads in one product and rectified a copy of it, so two [n_pool, n_idx_h, n_tokens] f32 tensors were live at once, the largest buffers of the graph at long context. Each head now gets its own product, rectified and summed in place into one [n_pool, n_tokens] score. qwen4exp: let the…
model: add support for clef decision model (text-only) (#29831) init support for clef (text only) more static graph clean up nits nits 2 Update gguf-py/gguf/constants.py Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/523856…
CUDA: fuse shared experts into MMVQ (#29184) CUDA: fuse shared experts into MMVQ check if buffer is null move stride_col_dst to fusion args Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52371643 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu…
spec : add probabilistic sampling for simple draft and MTP (#27694) Make the drafter probabilistic and the target verify by rejection sampling Drop stale spec_draft_q before drafting Fallback to argmax sampling for grammar-constrained requests and adding flag for enabling probabilistic draft sampling. Default flag value is greedy. Support grammar-constrained requests in rejection sampling Fix - re…
ggml-quants : avoid invalid rounding in qkx3 scale search (#29817) ggml-quants : avoid invalid rounding in qkx3 scale search The imatrix scale search can produce an infinite, NaN, or otherwise out-of-range value when the fitted minimum collapses to the maximum or makes the range extremely small. That value is then passed to nearest_int and can trip its assertion in Debug builds. Clamp the quantiza…
ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096) ggml-cpu : fix soft_max_back wrong output when dst aliases src1 GGML_OP_SOFT_MAX_BACK is listed in ggml_op_can_inplace, so the graph allocator may assign dst to alias either src0 (dy) or src1 (y). The result was built in several steps: ggml_vec_cpy_f32 (nc, dx, dy); ggml_vec_acc1_f32 (nc, dx, -dot_y_dy); ggml_vec_mul_f32 (nc,…