L

llama.cpp

L
llama.cpp AI

b11398

ggml-cpu: x86 (#29806) での tinyBLAS での BF16/FP16/FP32 K テイルをサポート ggml-cpu: tinyBLAS テストで BF16 K テイルをベクトリ化: use_ref が有効になっているとき tinyBLAS をスキップして、CPU テストを vec_dot パスと比較する。ggml-cpu: tinyBLAS F16/F32 の尾をベクトリ化 ウェブサイト: https://llama.app 証明: https://github.com/ggml-org/llama.cpp/attestations/52638584 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple...

L
llama.cpp AI

b11397

cuda: move neu_padded to where it is used (#29940) サインオフ:Adrien Gallouët angt@huggingface.co ウェブサイト: https://llama.app 認定: https://github.com/ggml-org/llama.cpp/attestations/52636582 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64, KleidiAI有効) ディSABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu...

L
llama.cpp AI

b11396

ci: windows llvmビルドにはninjaマルチコンフィギュレーション (#29959) が必要です ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52619546 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64,KleidiAI有効) ディSABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu...

L
llama.cpp AI

b11393

chat-peg-parser : pending_tool_call がリセットされたときに current_tool をクリアする (#29942) TOOL_ID ノードが、TOOL_CLOSE が current_tool を介して書き込んだ後に到着し、まだ破壊された pending_tool_call オプション (使用後フリー、その後 id バッファーの2つ目のフリー) に指向します。リセットでポインタを消すウェブサイト: https://llama.app 証明: https://github.com/ggml-org/llama.cpp/attestations/526...

L
llama.cpp AI

b11392

ci: デフォルトの権限を設定 (#29945) ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52595372 macOS/iOS: macOS アップル シリコン (arm64) macOS アップル シリコン (arm64, KleidiAI有効) ディSABLED macOS インテル (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 1...

L
llama.cpp AI

b11391

cuda: blocks_per_col を使用する場所へ移動する (#29939) サインオフ: Adrien Gallouët angt@huggingface.co ウェブサイト: https://llama.app 証明: https://github.com/ggml-org/llama.cpp/attestations/52591217 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64, KleidiAI有効) 障害 macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu s390x (CPU) Ubuntu...

L
llama.cpp AI

b11390

CUDA: n_expert >> n_ubatch (#29941) の場合MMQメモリエラーを修正する ウェブサイト: https://llama.app 証明: https://github.com/ggml-org/llama.cpp/attestations/52581427 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64, KleidiAI有効) 障害 macOS Intel (x64) iOS XFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu...

L
llama.cpp AI

b11389

vulkan: fix rdna4 mat_vec tuning (#29934) ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52579133 macOS/iOS: macOS アップル シリコン (arm64) macOS アップル シリコン (arm64, KleidiAI enabled) ディSABLED macOS インテル (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) x64 (CUDA 12) - CU...

L
llama.cpp AI

b11388

imatrix:新しい形式 (GGUF) のimatrices (#14891) のアクティベーションベースの統計を計算する アクティベーションを使用して統計を計算する 計算モードを決定する アクティベーションのためのエントロピーを計算する アクティベーションに基づくコシノス類似性を計算する 計算 l2 ノルム 計算_レイヤ_統計を追加する (,) 関数 総統計レポートレイアウトを更新 calc_mode = 1 のとき l2 ノルムの印刷を修正する リファクタ変数名 計算。..

L
llama.cpp AI

b11387

仕様: 短縮後 temp > 0 で拒否された n-グラムのドラフトを修正する (#29924) 共著者: Pranesh Gonegandla pgonegandla@nvidia.com ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52563003 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64, KleidiAI有効) 障害 macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) arm64 (CPU...

L
llama.cpp AI

b11386

common: prepare load_from_models_dir() for path conversion (#29674) これはfs::path近代化シリーズの一部です。fs_listを削除する機会もありました。署名: アドリアン・ガルーエ angt@huggingface.co ウェブサイト: https://llama.app 認定: https://github.com/ggml-org/llama.cpp/attestations/52560765 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64, KleidiAI e...

L
llama.cpp AI

b11385

サーバー: 既定の許可リスト (#29938) で死んだLLAMA_ARG_HF_REPO_FILEキーを修正 署名: アドリアン・ガルーエ angt@huggingface.co ウェブサイト: https://llama.app 認定: https://github.com/ggml-org/llama.cpp/attestations/52557999 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64, KleidiAI有効) 障害 macOS Intel (x64) iOS XFramework Linux: Ubuntu x64 (CPU) Ub64 (CPU) Ubuntu...

L
llama.cpp AI

b11384

ci: pushing tag needs deploy key (#29937) ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52545484 macOS/iOS: macOS アップル シリコン (arm64) macOS アップル シリコン (arm64, KleidiAI enabled) ディSABLED macOS インテル (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) arm64 (Vulkan) Ubuntu x64 (CUDA 12) - C...

L
llama.cpp AI

b11382

webgpu: fill/set_rows (#29897) にf16のサポートを追加する ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52497832 macOS/iOS: macOS アップル シリコン (arm64) macOS アップル シリコン (arm64, KleidiAI有効) ディSABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (VulkanCU) Ubuntu x64 (DA...

L
llama.cpp AI

b11381

mtmd : Windows (#29863) で使用停止されたstrdup警告を修正 アドリアン・ガルーエット (Adrien Gallouët) の署名: angt@huggingface.co ウェブサイト: https://llama.app 認証: https://github.com/ggml-org/llama.cpp/attestations/52479259 macOS/iOS: macOS アップル シリコン (arm64) macOS アップル シリコン (arm64, KleidiAI有効) 禁用 macOS インテル (x64) iOS XCフレームワーク Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) s390x (CPU) U...

L
llama.cpp AI

b11380

販売者: cpp-httplib を 0.59.0 に更新 (#29886) ウェブサイト: https://llama.app 証明: https://github.com/ggml-org/llama.cpp/attestations/52476117 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64, KleidiAI有効) 障害 macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12)...

L
llama.cpp AI

b11379

サーバー: n_batch を n_batch に限定して laya を中止する修正 (#29903) サーバー: n_batch を n_batch に制限して laya を中止する修正 #29902 サポート: Claude fix(レビュー) : rm テスト、埋め込み cond ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52435853 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64,DIAI有効) KleidiSABLED macOS Intel...

L
llama.cpp AI

b11378

common: add common_is_tty() helper and fix deprecated warnings on Windows (#29860) サインオフ: Adrien Gallouët angt@huggingface.co ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52429962 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64,KleidiAI有効) 障害 macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64...

L
llama.cpp AI

b11377

chat: honor json_schema in Ling 3.0 parser (#29813) chat: honor json_schema in Ling 3.0 parser Ling 3.0はツールコール用の文法のみを構築し、inputs.json_schemaを処理しなかったため、response_format要求は制限されていませんでした。既存のパーサーパターンに従って、ツールよりも優先順位を持つ熱心なレスポンス形式の文法経路を追加します。JSON の前に要求します。 思考が有効になったとき、

L
llama.cpp AI

b11376

ci: ADD_ADD f16 のフラッシュを修正する。 合併された ADD 耐性 (#29904) を使用する。 ウェブサイト: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52418791 macOS/iOS: macOS Apple シリコン (arm64) macOS Apple シリコン (arm64, KleidiAI有効) 障害 macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan)...

L
llama.cpp AI

b11375

グラフ: 繰り返し状態を1回集めて、そのため、リザーブはすべての分割 (#29856) をカバーします。 build_rs は、余分な状態 (n_rs - n_seqs 行) をそれぞれの get_row で集めた。最悪のケースのリザーブは n_rs == n_seqs を有し、ノードが0行にサイズされ、細胞が隣接していない任意のババッチは、GGML_SCHED_NO_REALLOCの下で中止される変化のないノード数でグラフ再配置を強制した。一回だけ。..

L
llama.cpp AI

b11374

ggml-openvino: 2026.4.1に更新し、パフォーマンスを最適化し、オペレーションを拡張し、デバイスリストを改善します。(#29852) ggml-openvino : Qwen3.5 MoE perf (#312) ravi9#312のスクワッシュ: ggml-openvino: 詳細な推論プロファイリングを追加 (Yu, Zijun) ggml-openvino: デフォルトでリモート出力テンサーを使用 (Yu, Zijun) ggml-openvino: 単一シーケンスリキュレント状態を最適化 (Yu, Zijun) opt1: 単一シーケンスのためのリキュレントリセットを削除。..

L
llama.cpp AI

b11372

qwen4 :exp 半減インデクサースコアメモリ (#29825) qwen4 exp:半減インデクサースコアメモリ インデクサーは1つのプロダクトのすべてのヘッドをスコアし、そのコピーを修正したため、2つの [n_pool,n_idx_h,n_tokens] f32テンソーは一度に起動し、長い文脈でグラフの最大のバッファです。各ヘッドが自分の製品を得て、直し、 [n_pool,n_tokens] スコアにまとめられます。QWEN4EXP: ほら。..

L
llama.cpp AI

b11371

モデル: キー決定モデル (テキストのみ) のサポートを追加 (#29831) init キー (テキストのみ) のサポート より静的グラフ ニットニット 2をクリーンアップ gguf-py/gguf/constants.py 共同作成者: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co 共同作成者: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co ウェブサイト: https://llama.appestations: https://github.com/ggml-org/llama.cpp/attestations/523856...

L
llama.cpp AI

b11370

CUDA: MMVQ (#29184) に共有専門家をフィューズする CUDA: MMVQ に共有専門家をフィューズする バッファがゼロかどうかをチェックする 移転stride_col_dst を fusion args に 移動する ウェブサイト: https://llama.app 証明書: https://github.com/ggml-org/llama.cpp/attestations/52371643 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI有効) 障害 macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu...

L
llama.cpp AI

b11368

spec: simple draft と MTP (#27694) の確率サンプリングを追加する 抽出者が確率サンプリングで、ターゲットは拒絶サンプリングで検証する 抽出する前に stale spec_draft_q を落とす Fallback to argmax 文法制限要求のサンプリングと確率サンプリングを可能にするフラグを追加するデフォルトのフラグ値は、貪欲です。文法に制限された要求を拒絶サンプルでサポートする 修正 - re...

L
llama.cpp AI

b11366

ggml-quants: qkx3スケール検索 (#29817) で無効な丸めを避ける ggml-quants: qkx3スケール検索で無効な丸めを避ける イマトリックススケール検索は、適正最小値が最大値に崩壊するか、範囲が非常に小さい場合に無限、NaN,または他の範囲外の値を生成することができます。その値は、次に、最も近い_int に渡され、デバッグビルドでアサーションをトリップできます。クアンティザを握って。..

L
llama.cpp AI

b11365

ggml-cpu: dstが src1 (#27096) を代名するときに soft_max_backの出力が誤っている。 ggml-cpu: dstが src1を代名するときに soft_max_backの出力が誤っている。 GGML_OP_SOFT_MAX_BACKは ggml_op_can_inplaceにリストされている。 したがって、グラフアロケータは、dstを src0 (dy) または src1 (y) の別名に代名することができます。結果はいくつかのステップで構築されました: ggml_vec_cpy_f32 (nc, dx, dy); ggml_vec_acc1_f32 (nc, dx, -dot_y_dy); ggml_vec_mul_f32 (nc,...