trunk/09232df64e161efc316a66bdbc3755991a1f594a
[Testcase Refactoring] Add hw_classification in test/test_fx.py (#192…
[Testcase Refactoring] Add hw_classification in test/test_fx.py (#192…
CUDA: refactor swizzling code fix templates/loop bounds
Fixes #196631 Summary TorchInductor already optimizes index_add when the index is a sliced permutation: index = torch.randperm(x.shape[0], device=x.device)[: y.shape[0]] result = torch.index_add(x, dim=0, source=y, index=index) Because the indices are unique and in bounds, the accumulating update can be rewritten using unsafe indexing and a non-accumulating index_put. The full-permutation case: in…
On MPS, torch.linalg.lstsq raised IndexError: max(): Expected reduction dim 0 to have non-zero size whenever min(m, n) == 0. Root cause: the MPS kernel solves the system through an SVD and computes the rank cutoff from the largest singular value, S.max(-1). An empty matrix has no singular values, and max over an empty dimension has no identity value, so it raises (the same as on CPU). Fix: when mi…
ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (#29806) ggml-cpu: vectorize BF16 K tails in tinyBLAS tests: Skip tinyBLAS when use_ref is enabled so CPU tests compare against the vec_dot path. ggml-cpu: vectorize tinyBLAS F16/F32 tails Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52638584 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple…
Python 3.10 reaches EOL this month and was dropped from the nightly binary build matrix in #198233, but 13 of the 31 images in .ci/docker/build.sh were still pinned to it. Most did not say so: only two carry the version in their tag, so images like pytorch-linux-jammy-cuda13.2-cudnn9-py3-gcc11 -- 28 workflow references -- were quietly 3.10. Bumps the eleven whose tags do not encode a version, so n…
cuda : move neu_padded to where it is used (#29940) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52636582 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu…
Summary: test_ordered_set.py imports CPython's test.support for two helpers, check_free_after_iterating and gc_collect. Not every Python distribution ships the test package, and where it is missing the whole module fails to import, so no test in it runs: ImportError: cannot import name 'support' from 'test' (unknown location) Copy check_free_after_iterating into the test file and call gc.collect()…
Safari rendering broke at some point with the trace plot overlapping. Fixed with some new divs. AI slop says: Safari resolves height: 100% on the timeline SVG grid items against the entire grid container instead of their assigned rows. This makes the main timeline and minimap overflow their tracks and overlap the stack trace panel. The stretched SVG also distorts the y-axis labels. Place each time…
ci : windows llvm build requires ninja multi-config (#29959) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52619546 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu…
PR Build Information Version: 0.0.0-nightly.pr20359.32835 Release Time: 2026-10-04T17:41:22.895Z PR: #20359 ⚠️ Important Notice This is a development build specifically created for testing purposes. Please note: This build is NOT intended for production use Features may be incomplete or unstable Use only for validating PR changes in a desktop environment May contain experimental code that hasn't b…
Summary Fix the remaining base types.DynamicClassAttribute failure from the CPython 3.13 port. DynamicClassAttribute is a pure-Python descriptor. When one is created inside a compiled class body and later read raw through C.__dict__["attr"], Dynamo sends that ephemeral descriptor through SourcelessBuilder. The builder had no case for the exact DynamicClassAttribute type, so it graph-broke with "Un…
[Inductor] Skip ExternKernelSchedulerNode in create_foreach_nodes (#1…
[Inductor] Skip ExternKernelSchedulerNode in create_foreach_nodes (#1…
add windows vulkan arm64 release add link
What's Changed Auto-generated; a maintainer may hand-edit. 🎮 Graphics When the graphics device is lost (a driver crash, driver update, or the GPU being removed), Stride no longer tries to silently recover, which used to leave things in a broken state. It now fails cleanly with a clear error, the same way Unreal and Godot handle it, and Game Studio shows the problem, offers to save, and restarts yo…
What's Changed Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX. ollama pull qwen3.8 ollama run qwen3.8 During the pre-release we will be testing and enabling additional models. Full Changelog: v0.34.4...v0.40.0-rc0
chat-peg-parser : clear current_tool when pending_tool_call is reset (#29942) A TOOL_ID node that arrives after TOOL_CLOSE wrote through current_tool, which still pointed into the just-destroyed pending_tool_call optional (use-after-free, then a second free of the id buffer). Clear the pointer on reset. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/526…
PR Build Information Version: 0.0.0-nightly.pr20358.32820 Release Time: 2026-10-04T16:13:46.946Z PR: #20358 ⚠️ Important Notice This is a development build specifically created for testing purposes. Please note: This build is NOT intended for production use Features may be incomplete or unstable Use only for validating PR changes in a desktop environment May contain experimental code that hasn't b…
For Stride engine 4.4.0-beta9. What's changed in the Stride samples Add missing animation assets for Guard Punch and Guard Idle (123c18a) Full samples changelog: samples/4.4.1...samples/4.4.2
What's Changed feat: added Sofya as a search provider by @yusufgurdogan in #965 feat: added AIHubMix as a provider fix: fixed bugs New Contributors @yusufgurdogan made their first contribution in #965 Full Changelog: v1.5.85...v1.5.86
Release 0.162.0-alpha.13
Caution ⚠️ Please read this before you install If you point v7 at your existing InvokeAI root directory, it will upgrade your database, and that database will no longer work with InvokeAI v6. The migrations only go one way, and v6 refuses to open a database that contains migrations it doesn't recognise. To try v7, create a brand-new InvokeAI root directory and experiment there. In the Invoke Launc…
mlx: match publisher tokenizer semantics Honor pretokenizer stage order, split behavior, Unicode boundaries, added-token normalization, and ranked BPE merges. Handle empty added tokens and empty Metaspace input consistently. Add shared Go/Python reference cases using published tokenizers, pulling missing models directly and failing on errors, plus focused regressions for configuration precedence,…
ci : set default permissions (#29945) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52595372 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CUDA 1…
cuda : move blocks_per_col to where it is used (#29939) Signed-off-by: Adrien Gallouët angt@huggingface.co Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52591217 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ub…
What's Changed Auto-generated; a maintainer may hand-edit. 🖥️ Launcher The launcher has been rewritten from WPF to Avalonia, a cross-platform UI framework. This release is still Windows-only (@Kryptos-FR, #3276) A visual refresh brings a proper title bar, new icons for the Getting Started and News tabs, clearer primary buttons, consistent colors, rounded cards, and a status message while the launc…
CUDA: fix MMQ memory fault if n_expert >> n_ubatch (#29941) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52581427 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu…
vulkan: fix rdna4 mat_vec tuning (#29934) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/52579133 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64 (Vulkan) Ubuntu arm64 (Vulkan) Ubuntu x64 (CUDA 12) - CU…
imatrix: calculate activation-based statistics for new format (GGUF) imatrices (#14891) Use activations to calculate the stats Determine calculation mode Compute entropy for activations Compute cosine similarity based on activations Compute l2 norm Add compute_layer_statistics() function Update aggregated statistic report layout Fix printing l2 norm when calc_mode = 1 Refactor variable name Comput…