O

Ollama

O
Ollama AI v0.40.0

v0.40.0

What's Changed Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX. ollama pull qwen3.8 ollama run qwen3.8 During the pre-release we will be testing and enabling additional models. Full Changelog: v0.34.4...v0.40.0-rc0

O
Ollama AI v0.40.0-rc1

v0.40.0-rc1: mlx: match publisher tokenizer semantics (#18779)

mlx: match publisher tokenizer semantics Honor pretokenizer stage order, split behavior, Unicode boundaries, added-token normalization, and ranked BPE merges. Handle empty added tokens and empty Metaspace input consistently. Add shared Go/Python reference cases using published tokenizers, pulling missing models directly and failing on errors, plus focused regressions for configuration precedence,…

O
Ollama AI v0.35.1

v0.35.1

Clef decision models Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone. Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it. curl http://localhost:11434/v1/systemone -d '{ "model": "clef-flash", "state": "The user took this screenshot.",…

O
Ollama AI v0.35.0

v0.35.0

Decision models Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API. Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification. Available models: Nimble from Bespoke Labs Tev1 from Together AI ollama pull nimble Send context and one or more questions: curl http://…

O
Ollama AI v0.35.1-rc0

v0.35.1-rc0: create: support explicit model capabilities (#18708)

Add CAPABILITY declarations to Modelfiles and an additive capabilities field to create requests. Preserve declarations across GGUF and safetensors creation, inheritance, and Modelfile export. Require decision capability before scheduling System One requests instead of matching Qwen architecture/renderer metadata. Retain main's GGUF-only scoring restriction until the separate MLX runtime work lands…

O
Ollama AI v0.35.0-rc1

v0.35.0-rc1

mlx: bound pull stall retries and let the watchdog interrupt them (#1…

O
Ollama AI v0.35.0-rc0

v0.35.0-rc0

feat: add System One scoring API (#18606)

O
Ollama AI v0.40.0-rc0

v0.40.0-rc0: llama-server: prepare to remove compatibility patch

Add manifest-list storage so runner-specific manifests can coexist under one tag while preserving existing v1 tags as best-effort downgrade anchors. Show/list/copy/remove/pull/push now understand runner and digest selection and transfer referenced child manifests and layers. Add lazy local compatibility migration for legacy Ollama GGUFs into llama.cpp-compatible children, covering the patched mode…

O
Ollama AI v0.34.4

v0.34.4

What's Changed Structured outputs on thinking models now apply in a single pass, making them faster and more reliable. Fixed intermittent "model not found" errors with a large local library Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running. Qwen 3.8 prompt processing is faster on Apple Silicon. Gemma 4 on Apple Silicon now picks the best image resolution per im…

O
Ollama AI v0.34.3

v0.34.3

What's Changed GET /api/show now advertises each model's thinking controls and default: Available in the CLI with: ollama show gemma4 thinking levels false, true default true Available in the API with: curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}' { "thinking": { "values": ["low", "high", "max"], "default": "max" } } Also available on ollama.com directly for cloud mode…

O
Ollama AI v0.34.3-rc1

v0.34.3-rc1

server: allow registry cross-host redirects among allowlisted hosts (…

O
Ollama AI v0.34.3-rc0

v0.34.3-rc0

api: expose model thinking levels and defaults (#18473)

O
Ollama AI v0.34.2

v0.34.2

What's Changed Added first-run setup when running ollama, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows. Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows. Fixed excessive memory growth during long generations with MLX speculative decoding. Updated llama.cpp. Full Changelog: v0.34.1...v0.34.2

O
Ollama AI v0.34.2-rc3

v0.34.2-rc3

cli: add first-run onboarding shared with the desktop app (#18495)

O
Ollama AI v0.34.2-rc2

v0.34.2-rc2: mlxrunner: Release freed KV buffers during speculative decode

The decode loop releases MLX's pool of freed buffers every 256 generated tokens, which is also how often the KV cache grows and drops its previous, smaller buffers. The check fires only when the token count lands exactly on a multiple of 256. Speculative decoding emits several tokens per round, so most rounds step over the boundary and the pool is never released. Each growth at a long context leav…

O
Ollama AI v0.34.2-rc1

v0.34.2-rc1: mlxrunner: lay out model by contract, checkpoint and construction

model is one package with three jobs: the contract between the runner and the architectures, the opened checkpoint, and building nn layers from checkpoint tensors. Its files did not say which was which. base.go carried the folded package's name over the interfaces and the registry, root.go held the safetensors header scan next to Root, and quant.go mixed the nvfp4 global-scale helpers with quant p…

O
Ollama AI v0.34.1

v0.34.1

What's Changed MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. Improved MLX memory handling on Apple Silicon Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR) /api/tags is much faster on large model libraries (3.1 s → 294 ms cold in testing),…

O
Ollama AI v0.34.1-rc1

v0.34.1-rc1

mlx: add mlx patch to docker build context (#18440)

O
Ollama AI v0.34.0

v0.34.0

Use Ollama models in ChatGPT Desktop Ollama models can now be used directly in ChatGPT Desktop, so you can keep your existing workflow while running open models. Setup is available from the Ollama app on MacOS. This release also improves structured output performance on Apple Silicon, adds support for OpenAI-compatible client tool search and response compaction. Full Changelog: v0.33.3...v0.34.0

O
Ollama AI v0.34.0-rc5

v0.34.0-rc5

openai: support standalone named function outputs (#18348)

O
Ollama AI v0.34.0-rc4

v0.34.0-rc4

proxy: normalize namespaced commands in Full Access (#18331)

O
Ollama AI v0.34.0-rc3

v0.34.0-rc3

openai: accept plaintext-labeled Codex agent messages (#18329)

O
Ollama AI v0.34.0-rc2

v0.34.0-rc2

openai: finalize responses at the web search limit (#18328)