Ollama 0.32.4 Brings Laguna to Apple GPUs, Speeds Up Qwen3 MoE, and Refines Speculative Decoding
The latest point release of Ollama, version 0.32.4, expands hardware support by enabling the Laguna model series on Apple GPUs through the MLX engine. This means self-hosted users with Apple Silicon can now run these models with full hardware acceleration, tapping into the performance and efficiency of the Metal-backed MLX runtime that has already proven […]

