Ollama 0.32.4 Brings Laguna to Apple GPUs, Speeds Up Qwen3 MoE, and Refines Speculative Decoding

imagem 34

The latest point release of Ollama, version 0.32.4, expands hardware support by enabling the Laguna model series on Apple GPUs through the MLX engine. This means self-hosted users with Apple Silicon can now run these models with full hardware acceleration, tapping into the performance and efficiency of the Metal-backed MLX runtime that has already proven valuable for other architectures.

On the inference front, speculative decoding gets a targeted fix: draft-model output heads are now quantized at the requested precision when creating draft tokens. This ensures more consistent quality and accuracy when using speculative models to accelerate generation. Additionally, Qwen3 Mixture-of-Experts (MoE) models benefit from two important corrections—decoding now works correctly even when experts are quantized at different bit widths, and a faster packed projection for gate and up layers delivers a 4–9% speed improvement on the M5 Max (and potentially other Apple hardware). Together, these changes make both everyday inference and advanced model serving more robust on local deployments.

Leave a Comment

Your email address will not be published. Required fields are marked *