speculative decoding

imagem 23

Ollama v0.32.6-rc0 Brings Faster Qwen3.5 and Better OpenAI Streaming

Performance Boost for Qwen3.5 on Apple GPUs Ollama’s latest release candidate speeds up Qwen3.5 on Apple hardware by automatically enabling speculative decoding through the MLX engine, which now leverages the model’s MTP head. The MLX and llama.cpp engines have also been updated to their latest versions, ensuring broader compatibility and performance improvements. Smoothed API Compatibility […]

Ollama v0.32.6-rc0 Brings Faster Qwen3.5 and Better OpenAI Streaming Read More »

imagem 34

Ollama 0.32.4 Brings Laguna to Apple GPUs, Speeds Up Qwen3 MoE, and Refines Speculative Decoding

The latest point release of Ollama, version 0.32.4, expands hardware support by enabling the Laguna model series on Apple GPUs through the MLX engine. This means self-hosted users with Apple Silicon can now run these models with full hardware acceleration, tapping into the performance and efficiency of the Metal-backed MLX runtime that has already proven

Ollama 0.32.4 Brings Laguna to Apple GPUs, Speeds Up Qwen3 MoE, and Refines Speculative Decoding Read More »