Ollama v0.32.6-rc0 Brings Faster Qwen3.5 and Better OpenAI Streaming
Performance Boost for Qwen3.5 on Apple GPUs Ollama’s latest release candidate speeds up Qwen3.5 on Apple hardware by automatically enabling speculative decoding through the MLX engine, which now leverages the model’s MTP head. The MLX and llama.cpp engines have also been updated to their latest versions, ensuring broader compatibility and performance improvements. Smoothed API Compatibility […]
Ollama v0.32.6-rc0 Brings Faster Qwen3.5 and Better OpenAI Streaming Read More »


