New Feature: SYCL q2_0 Mul_Mat
The latest build of llama.cpp, version b10203, introduces support for the q2_0 quantization scheme in the matrix multiplication (mul_mat) operation when using the SYCL backend. This enables efficient inference with 2-bit quantized models on Intel GPUs and other SYCL-compatible hardware, reducing memory footprint and potentially boosting throughput. Both the basic q2_0 mul_mat case and additional q2_0 operations are now covered, broadening the applicability of extremely low-bit models for self-hosted setups.
Downloads
Pre-built binaries for this release are available for all major operating systems. An online demo is accessible at llama.app, and a standalone UI package can be downloaded separately. Below you’ll find direct links for each platform and accelerator combination.
macOS/iOS
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (ROCm 7.2)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
Android
Windows
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) – CUDA 12.4 DLLs
- Windows x64 (CUDA 13) – CUDA 13.3 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (HIP)
openEuler
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)



