llama.cpp Build b10203 Adds SYCL q2_0 Mul_Mat Support and Wide Platform Binaries

imagem 55

New Feature: SYCL q2_0 Mul_Mat

The latest build of llama.cpp, version b10203, introduces support for the q2_0 quantization scheme in the matrix multiplication (mul_mat) operation when using the SYCL backend. This enables efficient inference with 2-bit quantized models on Intel GPUs and other SYCL-compatible hardware, reducing memory footprint and potentially boosting throughput. Both the basic q2_0 mul_mat case and additional q2_0 operations are now covered, broadening the applicability of extremely low-bit models for self-hosted setups.

Downloads

Pre-built binaries for this release are available for all major operating systems. An online demo is accessible at llama.app, and a standalone UI package can be downloaded separately. Below you’ll find direct links for each platform and accelerator combination.

macOS/iOS

Linux

Android

Windows

openEuler

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI

Leave a Comment

Your email address will not be published. Required fields are marked *