llama.cpp Build b10203 Adds SYCL q2_0 Mul_Mat Support and Wide Platform Binaries
New Feature: SYCL q2_0 Mul_Mat The latest build of llama.cpp, version b10203, introduces support for the q2_0 quantization scheme in the matrix multiplication (mul_mat) operation when using the SYCL backend. This enables efficient inference with 2-bit quantized models on Intel GPUs and other SYCL-compatible hardware, reducing memory footprint and potentially boosting throughput. Both the basic […]
llama.cpp Build b10203 Adds SYCL q2_0 Mul_Mat Support and Wide Platform Binaries Read More »

