This repository contains the necessary setup instructions, docker configuration, and patch modifications to run highly optimized LLM inference on Intel XPU architectures using vLLM.
Everyone is buzzing that only NVIDIA RTX or AMD hardware can achieve elite inference speeds is outdated and Intel's software is not good enough (only if you use in 2026 software from 2025)
By leveraging the oneAPI 2026.1 toolchain, Triton XPU, and direct Intel XMX hardware acceleration, Intel GPUs deliver exceptional token-per-second throughput numbers on modern hybrid architectures.