CONNECT WITH US

Intel tests system memory to ease AI KV-cache pressure

, DIGITIMES, Taipei
0

Intel Platform Validation Engineer Jeff Ke presents KV-cache offload test results at the 2026 OCP APAC Summit in Taipei. Credit: Sherri Wang

Intel tests presented at the OCP APAC Summit 2026 that moving the key-value cache used in large language model (LLM) inference from GPU memory to system DRAM can raise serving throughput and support more concurrent requests when VRAM is the bottleneck,...

The article requires paid subscription. Subscribe Now