Awesome Infra for AI › Inference Optimization

ovg-project/kvcached

⭐ 1515 Python repository created 2025-05-27

kvcached (KV cache daemon) is a specialized library designed to optimize LLM serving and training on shared Graphics Processing Units (GPUs). Its core innovation lies in applying an OS-style virtual memory abstraction to LLM systems, specifically for Key-Value (KV) caches. This approach enables elastic and demand-driven allocation of KV cache memory, which significantly improves GPU utilization, especially in environments with dynamic and mixed workloads. The project achieves this by decoupling the GPU's virtual addressing from the physical memory allocation for KV caches. Serving engines can initially reserve only virtual memory and later back it with physical GPU memory as the cache is actively used. This decoupling facilitates on-demand allocation and flexible sharing, leading to better GPU memory utilization. Key features include elastic KV cache management, GPU virtual memory for dynamic mapping, a memory control CLI for enforcing limits, a frontend router with sleep mode for idle models, and integration with popular serving engines like SGLang and vLLM. It also supports automatic prefix caching (APC) and RadixCache for cross-request prefix reuse, enhancing memory efficiency. kvcached is particularly beneficial for scenarios such as multi-LLM serving on a single GPU, serverless LLM deployments, compound AI systems with multiple specialized models, and colocation of LLM inference with other GPU workloads, like training jobs or vision models.

https://github.com/ovg-project/kvcached

elastic-kvcachegpu-multiplexinggpu-sharinginference-enginekvcachekvcache-optimizationkvcachedllmllm-inferencellm-servingonline-offline-coserveserverlesssglangvllmvirtual-memory

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.

alibaba/rtp-llm

RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.