Awesome Infra for AI › Inference Optimization

alibaba/rtp-llm

⭐ 1356 Cuda repository created 2023-12-27

RTP-LLM is a high-performance inference acceleration engine developed by Alibaba's Foundation Model Inference Team, widely used within the Alibaba Group to support LLM services across various business units. It focuses on optimizing the serving phase of Large Language Models (LLMs), not the training phase. Key features include the utilization of advanced CUDA kernels like PagedAttention and FlashAttention for high throughput and low latency, as well as various quantization techniques such as WeightOnly INT8/INT4 (with GPTQ and AWQ support) and Adaptive KVCache Quantization to reduce memory footprint and improve performance. The engine incorporates detailed optimization of dynamic batching overhead at the framework level and is specifically optimized for V100 GPUs, with ongoing development for multi-hardware support including AMD ROCm, Intel CPU, and ARM CPU. RTP-LLM offers flexibility by integrating seamlessly with HuggingFace models, supporting multiple weight formats, deploying multiple LoRA services from a single model instance, handling multimodal inputs, and enabling multi-machine/multi-GPU tensor parallelism. It also includes advanced acceleration techniques like Contextual Prefix Cache, System Prompt Cache, and Speculative Decoding for efficient multi-turn dialogues and reduced inference costs. The project is built upon and inspired by other leading inference projects such as FasterTransformer, TensorRT-LLM, vLLM, and HuggingFace Transformers, making it a robust solution for production-grade LLM serving.

https://github.com/alibaba/rtp-llm

LLM inferencemodel servingperformance optimizationquantizationCUDAdistributed inferenceAlibabaLLMOps

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.