Awesome Infra for AI › Inference Optimization

EfficientMoE/MoE-Infinity

⭐ 367 Python repository created 2024-01-22

MoE-Infinity provides a specialized PyTorch library engineered for the efficient serving of Mixture-of-Experts (MoE) models, a type of large language model architecture. It tackles the significant memory demands of MoE models by implementing expert offloading to host memory, enabling their deployment on GPUs with limited VRAM. The library minimizes offloading overheads through advanced techniques such as expert activation tracing, activation-aware expert prefetching, and activation-aware expert caching. This allows it to achieve state-of-the-art latency performance for MoE inference in resource-constrained environments, outperforming other popular inference frameworks like vLLM, HuggingFace Accelerate, DeepSpeed, Mixtral-Offloading, and Ollama/LLama.cpp in specific benchmarks. Designed for ease of use, MoE-Infinity is fully compatible with HuggingFace models and APIs, supporting a wide range of MoE checkpoints including Deepseek-V2, Google Switch Transformers, Meta NLLB-MoE, and Mixtral. It integrates with LLM acceleration techniques like FlashAttention to further boost performance. The library also supports multi-GPU environments with OS-level optimizations. For deployment, it provides an OpenAI-compatible API server, allowing users to query models via standard HTTP requests or the `openai` Python package. While the current open-source version prioritizes user-friendliness and HuggingFace compatibility, and does not yet include distributed inference, its core focus is on optimizing inference for MoE models in production.

https://github.com/EfficientMoE/MoE-Infinity

LLM inferenceMoE modelsinference optimizationmodel servingHuggingFace compatibilityPyTorchGPU memory managementdeep learning

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.