Awesome Infra for AI › Inference Optimization

interestingLSY/swiftLLM

⭐ 358 Python repository created 2024-05-11

SwiftLLM is an open-source, tiny yet powerful LLM inference system specifically tailored for research purposes. It aims to provide performance comparable to vLLM while drastically reducing the codebase to less than 2,000 lines of Python and OpenAI Triton code. This design philosophy makes it exceptionally easy to read, modify, debug, test, and extend, serving as a flexible foundation for researchers exploring novel ideas in LLM serving. The project supports key optimization features such as Iterational Scheduling and Selective Batching, PagedAttention, Piggybacking prefill and decoding, Flash Attention, and Paged Attention v2 (Flash-Decoding). It primarily focuses on LLaMA, LLaMA2, and LLaMA3 models. SwiftLLM differentiates itself from production-oriented frameworks by intentionally omitting features like quantization, LoRA, multimodal support, and a wide range of model/hardware compatibility to maintain its minimal footprint. Architecturally, SwiftLLM comprises a "control plane" responsible for scheduling and coordination (e.g., Engine, Scheduler, API server) and a "data plane" handling concrete computations (model layers, Triton kernels). This modular design allows researchers to either leverage both components or integrate their custom control plane with SwiftLLM's optimized data plane. While currently single-node, future plans include tensor and pipeline parallelism. It is explicitly not an all-in-one production solution but a research tool to simplify experimentation and development in LLM inference.

https://github.com/interestingLSY/swiftLLM

LLM inferencemodel servinginference engineresearch toolperformance optimizationLLaMATritonPagedAttentionFlashAttentionLLMOps

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.