Awesome Infra for AI › Model Serving Frameworks

ModelTC/LightLLM

⭐ 4305 Python repository created 2023-07-22

LightLLM is a Python-based framework engineered for efficient Large Language Model (LLM) inference and serving. It distinguishes itself through its lightweight design, ease of scalability, and high-speed performance, vital for deploying LLMs in production environments. The framework integrates and leverages advanced techniques from various open-source projects including FasterTransformer, Text Generation Inference (TGI), vLLM, and FlashAttention to achieve its performance targets. Its core purpose is to optimize the runtime execution of LLMs, enabling faster response times and higher throughput. LightLLM supports deployment of a wide array of LLMs and provides features like prefix KV cache transfer and sophisticated request schedulers, which are critical for maintaining service level agreements (SLAs) in LLM serving. The project emphasizes a pure-Python design and token-level KV cache management, making it a flexible and adaptable base for both production deployments and academic research in LLM serving optimization. It has been referenced and used by other significant projects in the LLM ecosystem, highlighting its impact and technical contributions to the inference and serving domain.

https://github.com/ModelTC/LightLLM

deep-learninggptllamallmmodel-servingnlpopenai-tritoninferenceservingoptimizationperformance

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...