vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Chitu (Red Rabbit) is a high-performance inference framework specifically designed for large language models (LLMs), emphasizing efficiency, flexibility, and availability in production environments. It addresses the needs of enterprise AI deployments, from small-scale experimentation to large-scale cluster deployments. The framework offers multi-platform compatibility, supporting not only a wide range of NVIDIA GPUs (from flagship to older models) but also optimized support for domestic chips like Ascend and Muxi. Chitu ensures scalability across various deployment scenarios, including pure CPU, single GPU, and large-scale clustered inference. A core focus is on long-term stable operation, enabling it to handle concurrent business traffic reliably in real-world production settings. Key features include efficient algorithm implementations for FP4 to FP8 and BF16 conversions, enabling efficient serving of large models like DeepSeek-R1 671B. The project provides pre-built Docker images for quick deployment on supported platforms and actively integrates contributions from the open-source community. It aims to provide a more efficient, flexible, compatible, and stable solution for LLM inference deployment.
https://github.com/thu-pacman/chitu
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...