vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
xLLM is an efficient inference framework specifically designed for Large Language Models (LLMs), Vision-Language Models (VLMs), Diffusion Transformers (DiT), and REC models. Its core purpose is to provide high-performance, low-latency, and high-throughput inference capabilities, catering to enterprise-grade deployments. A significant highlight of xLLM is its deep optimization for various AI accelerators, with a particular focus on Chinese domestic hardware such as Ascend NPU, Cambricon MLU, Moore Threads GPU, Hygon DCU, MetaX MACA, and Iluvatar CoreX GPU. This specialization enables enhanced efficiency and reduced operational costs for AI inference workloads. The framework features a service-engine decoupled architecture, where the service layer manages scheduling and availability, while the engine layer handles the computational aspects. This design promotes scalability and robustness. xLLM has been battle-tested at scale within JD.com's core retail business, demonstrating its capability for real-world production environments. It supports a wide array of mainstream LLM models, and its development includes advanced features like hybrid KV cache management (building upon projects like Mooncake) to further optimize memory usage and performance during inference. The project is actively maintained and has released technical reports detailing its architecture and implementation insights.
https://github.com/xLLM-AI/xllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...