vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Mosec is a robust and flexible framework designed for high-performance machine learning model serving. It enables developers to build ML-powered backend services and microservices by bridging the gap between trained models and efficient online inference APIs. The framework leverages Rust for its web layer and task coordination, ensuring blazing speed and efficient CPU utilization through asynchronous I/O. For ease of use, the user interface remains purely in Python, allowing ML engineers to serve models with the same code used for offline testing, agnostic to the underlying ML framework. Key features include dynamic batching, which aggregates requests for batched inference to improve system throughput, and support for pipelined stages, allowing multiple processes to handle mixed CPU/GPU/IO workloads sequentially. Mosec is cloud-friendly, designed with features like model warmup, graceful shutdown, and Prometheus monitoring metrics for seamless integration with Kubernetes or other container orchestration systems. It focuses specifically on the online serving aspect, allowing users to concentrate on model optimization and business logic without worrying about the serving infrastructure complexities. The framework supports various ML ecosystems like PyTorch, TensorFlow, JAX, and Hugging Face models, making it versatile for a wide range of AI applications.
https://github.com/mosecorg/mosec
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...