vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
SGLang-Omni is a serving runtime for multimodal, speech and text-to-speech models, built by the SGLang team as a companion to their text-focused LLM engine. Where a plain LLM server handles one model and one token stream, omni and speech workloads chain several stages together: an audio or text encoder, a language backbone, an acoustic or semantic token generator, and a vocoder that turns tokens into waveforms. SGLang-Omni is designed around that multi-stage shape, giving each stage its own scheduling so that the pipeline can stream audio out while later requests are still being prefilled. The server speaks OpenAI-compatible routes, including /v1/audio/speech for synthesis and chat endpoints for omni models that accept mixed text, image and audio input, so existing OpenAI SDK clients can point at it without code changes. It carries shared pipeline state, engine construction, reference-voice encoding, capability metadata and vocoder scheduling as reusable components, and ships day-zero support for new speech and music models as they are released — recent examples include MiniMax Music 3, MOSS-TTS Local Transformer and Higgs Audio v3, several of which emit native streaming speech at 32 or 48 kHz. Installation is a single Python package from PyPI, and the documentation ships a cookbook of per-model recipes covering the launch flags and request shapes each checkpoint expects. The project is aimed at teams that already run SGLang for text inference and now need production serving for speech and multimodal endpoints on the same operational footing: continuous batching, streaming, capability discovery and an HTTP surface that mirrors the rest of their stack.
https://github.com/sgl-project/sglang-omni
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...