Awesome Infra for AI › Model Serving Frameworks

vllm-project/vllm-omni

⭐ 7046 Python repository created 2025-09-11

vLLM-Omni extends the capabilities of the original vLLM framework to support a broader range of AI models beyond just large language models. While vLLM initially focused on text-based autoregressive generation, vLLM-Omni expands this to include omni-modality inference and serving, handling text, image, video, and audio data. It specifically addresses non-autoregressive architectures like Diffusion Transformers (DiT) and other parallel generation models, enabling heterogeneous outputs from traditional text generation to various multimodal formats. The framework is engineered for speed, leveraging vLLM's efficient KV cache management for autoregressive models and implementing pipelined stage execution overlapping to achieve high throughput. It features full disaggregation based on 'OmniConnector' and dynamic resource allocation across different processing stages. Flexibility and ease of use are key, with a heterogeneous pipeline abstraction to manage complex model workflows, seamless integration with Hugging Face models, and support for various parallelism techniques (tensor, pipeline, data, expert) for distributed inference. It also offers streaming outputs and an OpenAI-compatible API server. vLLM-Omni supports popular open-source omni-modality, TTS, and diffusion models from HuggingFace, making it a comprehensive solution for deploying and serving advanced multimodal AI models.

https://github.com/vllm-project/vllm-omni

omnimodalmodel servinginferencemultimodaldiffusionaudio generationvideo generationimage generationtransformerpytorchdistributed inferencequantizationAI deployment

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...

superduper-io/superduper

SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.