vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Jina-serve is a comprehensive framework designed for building and deploying multimodal AI services, from local development to production at scale. It facilitates the creation of AI services that communicate via gRPC, HTTP, and WebSockets, emphasizing high performance with features like scaling, streaming, and dynamic batching for efficient model inference. The framework offers native support for all major ML frameworks and data types, making it versatile for various AI applications. Key capabilities include robust LLM serving with token-by-token streaming, which is crucial for responsive generative AI applications. Jina-serve also provides built-in Docker integration and an Executor Hub for easy service management. For deployment, it offers one-click options to Jina AI Cloud and enterprise-ready support for Kubernetes and Docker Compose, simplifying the operational aspects of AI services. It allows users to define custom AI services using familiar Python classes (Executors) and chain them into pipelines (Flows) for complex AI workflows. The framework's core concepts revolve around BaseDoc/DocList for data handling, Executors for processing, Gateways for connectivity, and Deployments/Flows for orchestration.
https://github.com/jina-ai/serve
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...
SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.