Awesome Infra for AI › Model Serving Frameworks

jina-ai/serve

⭐ 21865 Python repository created 2020-02-13

Jina-serve is a comprehensive framework designed for building and deploying multimodal AI services, from local development to production at scale. It facilitates the creation of AI services that communicate via gRPC, HTTP, and WebSockets, emphasizing high performance with features like scaling, streaming, and dynamic batching for efficient model inference. The framework offers native support for all major ML frameworks and data types, making it versatile for various AI applications. Key capabilities include robust LLM serving with token-by-token streaming, which is crucial for responsive generative AI applications. Jina-serve also provides built-in Docker integration and an Executor Hub for easy service management. For deployment, it offers one-click options to Jina AI Cloud and enterprise-ready support for Kubernetes and Docker Compose, simplifying the operational aspects of AI services. It allows users to define custom AI services using familiar Python classes (Executors) and chain them into pipelines (Flows) for complex AI workflows. The framework's core concepts revolve around BaseDoc/DocList for data handling, Executors for processing, Gateways for connectivity, and Deployments/Flows for orchestration.

https://github.com/jina-ai/serve

cloud-nativecncfdeep-learningdockerfastapiframeworkgenerative-aigrpcjaegerkubernetesllmopsmachine-learningmicroservicemlopsmultimodalneural-searchopentelemetryorchestrationpipelineprometheusAI servicesmodel servingLLM deploymentinferencescalingstreamingdynamic batching

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...

superduper-io/superduper

SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.