Awesome Infra for AI › Model Serving Frameworks

openvinotoolkit/model_server

⭐ 941 C++ repository created 2018-09-26

OpenVINO Model Server (OVMS) is a scalable inference server designed for deploying and serving AI/ML models efficiently. It enables remote inference, allowing lightweight clients to perform API calls to edge or cloud deployments without direct exposure to model frameworks or hardware specifics. OVMS supports various model servers via standard network protocols like gRPC and REST API, making it easy to integrate with applications written in any programming language. It is optimized for Intel architectures and utilizes OpenVINO for inference execution, ensuring high performance. Key features include support for OpenAI-compatible generative APIs, KServe, and TensorFlow Serving, facilitating the deployment of a wide range of models including LLMs and computer vision tasks. OVMS also offers robust model management capabilities such as model versioning, dynamic input shapes, and runtime updates. It supports deployment in diverse environments including Docker containers, bare metal, and Kubernetes clusters. Additionally, it provides a Directed Acyclic Graph (DAG) scheduler for pipeline orchestration, metrics compatible with Prometheus, and support for multiple frameworks like TensorFlow, PaddlePaddle, and ONNX. The server is designed for efficient resource utilization through horizontal and vertical inference scaling, making it suitable for microservices-based applications and cloud deployments.

https://github.com/openvinotoolkit/model_server

aiclouddeep-learningedgegenaiinferencekubernetesmachine-learningmodel-servingopenvinoservingllm-servingopenai-api-compatiblemodel-deployment

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...