Awesome Infra for AI › Model Serving Frameworks

BMW-InnovationLab/BMW-YOLOv4-Inference-API-CPU

⭐ 217 Python repository created 2019-12-11

This project provides a robust, CPU-only inference API for YOLOv3 and YOLOv4 object detection models, enabling efficient deployment and serving of trained models. It's built to run on both Windows and Linux, leveraging Docker and Docker Swarm for streamlined containerization and scaling. The API offers various RESTful endpoints for loading models, performing object detection on single images or batches, retrieving labels, and querying model configurations. A key feature is its 'no-code' approach, simplifying the deployment process for users without extensive programming knowledge. It supports loading multiple object detection models concurrently and provides mechanisms for managing model configurations, including confidence thresholds and NMS thresholds. The inclusion of Docker Swarm support allows for redundancy, load balancing, and scaling of the inference service, which is crucial for handling higher traffic and ensuring service availability. The project outlines a clear model structure requirement, specifying the need for configuration files, weights, class names, and a JSON configuration for each model. This tool is purpose-built for the operational phase of AI, focusing exclusively on inference and serving of pre-trained models rather than the training aspect.

https://github.com/BMW-InnovationLab/BMW-YOLOv4-Inference-API-CPU

object detectioninference APIYOLOCPU inferenceDockerDocker SwarmREST APIcomputer visiondeep learningmodel serving

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...