vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Roboflow Inference allows users to deploy and manage computer vision models, including custom fine-tuned models and foundation models like Florence-2, CLIP, and SAM2, on any computer or intelligent edge device. It acts as a command center for computer vision projects, offering functionalities like real-time inference, video stream management, and integration with traditional computer vision methods (e.g., OCR, barcode reading). A core feature is 'Workflows,' which enables users to compose various ML models and traditional CV tasks into complex, chained pipelines. These workflows can be used to track, count, time, measure, and visualize objects, add business logic, and send notifications. The platform supports deploying production systems at scale and provides an API for programmatic interaction. Users can install it via `pip` and Docker, with optional GPU acceleration. It facilitates scenarios such as active learning, multi-model consensus, and building custom AI agents for video streams, essentially turning any camera into an AI camera for monitoring, recording, and analyzing predictions.
https://github.com/roboflow/inference
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...