vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Pipeless is an open-source framework designed to simplify the development and deployment of real-time computer vision (CV) applications. It handles various complexities inherent in CV systems, including code parallelization, multimedia pipeline management (supporting Gstreamer), memory management, and efficient model inference. The framework is particularly adept at managing multiple video streams simultaneously, allowing dynamic configuration and processing steps for each stream without requiring code changes or restarts. Inspired by serverless architectures, Pipeless enables developers to define "stages" as micro-pipelines, each comprising a pre-process function, a model, and a post-process function. This modular approach allows for flexible composition of processing workflows. It supports industry-standard models like YOLO and custom models, integrating with popular inference runtimes such as ONNX Runtime, CUDA, TensorRT, OpenVINO, and CoreML, enabling high-performance inference on both CPUs and GPUs. The framework supports multi-language hooks (including Python), built-in restart policies for stream resilience, and highly parallelized execution, abstracting away threading and multiprocessing complexities. Deployment options include edge, IoT devices, and cloud environments, facilitated by container images. Pipeless aims to significantly reduce the time needed to ship production-ready computer vision applications from weeks or months to minutes.
https://github.com/pipeless-ai/pipeless
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...