vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Pinferencia aims to simplify the deployment of machine learning models by providing a Python-based inference server. It allows users to quickly expose their models via auto-generated GUI and REST APIs with minimal code. The library supports a wide range of machine learning frameworks including Hugging Face, PyTorch, and TensorFlow, by allowing users to register their trained models or even simple Python functions as services. A key feature is its fast setup, enabling models to go "online" with just a few lines of code. It includes an out-of-the-box GUI and automatic API documentation with an interactive try-out feature. Pinferencia also boasts 100% test coverage and is designed to be lightweight, avoiding heavy-weight solutions often found in ML model serving. Furthermore, it offers compatibility with the Kserve API, making it suitable for integration into environments like Kubeflow alongside other serving solutions like TF Serving, Triton, and TorchServe, while being particularly fast for prototyping.
https://github.com/underneathall/pinferencia
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...