vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
vllm-mlx is an inference server that brings vLLM-style serving mechanics to Apple Silicon Macs through the MLX framework. Compared with using Ollama or mlx-lm directly, it adds the throughput machinery that a serving stack needs: continuous batching so concurrent requests share the GPU efficiently, a paged KV cache with prefix sharing, a trie-based prefix cache shared across requests, an SSD-tiered cache that spills prefixes to disk for long-context agents, and warm prompts that preload popular prefixes at startup to cut time to first token. It exposes two API families from one process — the OpenAI routes for chat completions, completions, embeddings, rerank and responses, and Anthropic's /v1/messages with streaming, tool use and system prompts — so an OpenAI SDK client and Claude Code can both point at the same local server with only environment variables changed. Models run on Metal with unified memory and no conversion step. Coverage is multimodal: text, image, video and audio input from one server, vision models including Gemma 3 and 4, Qwen3-VL, Pixtral and Llama vision, audio content blocks in chat, native text-to-speech with a dozen voices across many languages, and Whisper-family speech recognition. Tool calling is parsed for a dozen model families, structured output is enforced against a JSON schema, and reasoning traces can be extracted for models such as Qwen3 and DeepSeek-R1. Advanced options include mixture-of-experts top-k reduction, speculative decoding and sparse prefill. Operationally it exports Prometheus metrics and ships a built-in benchmarking command for prompt sweeps with CSV or JSON output. Installed from PyPI, Apache 2.0 licensed, and restricted to Apple Silicon.
https://github.com/waybarrios/vllm-mlx
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...