vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
PegaInfer is an LLM inference server built from scratch in Rust and CUDA. It carries no Python framework runtime at all: every GPU kernel and every scheduler is hand-written, and the default build serves Qwen3-4B and Qwen3-8B without any Python present. Model support is gated behind cargo features — qwen3 is on by default, while Qwen3.5 (hybrid Gated DeltaNet plus full attention), DeepSeek-V2-Lite (MoE with expert parallelism across two GPUs) and Kimi-K2-Instruct (MLA plus MoE with Marlin INT4 on an eight-GPU expert-parallel path) each require rebuilding the server with the matching feature flag. Model type is auto-detected from the checkpoint config.json, so the server only needs a --model-path. The Qwen3.5 and GLM5.2 feature builds invoke Triton or TileLang at build time to generate ahead-of-time kernels, but neither needs Python at runtime. The HTTP surface is an OpenAI-compatible /v1/completions endpoint supporting max_tokens, temperature, top_k, top_p and SSE streaming; sampling and logprob coverage varies by model line. The project publishes a head-to-head benchmark against vLLM 0.22.1 on a single RTX 5090 with Qwen3-4B in BF16, driven by the same vllm bench serve client: roughly 771 MB resident memory against vLLM 3814 MB, and a three-second cold start to HTTP-ready against vLLM 70 seconds cold or 32.7 seconds with a warm torch.compile cache. Requirements are a CUDA-capable GPU, the CUDA Toolkit and an NVIDIA R545 driver or newer; release builds are mandatory because debug CUDA builds are unusably slow. Windows is supported alongside Linux. It suits teams that want a small, fast-starting single-process inference server without a Python model framework in the deployment.
https://github.com/pegainfer-project/pegainfer
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...