vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
Shimmy is a lightweight, single-binary inference server written entirely in Rust, designed to serve GGUF (GGML Universal File Format) large language models (LLMs) with 100% OpenAI-compatible API endpoints. This allows existing AI tools and applications to seamlessly integrate with locally run LLMs without modification. A key feature is its custom **Airframe engine**, a pure-Rust WebGPU (WGSL) transformer runtime that eliminates the need for C++ toolchains or complex compilation, ensuring high-quality, deterministic output on any GPU (NVIDIA, AMD, Intel, integrated) through WebGPU. Airframe automatically determines model specifications from GGUF metadata and supports extended context via YaRN RoPE scaling. Further enhancing efficiency, Shimmy includes **TurboShimmy INT4 KV**, an on-GPU KV-cache compression system. This innovation reduces KV cache VRAM consumption by approximately 7x without significant quality degradation, making it possible to run larger models or longer contexts on GPUs with limited memory (e.g., Llama-3.2-3B on 4GB GPUs or 7B models on 6GB GPUs). TurboShimmy operates entirely within WGSL compute shaders, converting 32-bit floats to 4-bit integers for the KV cache and dequantizing on-the-fly, ensuring no CPU roundtrips. Shimmy supports a wide range of GGUF models across various architectures and quantizations, read directly from model metadata for flexible deployment. The project emphasizes being 'free forever,' relying on community sponsorship for continuous development.
https://github.com/Michael-A-Kuykendall/shimmy
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.