jundot/omlx
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
Awesome Infra for AI › Model Serving Frameworks
vLLM is an open-source library designed for efficient inference and serving of large language models (LLMs). It was initially developed in the Sky Computing Lab at UC Berkeley and has since grown into a project supported by a large community. The engine focuses on maximizing throughput and optimizing memory usage, primarily through its innovative PagedAttention mechanism. This technology efficiently manages the key and value memory for attention mechanisms, leading to significant performance gains. Key features of vLLM include continuous batching for incoming requests, chunked prefill, and prefix caching to accelerate processing. It offers fast and flexible model execution leveraging piecewise and full CUDA/HIP graphs. For quantization, vLLM supports a wide array of formats including FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, and more, as well as optimized attention kernels like FlashAttention and FlashInfer. It also incorporates optimized GEMM/MoE kernels and various speculative decoding techniques, alongside automatic kernel generation and graph-level transformations. vLLM emphasizes ease of use and flexibility, offering seamless integration with Hugging Face models, high-throughput serving with diverse decoding algorithms (including parallel sampling and beam search), and distributed inference capabilities (tensor, pipeline, data, and expert parallelism). It supports streaming outputs, structured output generation via xgrammar or guidance, and tools for calling and reasoning parsers. The engine provides an OpenAI-compatible API server, Anthropic Messages API, gRPC support, and efficient multi-LoRA support. It is compatible with a wide range of hardware, including NVIDIA, AMD GPUs, x86/ARM/PowerPC CPUs, and various other hardware plugins like Google TPUs and Intel Gaudi. vLLM supports over 200 model architectures, including decoder-only LLMs, Mixture-of-Expert LLMs, hybrid attention models, multi-modal models, and embedding/retrieval models.
https://github.com/vllm-project/vllm
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...
SuperDuperDB is an end-to-end framework for building AI applications and agents by integrating AI models directly into databases, facilitating inference, RAG, and AI agent orchestration.