Awesome Infra for AI › Model Serving Frameworks

dphnAI/sonar

⭐ 1869 C++ added to this list on 2026-07-06 repository created 2023-06-23

Sonar, formerly Aphrodite Engine, is an advanced inference engine designed for efficient and scalable serving of large language models. It builds upon and integrates technologies from projects like vLLM, particularly leveraging its PagedAttention mechanism for optimized K/V cache management. The engine supports continuous batching, significantly improving throughput for concurrent users and ensuring high-performance model inference. Key features include highly optimized CUDA kernels, a wide range of quantization techniques (e.g., AQLM, AWQ, GPTQ, ExLlamaV3, Marlin, MXFP4, BitsAndBytes) to reduce memory footprint and speed up inference, and support for distributed inference deployments. Sonar also incorporates modern sampling techniques like DRY, XTC, and Mirostat, disaggregated inference, and various speculative decoding methods such as EAGLE and DFlash. It provides multimodal and multi-LoRA support, allowing for greater flexibility in model deployment. The engine offers an OpenAI-compatible API server for easy integration with existing UIs and platforms, making it a robust solution for large-scale LLM operationalization.

https://github.com/dphnAI/sonar

large language modelsLLM inferencemodel servinginference enginevLLMPagedAttentionquantizationspeculative decodingdistributed inferenceCUDAHuggingFaceAPI servercontinuous batching

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...