Awesome Infra for AI › Model Serving Frameworks

predibase/lorax

⭐ 3835 Python repository created 2023-10-20

LoRAX (LoRA eXchange) is a specialized inference server that dramatically reduces the cost and improves the efficiency of serving numerous fine-tuned Large Language Models. Its core innovation lies in its ability to manage and serve thousands of LoRA (Low-Rank Adaptation) adapters on a single GPU without compromising on throughput or latency. Key features include dynamic adapter loading, which allows adapters from HuggingFace, Predibase, or local filesystems to be loaded just-in-time without blocking concurrent requests, and the ability to merge adapters per request. The server employs heterogeneous continuous batching to pack requests for different adapters into the same batch, thereby maintaining consistent latency and throughput. Adapter exchange scheduling asynchronously prefetches and offloads adapters between GPU and CPU memory, optimizing aggregate system throughput. LoRAX also incorporates various inference optimizations such as tensor parallelism, pre-compiled CUDA kernels (flash-attention, paged attention, SGMV), quantization (bitsandbytes, GPT-Q, AWQ), and token streaming. It is production-ready with prebuilt Docker images, Helm charts for Kubernetes deployment, Prometheus metrics, distributed tracing with Open Telemetry, and an OpenAI-compatible API that supports multi-turn chat conversations and private adapters with per-request tenant isolation. The platform supports popular base models like Llama, Mistral, and Qwen, loading them in fp16 or quantized formats, and integrates with PEFT and Ludwig for LoRA adapter training. LoRAX is free for commercial use under the Apache 2.0 license.

https://github.com/predibase/lorax

LLM servingLoRAinference servermodel servingLLM inferenceGPU optimizationcontinuous batchingadapter loadingKubernetesAPIperformancedeploymentPredibase

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...