Awesome Infra for AI › Model Serving Frameworks

efeslab/Nanoflow

⭐ 975 Jupyter Notebook repository created 2024-08-19

Nanoflow is an advanced serving framework specifically engineered for Large Language Models (LLMs), focusing on maximizing throughput and efficiency. It introduces novel techniques such as intra-device parallelism, which allows for the simultaneous execution of compute-, memory-, and network-bound operations within a single GPU by using nano-batching and execution unit scheduling. This approach significantly enhances resource utilization compared to traditional sequential execution pipelines. Additionally, Nanoflow incorporates asynchronous CPU scheduling to efficiently manage GPU execution, batch formation, and KV-cache management, thereby reducing CPU overhead during inference. The framework also optimizes KV-cache handling by eagerly offloading finished requests to SSDs, enabling faster reuse for multi-round conversations. Benchmarking against state-of-the-art frameworks like vLLM, Deepspeed-FastGen, and TensorRT-LLM, Nanoflow demonstrates superior throughput performance and effectively sustains higher request rates with lower latency across various real-world and synthetic workloads. The project provides a Cpp-based backend and a Python-based demo frontend, integrating cutting-edge kernel libraries such as CUTLASS, FlashInfer, and MSCCL++ for optimized operations. It is designed to be flexible and supports a range of popular LLM architectures, including Llama2, Llama3, and Qwen2 models.

https://github.com/efeslab/Nanoflow

LLM servinginference optimizationhigh-performance computingdeep learningGPU utilizationparallel computingasynchronous schedulingKV-cache managementCudaLLMs

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...