Awesome Infra for AI › Inference Optimization

chengzeyi/ParaAttention

⭐ 434 Python repository created 2024-10-28

ParaAttention is a project designed to significantly speed up the inference process for Diffusion Transformer (DiT) models like FLUX, HunyuanVideo, and Mochi. It achieves this primary goal through two main mechanisms: * **Context Parallelism:** This technique partitions input tensors along the sequence dimension across multiple GPUs, allowing for parallel processing of neural network activations in all layers. Unlike sequence parallelism, which focuses on specific layers, context parallelism in ParaAttention offers a unified attention mechanism that combines Ulysses and Ring-style parallelism, optimizing performance across various models and hardware configurations. * **First Block Cache (FBCache):** Inspired by other denoising caching algorithms, FBCache reuses previous computation results if the difference in the residual output of the first transformer block is sufficiently small. This dynamic caching strategy can lead to substantial reductions in computation a significant speedup (up to 2x) while maintaining high accuracy, with an adjustable `residual_diff_threshold` to balance speed and fidelity. The project emphasizes ease of use, providing simple interfaces to enable these optimizations with minimal code changes, particularly for `diffusers` pipelines. It also highlights efficient `torch.compile` integration, aiming to minimize graph breaks for maximum optimization opportunities in the backend compiler. ParaAttention is primarily focused on the operational and performance aspects of deploying and serving pre-trained AI models, rather than model training or development.

https://github.com/chengzeyi/ParaAttention

inference optimizationcontext parallelismdynamic cachingdiffusion transformersattentionparallel computingdiffusersFLUXHunyuanVideoMochiGPU accelerationLLM inference

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.