Awesome Infra for AI › Inference Optimization

AI-Hypercomputer/JetStream

⭐ 464 Python repository created 2024-03-01

JetStream is a high-performance engine specifically designed for serving large language models (LLMs) on XLA-compatible hardware, with an initial focus on Google's Tensor Processing Units (TPUs) and planned support for GPUs. Its core purpose is to optimize LLM inference for maximum throughput and memory efficiency, critical aspects for deploying and scaling generative AI applications. The project provides reference engine implementations for both JAX and PyTorch models, allowing developers to leverage its optimizations for various LLM architectures. JetStream aims to simplify the process of deploying LLMs on specialized hardware by offering tools for setup, local testing, and performance benchmarking. It also includes features for observability and profiling, enabling users to monitor and fine-tune their inference deployments. While the project is in the process of migrating core functionality to `tpu-inference`, its principles and existing capabilities remain relevant for understanding and implementing high-efficiency LLM serving solutions on XLA devices.

https://github.com/AI-Hypercomputer/JetStream

gemmagptgpuinferencejaxlarge-language-modelsllamallama2llmllm-inferencellmopsmlopsmodel-servingpytorchtputransformerXLAthroughput optimizationmemory optimization

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.