Awesome Infra for AI › Inference Optimization

zengxiao-he/tessera

⭐ 565 Python added to this list on 2026-06-22 repository created 2026-06-05

Tessera is a from-scratch LLM stack designed for efficient distillation and serving of large language models. It provides an end-to-end solution for shrinking a large "teacher" model into a smaller "student" model and then serving it efficiently. The project includes custom GPU kernels written in Triton and CUDA for optimized operations like FlashAttention and fused computations. For the training phase, it features knowledge distillation losses, a custom FSDP/ZeRO-3 implementation for sharded training, and atomic, sharded checkpoints. The serving component is robust, offering a block-paged KV cache with a ref-counted allocator for prefix sharing, a continuous-batching scheduler with admission control and preemption, and speculative decoding. It also supports post-training quantization methods like int8 weight-only, AWQ, and FP8. A Rust-based gateway (tokio + axum) handles HTTP requests and integrates with the Python inference engine via PyO3. Additional features include a JAX/XLA reimplementation for parity checks and interpretability helpers like activation hooks and a logit lens, demonstrating a comprehensive approach to operationalizing distilled LLMs.

https://github.com/zengxiao-he/tessera

cudaflash-attentionfsdpinference-enginejaxknowledge-distillationkv-cachellmmechanistic-interpretabilityml-systemspaged-attentionpytorchquantizationrustspeculative-decodingtriton

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.