Awesome Infra for AI › Inference Optimization

Inference Optimization

37 projects

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

⭐ 8883 C added 2026-08-03

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

⭐ 8210 Python

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

⭐ 2920 C++

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

⭐ 2518 Python

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

⭐ 2006 Python

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

⭐ 1963 Jupyter Notebook

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

⭐ 1515 Python

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.

⭐ 1515 Rust

alibaba/rtp-llm

RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.

⭐ 1356 Cuda

StarlightSearch/EmbedAnything

EmbedAnything is a highly performant, modular, and memory-safe Rust-based pipeline for generating multimodal embeddings and streaming them to vector databases, supporting various sources and infere...

⭐ 1310 Rust

ShannonAI/service-streamer

Service Streamer is middleware that optimizes deep learning model inference by batching discrete web requests into mini-batches, significantly boosting GPU utilization and overall system performanc...

⭐ 1241 Python

Tencent/FeatherCNN

FeatherCNN is a high-performance, lightweight inference engine for convolutional neural networks, specifically optimized for ARM CPUs on mobile and embedded devices.

⭐ 1228 C++

qualcomm/ai-hub-models

Qualcomm AI Hub Models provides pre-optimized machine learning models for efficient deployment and inference on Qualcomm hardware, offering tools for compilation, quantization, profiling, and runni...

⭐ 1226 Python

PaddlePaddle/Paddle.js

Paddle.js is a browser-based deep learning inference engine for Baidu PaddlePaddle, enabling model loading and execution directly in web environments with WebGL, WebGPU, and WebAssembly support.

⭐ 1106 JavaScript

aiptimizer/TurboOCR

TurboOCR is a high-performance, GPU-accelerated OCR server designed for fast and accurate text extraction from images and PDFs, leveraging TensorRT and PP-OCRv5.

⭐ 1104 C++

PrithivirajDamodaran/FlashRank

FlashRank is an ultra-lite and super-fast Python library designed to re-rank search results in RAG pipelines using state-of-the-art LLMs and cross-encoders without requiring PyTorch or Hugging Face...

⭐ 1008 Python

zhihu/ZhiLight

ZhiLight is a highly optimized LLM inference acceleration engine for Llama and its variants, developed by Zhihu and ModelBest Inc., designed for efficient deployment on various NVIDIA GPUs.

⭐ 908 C++

insight-platform/Savant

Savant is an open-source framework for building high-performance, real-time multimedia AI applications, specifically computer vision and video analytics pipelines, on Nvidia hardware for both edge ...

⭐ 860 Python

Adlik/Adlik

Adlik is an end-to-end framework for optimizing and accelerating deep learning inference across cloud, edge, and device environments.

⭐ 805 C++

msnh2012/Msnhnet

Msnhnet is a lightweight, C++ inference framework for deploying PyTorch models, supporting various architectures like YOLO and ResNet on CPU and GPU with optimizations for embedded devices.

⭐ 738 C++

ARahim3/mlx-dspark

Native MLX port of the DSpark and DFlash speculative decoding drafters, giving lossless multi-fold faster LLM decoding on Apple Silicon with an OpenAI-compatible serving mode.

⭐ 697 Python added 2026-08-24

kossisoroyce/timber

Timber is an AOT compiler that transforms classical ML models (XGBoost, LightGBM, scikit-learn, CatBoost, ONNX) into native C99 inference code for extremely fast, portable, and low-overhead model s...

⭐ 689 Python

zengxiao-he/tessera

From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX or...

⭐ 565 Python added 2026-06-22

Tencent/Forward

Forward is a high-performance deep learning inference acceleration framework developed by Tencent, leveraging TensorRT for optimized deployment of models on NVIDIA GPUs with support for major frame...

⭐ 557 C++

SearchSavior/OpenArc

OpenArc is an inference engine for Intel devices, enabling the serving of various AI models like LLMs, VLMs, Whisper, and embedding models via OpenAI-compatible endpoints with OpenVINO acceleration.

⭐ 536 Python

brontoguana/krasis

Krasis is a hybrid LLM runtime focused on efficiently running large Mixture-of-Experts models on consumer-grade NVIDIA GPUs with limited VRAM.

⭐ 523 C++

ome-projects/ome

OME (Open Model Engine) is a Kubernetes operator designed for enterprise-grade management, deployment, and serving of Large Language Models (LLMs), optimizing resource utilization and supporting va...

⭐ 514 Go

AI-Hypercomputer/JetStream

JetStream is an optimized engine for large language model (LLM) inference on XLA devices, primarily TPUs, focusing on throughput and memory efficiency.

⭐ 464 Python

intel/xFasterTransformer

xFasterTransformer is an optimized inference solution for Large Language Models on Intel Xeon platforms, leveraging hardware capabilities for high performance and scalability.

⭐ 435 C++

chengzeyi/ParaAttention

ParaAttention accelerates Diffusion Transformer (DiT) model inference through context parallel attention and dynamic caching, supporting Ulysses and Ring-style parallelism.

⭐ 434 Python

carloslfu/slotstream

Native Swift inference engine that runs 100GB+ open LLMs on Macs with only 16-64GB of memory by streaming model weights from SSD as needed.

⭐ 413 Swift added 2026-09-28

kigner/audio.cpp-webui

High-performance C++ audio inference framework built on `ggml` for local AI models, supporting TTS, ASR, voice conversion, and more with a full-task WebUI and optimized CUDA performance.

⭐ 376 C++ added 2026-07-27

EfficientMoE/MoE-Infinity

MoE-Infinity is a PyTorch library for cost-effective, fast, and easy serving of Mixture-of-Experts (MoE) Large Language Models, optimizing inference on memory-constrained GPUs.

⭐ 367 Python

interestingLSY/swiftLLM

SwiftLLM is a compact, high-performance LLM inference system designed for research, offering vLLM-equivalent performance with a significantly smaller codebase for easy understanding and modification.

⭐ 358 Python

raketenkater/ggrun

GGRUN is an auto-tuning launcher for GGUF models on llama.cpp/ik_llama.cpp, providing an OpenAI-compatible server with multi-GPU tensor splitting, MoE expert placement, AI-tuned flag optimization, ...

⭐ 281 Go added 2026-06-22

giannisanni/pulsar

Rust and CUDA inference engine for giant Mixture-of-Experts models that streams routed experts from NVMe per token, keeping only attention and hot experts in VRAM on consumer GPUs.

⭐ 224 Rust added 2026-08-10

RightNow-AI/auto

Auto is an AGI compiler that records LLM agent behavior, identifies repeatable patterns, and compiles them into verified, sandboxed WebAssembly binaries for efficient execution.

⭐ 128 Rust added 2026-07-27