Awesome Infra for AI › Inference Optimization

FareedKhan-dev/kimi-k3-in-c

⭐ 8883 C added to this list on 2026-08-03 repository created 2026-08-01

kimi-k3-in-c is an highly optimized, zero-dependency C99 inference engine specifically engineered to run the massive 2.78-trillion-parameter Kimi K3 LLM. Its core innovation lies in its ability to perform inference on a single CPU, utilizing as little as 8.24 GB of RAM, a remarkable feat given the model's 1.56 TB checkpoint size. The project achieves this through a series of "four reductions" and advanced techniques such as streaming parts of the model (trunk) and routing experts directly from their packed 4-bit form on disk, effectively making the model's footprint on resident memory tiny. It boasts full portability, requiring no BLAS libraries, AI frameworks, or GPUs, making it suitable for resource-constrained environments. The engine focuses on memory-efficient inference, using quantization (MXFP4) and custom kernels. It includes tools for downloading, packing, and analyzing the model, alongside a comprehensive validation suite to ensure byte-identical output across different memory budgets. The project emphasizes the systems programming aspect, providing detailed insights into how it manages to fit such a large model into limited memory, a critical aspect for on-device or edge AI inference scenarios.

https://github.com/FareedKhan-dev/kimi-k3-in-c

llmllm-inferencecpu-inferencememory-efficientquantizationc99from-scratchsystems-programminginference-enginemoemixture-of-expertstransformerzero-dependencieskimi-k3avx2simd

Also in Inference Optimization

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.

alibaba/rtp-llm

RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.