FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
Pulsar is an inference engine designed to run Mixture-of-Experts models far larger than available VRAM on consumer hardware. Routed experts live on NVMe and stream per token while everything that makes decisions — attention, KV cache, routers — stays resident in GPU memory. It is written in Rust and CUDA with no llama.cpp anywhere in the stack, and is the successor to the author's earlier NeutronStar C fork. Multi-GPU placement is zero-config: the engine measures PCIe bandwidth and places attention and hot experts where they fit. Eleven model architectures are supported, including Hy3 295B, GLM-5.2 743B with MLA and DSA sparse attention, Kimi K2.7 at roughly 1T parameters, MiniMax M3, Gemma 4 26B-A4B with interleaved sliding-window attention, TML Inkling 1T, Qwen3-235B-A22B, DeepSeek-V4-Flash 284B, Qwen3.6-35B-A3B with Gated DeltaNet linear attention, gpt-oss 20B with per-head attention sinks, Laguna-S-2.1 118B and Ornith-397B. The reference machine is an RTX 5060 Ti 16GB paired with an RTX 4060 Ti 16GB, a Ryzen 9900X, 30 GB of RAM and one Gen5 NVMe; published figures include GLM-5.2 at 2.7 tokens per second, Hy3 295B at 6.0 and Qwen3.6-35B-A3B at 51.8, all as sustained warm decode at n=64 and temperature 0. A popularity census built over the first full run determines which experts stay in the resident tier, so throughput measured cold is not representative. Dense models take a different path entirely — per-layer card ownership with K-quant-native attention and speculative decode through the model's own MTP layer. Long-context decoding on GQA models uses a split-K attention kernel that fans the position scan across blocks past 4k visible rows. Benchmarks come from a single scripted harness and the documentation records known gaps and unbenchmarked models rather than omitting them.
https://github.com/giannisanni/pulsar
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.