FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
Auto is a novel AGI compiler system designed to optimize the execution of LLM agents by converting their observed behaviors into efficient, low-cost WebAssembly (Wasm) binaries. It functions by treating frontier large language models as "interpreters" and building a compilation layer on top. The system records an agent's actions, identifies symbolic and repeatable patterns, extracts these, and then distills the remaining fuzzy parts into smaller, specialized models. The entire process is verified against a behavioral contract, ensuring the compiled output (a `.cbin` artifact) is correct and capability-confined. The core of Auto involves a tiered runtime: Tier-1 executes the compiled fast path, while Tier-0 (a frontier model) handles novel inputs. If a novel input is encountered, the system "deopts" to Tier-0, captures the resulting trace, and recompiles it into Tier-1, effectively ensuring that "nothing is figured out twice" – a concept referred to as "the ratchet." This significantly reduces marginal cost and latency for repeated tasks. The compiled artifacts are content-addressed and include a manifest reporting measured evaluation scores, cost/latency bounds, capability requirements, and provenance, all grounded in actual measured numbers. Auto's compilation pipeline includes symbolic extraction, distillation into small specialists, verification (including differential testing), and optimization. Guards use calibrated abstention based on distance metrics to determine novelty. The project emphasizes capability confinement, with Wasm binaries declaring zero imports, ensuring they cannot exceed their declared effects. Measured results demonstrate substantial cost savings and latency improvements compared to relying solely on frontier models.
https://github.com/RightNow-AI/auto
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.