FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
Timber is a specialized Ahead-of-Time (AOT) compiler designed to significantly accelerate inference for classical machine learning models. It supports popular frameworks like XGBoost, LightGBM, scikit-learn, CatBoost, and ONNX (specifically tree ensembles, linear models, SVMs, k-NN, Naive Bayes, GPR, Isolation Forest). The core functionality involves converting these trained models into highly optimized, self-contained C99 inference artifacts with zero runtime dependencies. This process includes a multi-pass optimizing compiler that applies techniques like dead-leaf elimination, threshold quantization, constant-feature folding, and branch sorting. The generated C99 code is then compiled into a shared library, which can be served via a built-in HTTP server featuring an Ollama-compatible API. This allows for one-command model serving with microsecond-level latency, dramatically outperforming Python-based inference environments. Timber is particularly valuable for scenarios requiring fast, predictable, and portable inference, such as fraud detection systems, edge and IoT deployments, and regulated industries (finance, healthcare, automotive) needing auditable and deterministic inference. It also offers hardware acceleration backends for SIMD, GPU, FPGA, and embedded targets, along with certification reports and air-gapped deployment bundles, making it a comprehensive tool for high-performance and safety-critical ML model deployment.
https://github.com/kossisoroyce/timber
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.