FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
Slotstream is an open-source inference engine, written entirely in native Swift on Apple's MLX and Metal frameworks, designed to run large open-weight language models on Macs whose memory is too small to hold the full model. Its flagship demonstration runs a 125-billion-parameter model (Qwen3.8-Flash-Next) on Macs with 16 to 64GB of RAM by keeping most of the model on SSD storage and loading only the parts needed for each step as it computes, rather than requiring the whole model to be resident in memory. The authors report a measured 15.86 tokens per second on a 48GB M5 Pro Mac at a 22GB memory target, with the measurement method, raw data, and even failed experiments published for scrutiny. The engine exposes Ollama-, OpenAI-, and Anthropic-compatible APIs as well as a native Swift library, and a CLI command can launch existing coding agents such as Claude Code, Codex, Pi, opencode, and Hermes directly against the local model. After an initial model download, it runs fully offline with no Python dependency and no cloud account. The project explicitly targets the common case of 16-64GB Apple Silicon Macs rather than machines with 96GB or more, where the model already fits in memory and other engines that keep it fully resident report faster responses; it links to alternatives better suited to that hardware. The project is also the inference engine behind an in-development consumer app, Sevra, though Slotstream's CLI, APIs, and library remain independently usable.
https://github.com/carloslfu/slotstream
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.