FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
MTPLX is a specialized inference engine designed for running large language models (LLMs) locally on Apple Silicon (M1, M2, M3, M4, M5 Macs). Its core innovation lies in implementing multi-token prediction (MTP) without an external drafter model, utilizing modern LLMs' built-in MTP heads. This technique, based on Leviathan and Chen's rejection sampling theorem, allows the model to draft and verify multiple tokens ahead in a single batched forward pass, resulting in significantly faster decoding while maintaining output distribution accuracy. The project boasts measured speedups of 1.6x to 2.24x compared to traditional autoregressive decoding. MTPLX offers both a user-friendly Mac application and a command-line interface (CLI) for ease of use. The application simplifies model setup, performs hardware checks, recommends suitable models, and includes a fan control mechanism for optimal performance. It features a dashboard to monitor live performance metrics like tokens per second and acceptance rates. MTPLX provides an OpenAI-compatible and Anthropic-compatible API server, enabling integration with various tools and applications like OpenCode, Pi, and Open WebUI. It also includes "Forge," a utility to convert Hugging Face models into MTPLX-ready MTP models and train MTP adapters, with an honest verification step to ensure actual performance gains. The tool prioritizes Apple Silicon-native optimization, using MLX rather than CUDA.
https://github.com/youssofal/MTPLX
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.
RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.