FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
GGRUN (gguf run) is a powerful launcher designed to simplify and optimize the deployment of GGUF-formatted LLMs using llama.cpp or its faster ik_llama.cpp fork. It serves as an alternative to Ollama, particularly for complex multi-GPU setups. The tool automates the intricate process of configuring model loading and execution by intelligently analyzing available hardware (GPUs, RAM, PCIe topology) to determine optimal multi-GPU tensor-split and Mixture-of-Experts (MoE) expert placement. Instead of manual flag tuning, GGRUN features "AI Tune," which benchmarks and caches the fastest valid flag sets per model and hardware configuration. It includes a Hugging Face downloader with hardware-aware quantization selection and supports speculative decoding (MTP, EAGLE-3) and vision projectors. GGRUN exposes an OpenAI-compatible API, enabling seamless integration with existing LLM applications. It also provides a user-friendly TUI for browsing and downloading models, adjusting settings, and launching, making it accessible for both CLI and interactive use. With crash recovery and backend fallback mechanisms, GGRUN enhances the robustness and performance of local LLM inference, especially on heterogeneous multi-GPU systems where Ollama's conservative heuristics might fall short. The project emphasizes performance, demonstrating significant speed improvements over Ollama and raw llama.cpp for various models and configurations.
https://github.com/raketenkater/ggrun
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.