FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
ZhiLight is an open-source inference engine co-developed by Zhihu and ModelBest Inc., specifically engineered to accelerate the inference of large language models (LLMs) such as Llama and its derivatives. It is particularly optimized for PCIe-based GPUs and demonstrates significant performance advantages over other mainstream open-source inference engines like vLLM and SGLang. The engine incorporates a suite of advanced features to enhance LLM serving, including an asynchronous OpenAI-compatible API, custom tensor management, unified global memory management, and "dual streams" for encode and all-reduce overlap, supporting INT8-quantized all-reduce. It also features host all-reduce based on SIMD instructions, optimized fused kernels (QKV, residual & layernorm), fused batch attention for decoding with tensor core instructions, and support for dynamic batching, flash attention prefill, chunked prefill, and prefix caching. ZhiLight supports various quantization techniques such as Native INT8, SmoothQuant, FP8, AWQ, and GPTQ, along with the Marlin kernel for GPTQ. It also boasts support for Mixture of Experts (MoE) models, including DeepseekV2 MoE and DeepseekV2 MLA, and is compatible with Llama/Llama2, Mixtral, Qwen2 series, and similar architectures. The project provides Docker images for easy deployment and comprehensive performance benchmarks comparing it against other engines on different NVIDIA GPUs and model sizes.
https://github.com/zhihu/ZhiLight
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.