FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
Claw Compactor is an open-source, multi-stage LLM token compression engine designed to significantly reduce the cost and improve the efficiency of large language model inference. It leverages a 14-stage Fusion Pipeline, where each stage specializes in a particular compression technique – ranging from AST-aware code analysis, JSON statistical sampling, to simhash-based deduplication. The pipeline processes content through an immutable data flow, ensuring that each stage's output feeds the next. Key features include content-aware routing, which automatically detects content types (code, JSON, logs) and languages to apply appropriate compression strategies, and reversible compression, allowing for the retrieval of original content. Unlike other methods that discard tokens based on perplexity, Claw Compactor maintains content integrity, making it suitable for structured data like code and JSON. This tool offers substantial token reduction (15-82%) with zero additional LLM inference cost and low latency, making it an effective solution for optimizing LLM context windows and reducing operational expenses for AI applications and agents.
https://github.com/open-compress/claw-compactor
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.
RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.