FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
OneDiff is an open-source library designed to significantly accelerate the inference speed of diffusion models. It offers out-of-the-box acceleration for widely used platforms such as Hugging Face Diffusers, ComfyUI, and Stable Diffusion Web UI. The library achieves performance gains through PyTorch code compilation tools and highly optimized GPU kernels specifically tailored for diffusion models, including support for various architectures like SDXL, SVD, DiT, Kolors, PixArt, and Latte. OneDiff can deliver substantial speedups, enabling faster image and video generation and reducing computational costs. It supports NVIDIA GPUs and leverages backends like OneFlow and Nexfort for compilation. The project also provides tools for quality evaluation after acceleration, ensuring that performance improvements do not compromise output fidelity. OneDiff is aimed at users and developers who need to deploy and run diffusion models in production environments with high efficiency and low latency, offering solutions to optimize various aspects of the inference pipeline, including handling new input shapes and online serving scenarios.
https://github.com/siliconflow/onediff
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.
RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.