FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
xFasterTransformer provides an exceptionally optimized solution for Large Language Model (LLM) inference on X86 platforms, specifically targeting Intel Xeon processors. It leverages the hardware capabilities of these platforms to achieve high performance and scalability for LLM inference on both single-socket and multi-socket/multi-node configurations. The project aims to accelerate the serving of various popular LLM models, including DeepSeek, ChatGLM, Llama, Baichuan, Qwen, Opt, Gemma, and Mixtral, supporting different data types like FP16, BF16, INT8, and INT4. It offers both C++ and Python APIs, ranging from high-level to low-level interfaces, to facilitate easy adoption and integration into existing solutions or services. xFasterTransformer also includes utilities for converting Hugging Face models to its optimized format and provides support for integration with serving frameworks like vLLM, FastChat, and MLServer, including an OpenAI-compatible server. The project explicitly focuses on the inference and serving phase of LLMs, providing benchmarks and examples to showcase performance and usage.
https://github.com/intel/xFasterTransformer
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.