FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
vLLM Ascend (`vllm-ascend`) is a specialized hardware plugin designed to integrate vLLM with Huawei's Ascend Neural Processing Units (NPUs). This project's core purpose is to enable efficient and seamless deployment of large language models (LLMs) and other Transformer-like models on Ascend hardware, addressing the specific challenges of accelerated inference on this architecture. It operates as a hardware-pluggable interface, decoupling the NPU integration from the main vLLM framework, adhering to a design principle that promotes modularity and hardware abstraction. By leveraging `vllm-ascend`, users can deploy popular open-source models, including Transformer-based, Mixture-of-Experts (MoE), Embedding, and Multi-modal LLMs, directly on Ascend NPUs with optimized performance. The project provides documentation, community support channels, and regular updates, ensuring compatibility with the latest vLLM versions and Ascend CANN software stacks. It supports various Ascend hardware series, such as Atlas 800I and Atlas A3, targeting both inference and training use cases. This plugin is crucial for organizations and researchers looking to utilize Ascend's computational power for AI inference by providing a performance-optimized, production-ready serving solution.
https://github.com/vllm-project/vllm-ascend
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.
RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.