FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
Qualcomm AI Hub Models is a collection of state-of-the-art machine learning models specifically optimized for performance (latency, memory) and designed for deployment on Qualcomm devices. The project facilitates the end-to-end process of preparing and running AI models on embedded and mobile hardware. It includes functionalities for compiling models for target runtimes like Qualcomm AI Engine Direct, LiteRT (TensorFlow Lite), and ONNX, with support for various quantizations (FP32, FP16, INT16, INT8) suitable for CPU, GPU, and NPU architectures found in Snapdragon chipsets. Users can utilize the associated Qualcomm AI Hub Workbench to compile, quantize, and profile models on cloud-hosted physical Qualcomm devices, ensuring optimal performance and power efficiency. The platform also enables running inference with sample inputs and comparing on-device outputs with PyTorch for validation. Beyond compilation, the repository offers end-to-end model demos and Python applications that wrap model inference with pre- and post-processing steps, simplifying the integration of these optimized models into real-world applications. The primary focus is on enabling efficient, high-performance AI inference on Qualcomm's broad range of mobile, automotive, and IoT platforms.
https://github.com/qualcomm/ai-hub-models
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.