FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
audio.cpp-webui is a high-performance C++ audio inference framework designed for efficient and portable operation of local audio AI models, such as TTS (text-to-speech), ASR (automatic speech recognition), voice conversion, speaker diarization, VAD (voice activity detection), and music generation. Built on `ggml`, it offers significant performance advantages over Python-based solutions, often achieving 1.8x-5.0x faster execution and 45%-80% latency reduction for TTS tasks, with strong CUDA optimization. The project provides a full-featured WebUI, making it accessible and user-friendly for various audio tasks without complex Python dependencies. It also features Windows-friendly local launcher scripts, an OpenAI-compatible HTTP API server for integration with other applications, and experimental JSON pipeline support for multi-step workflows. Key highlights include strong parity with Python reference paths, performance-focused execution with reusable sessions and batch inference, portability across different systems, and integrated audio utilities for noise reduction, enhancement, and resampling. The framework focuses on providing optimized and reusable building blocks for rapid integration of new audio models and efficient end-to-end inference paths.
https://github.com/kigner/audio.cpp-webui
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.