FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
OpenArc is an inference engine specifically designed for Intel hardware, leveraging OpenVINO to optimize and accelerate the serving of diverse AI models. It provides OpenAI-compatible API endpoints for large language models (LLMs), vision-language models (VLMs), speech-to-text models (Whisper, Qwen-ASR), text-to-speech models (Kokoro-TTS, Qwen-TTS), embedding models, and rerankers. Key features include multi-GPU pipeline parallelism, CPU offload/hybrid device support, and NPU device compatibility. The engine supports speculative decoding and streaming cancellation, and allows for model concurrency, enabling multiple models to be loaded and inferred simultaneously. It also includes metrics for benchmarking inference performance, such as time to first token, prefill throughput, and decode throughput. OpenArc aims to simplify access, deployment, and utilization of OpenVINO acceleration for various AI use cases, offering containerization with Docker and jinja templating with AutoTokenizers. It stands on the shoulders of projects like Optimum-Intel, OpenVINO, llama.cpp, vLLM, and Transformers.
https://github.com/SearchSavior/OpenArc
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.