FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.
RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.
EmbedAnything is a highly performant, modular, and memory-safe Rust-based pipeline for generating multimodal embeddings and streaming them to vector databases, supporting various sources and infere...
Service Streamer is middleware that optimizes deep learning model inference by batching discrete web requests into mini-batches, significantly boosting GPU utilization and overall system performanc...
FeatherCNN is a high-performance, lightweight inference engine for convolutional neural networks, specifically optimized for ARM CPUs on mobile and embedded devices.
Qualcomm AI Hub Models provides pre-optimized machine learning models for efficient deployment and inference on Qualcomm hardware, offering tools for compilation, quantization, profiling, and runni...
Paddle.js is a browser-based deep learning inference engine for Baidu PaddlePaddle, enabling model loading and execution directly in web environments with WebGL, WebGPU, and WebAssembly support.
TurboOCR is a high-performance, GPU-accelerated OCR server designed for fast and accurate text extraction from images and PDFs, leveraging TensorRT and PP-OCRv5.
FlashRank is an ultra-lite and super-fast Python library designed to re-rank search results in RAG pipelines using state-of-the-art LLMs and cross-encoders without requiring PyTorch or Hugging Face...
ZhiLight is a highly optimized LLM inference acceleration engine for Llama and its variants, developed by Zhihu and ModelBest Inc., designed for efficient deployment on various NVIDIA GPUs.
Savant is an open-source framework for building high-performance, real-time multimedia AI applications, specifically computer vision and video analytics pipelines, on Nvidia hardware for both edge ...
Adlik is an end-to-end framework for optimizing and accelerating deep learning inference across cloud, edge, and device environments.
Msnhnet is a lightweight, C++ inference framework for deploying PyTorch models, supporting various architectures like YOLO and ResNet on CPU and GPU with optimizations for embedded devices.
Native MLX port of the DSpark and DFlash speculative decoding drafters, giving lossless multi-fold faster LLM decoding on Apple Silicon with an OpenAI-compatible serving mode.
Timber is an AOT compiler that transforms classical ML models (XGBoost, LightGBM, scikit-learn, CatBoost, ONNX) into native C99 inference code for extremely fast, portable, and low-overhead model s...
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX or...
Forward is a high-performance deep learning inference acceleration framework developed by Tencent, leveraging TensorRT for optimized deployment of models on NVIDIA GPUs with support for major frame...
OpenArc is an inference engine for Intel devices, enabling the serving of various AI models like LLMs, VLMs, Whisper, and embedding models via OpenAI-compatible endpoints with OpenVINO acceleration.
Krasis is a hybrid LLM runtime focused on efficiently running large Mixture-of-Experts models on consumer-grade NVIDIA GPUs with limited VRAM.
OME (Open Model Engine) is a Kubernetes operator designed for enterprise-grade management, deployment, and serving of Large Language Models (LLMs), optimizing resource utilization and supporting va...
JetStream is an optimized engine for large language model (LLM) inference on XLA devices, primarily TPUs, focusing on throughput and memory efficiency.
xFasterTransformer is an optimized inference solution for Large Language Models on Intel Xeon platforms, leveraging hardware capabilities for high performance and scalability.
ParaAttention accelerates Diffusion Transformer (DiT) model inference through context parallel attention and dynamic caching, supporting Ulysses and Ring-style parallelism.
Native Swift inference engine that runs 100GB+ open LLMs on Macs with only 16-64GB of memory by streaming model weights from SSD as needed.
High-performance C++ audio inference framework built on `ggml` for local AI models, supporting TTS, ASR, voice conversion, and more with a full-task WebUI and optimized CUDA performance.
MoE-Infinity is a PyTorch library for cost-effective, fast, and easy serving of Mixture-of-Experts (MoE) Large Language Models, optimizing inference on memory-constrained GPUs.
SwiftLLM is a compact, high-performance LLM inference system designed for research, offering vLLM-equivalent performance with a significantly smaller codebase for easy understanding and modification.
GGRUN is an auto-tuning launcher for GGUF models on llama.cpp/ik_llama.cpp, providing an OpenAI-compatible server with multi-GPU tensor splitting, MoE expert placement, AI-tuned flag optimization, ...
Rust and CUDA inference engine for giant Mixture-of-Experts models that streams routed experts from NVMe per token, keeping only attention and hot experts in VRAM on consumer GPUs.
Auto is an AGI compiler that records LLM agent behavior, identifies repeatable patterns, and compiles them into verified, sandboxed WebAssembly binaries for efficient execution.