Awesome Infra for AI › Inference Optimization

brontoguana/krasis

⭐ 523 C++ repository created 2026-02-09

Krasis is an LLM runtime engineered to facilitate the execution of multi-hundred-billion-parameter Mixture-of-Experts (MoE) models on commodity NVIDIA GPUs, even those with constrained VRAM. It achieves this through a sophisticated hybrid architecture that prioritizes fast GPU prompt processing and GPU-executed decode. A key innovation is the Hybrid Cache-aware Scheduling (HCS) system, which intelligently manages the residency of 'hot' and 'cold' experts between VRAM and CPU RAM, enabling models much larger than available VRAM to run locally. The performance-critical runtime path is built in Rust/CUDA, leveraging optimized CUDA kernels, cached quantized weights, and precise VRAM budgeting, while Python handles setup and model loading. The project supports BF16 safetensors from Hugging Face models and includes features like compact KV cache modes (k6v6, k4v4) and HQQ attention support for efficient memory usage. It provides an interactive launcher, an OpenAI-compatible API, and boasts robust VRAM safety systems to prevent OOM errors. Krasis offers a streamlined installation and update process, along with repeatable benchmarks and correctness validation, making it a powerful tool for accelerating LLM inference on resource-limited hardware.

https://github.com/brontoguana/krasis

cpu-inferencegguf-model-supportgpu-inferencehigh-performance-inferencehybrid-inferenceinference-engineinference-optimizationlarge-language-modelsllama-cpp-alternativellm-inferencemixture-of-expertstransformer

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.