FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
Awesome Infra for AI › Inference Optimization
FlashRank is a Python library focused on optimizing the re-ranking phase within search and retrieval pipelines, particularly for Retrieval Augmented Generation (RAG) applications. It provides an efficient and lightweight solution for reordering initial search results to improve relevance before feeding them into Large Language Models (LLMs). The library emphasizes minimal resource consumption, offering models as small as ~4MB, and is designed for speed, making it suitable for latency-sensitive, user-facing scenarios. FlashRank supports both pairwise/pointwise (cross-encoder based) and listwise (LLM based) re-rankers, including various pre-trained models like TinyBERT, MiniLM, and specialized models for multilingual or domain-specific use cases. A key feature is its independence from heavy dependencies like PyTorch or Hugging Face Transformers, allowing it to run efficiently on CPUs and in serverless environments, reducing cold start times and operational costs. The project aims to provide competitive re-ranking performance with a focus on efficiency and cost-effectiveness, making it a valuable tool for optimizing inference throughput in LLM applications.
https://github.com/PrithivirajDamodaran/FlashRank
kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...
GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.
vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.
MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.
Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.
OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.
kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...
NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.