Awesome Infra for AI › Inference Optimization

zilliztech/GPTCache

⭐ 8210 Python repository created 2023-03-24

GPTCache is designed to optimize the performance and cost-efficiency of applications leveraging Large Language Models (LLMs). The library implements a semantic cache that stores responses from LLM API calls. When a new query arrives, GPTCache first checks its cache for semantically similar previous queries. If a sufficiently similar query and its response are found, the cached response is returned immediately, bypassing the need to call the LLM API again. This mechanism drastically cuts down on API costs associated with repeated or similar queries and substantially reduces the response latency, as retrieving from a local cache is much faster than an external API call. The project integrates seamlessly with popular LLM frameworks like LangChain and LlamaIndex, making it easy to incorporate into existing AI applications. It supports both exact and similar query matching for caching, utilizing embedding models (like ONNX) and vector databases (such as FAISS, integrated via Milvus in some contexts) to manage the semantic search for cached responses. Users can configure parameters like 'temperature' to control the caching behavior, influencing the likelihood of re-querying the LLM versus using a cached response. GPTCache supports various storage layers for its cache, including SQLite for metadata and different vector databases for embeddings. By providing a server docker image, GPTCache also enables integration with applications developed in any language, extending its utility beyond Python environments.

https://github.com/zilliztech/GPTCache

semantic cacheLLM cachingopenailangchainllama_indexcost optimizationperformancevector searchembeddingmilvusredisfaiss

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.

alibaba/rtp-llm

RTP-LLM is Alibaba's high-performance inference engine for Large Language Models, designed to accelerate the serving of diverse LLM applications in production environments.