Awesome Infra for AI › Inference Optimization

StarlightSearch/EmbedAnything

⭐ 1310 Rust repository created 2024-03-31

EmbedAnything provides a minimalist, yet high-performance, Rust-built pipeline for generating embeddings from diverse data sources, including text, images, audio, and documents. Its core strength lies in its ability to efficiently process and stream these embeddings to vector databases, leveraging Rust for memory safety and speed. Key features include multi-modality support, enabling the processing of various data types, and compatibility with different inference backends like Candle, ONNX Runtime, and cloud models. The project emphasizes low memory footprint and ease of deployment, partly due to its lack of a PyTorch dependency. It incorporates advanced functionalities such as in-built chunking methods (semantic, late-chunking), GPU acceleration, and direct integration with AWS S3 buckets. A notable aspect is 'Vector Streaming,' which separates document preprocessing from model inference to reduce latency and improve throughput, achieving an efficient, concurrent workflow. This tool is designed for production-ready environments, offering modularity for various vector database integrations and supporting dense, sparse, ONNX, model2vec, and late-interaction embeddings. It also provides examples for utilizing indexes in search agents.

https://github.com/StarlightSearch/EmbedAnything

aicloudgenerative-aihacktoberfesthigh-performanceindexinginferenceinformation-retrievallarge-language-modelslocalmachine-learningonnxruntimepipelineproduction-readypythonragrustsearchservervector-databaseembeddingsmultimodal

Also in Inference Optimization

FareedKhan-dev/kimi-k3-in-c

kimi-k3-in-c is a portable C99 inference engine designed to run the Kimi K3 2.78-trillion-parameter LLM on a single CPU with minimal RAM, focusing on extreme memory efficiency without GPUs or exter...

zilliztech/GPTCache

GPTCache is a library that creates a semantic cache for LLM queries, significantly reducing API costs and improving response times by storing and reusing previous LLM responses.

vllm-project/vllm-ascend

vLLM Ascend is a community-maintained hardware plugin for seamlessly running vLLM and various large language models on Ascend NPUs, optimizing inference performance.

youssofal/MTPLX

MTPLX is a macOS-native inference engine for Apple Silicon that leverages multi-token prediction (MTP) for accelerated local LLM serving, offering significant speed improvements.

open-compress/claw-compactor

Claw Compactor is an LLM token compression engine that uses a 14-stage Fusion Pipeline for content-aware, reversible compression to reduce LLM inference costs and optimize context windows.

siliconflow/onediff

OneDiff is an acceleration library for diffusion models, providing out-of-the-box performance optimizations for popular UIs and libraries like Hugging Face Diffusers and ComfyUI.

ovg-project/kvcached

kvcached is a KV cache library that introduces virtual memory abstraction for LLM serving on shared GPUs, enabling elastic and demand-driven KV cache allocation for improved GPU utilization under d...

nobodywho-ooo/nobodywho

NobodyWho is an efficient, on-device inference engine that enables local execution of LLMs and SLMs across various platforms without requiring API keys.