Awesome Infra for AI › Model Serving Frameworks

anush008/fastembed-rs

⭐ 1030 Rust added to this list on 2026-09-21 repository created 2023-10-01

fastembed-rs is a Rust library for producing vector embeddings and reranking scores entirely on-device, without calling an external API. It wraps the ort ONNX Runtime bindings for model inference and the Hugging Face tokenizers library for fast text encoding, and exposes a synchronous API with no dependency on an async runtime like Tokio, making it easy to embed inside CLIs, servers, or edge applications. It supports four kinds of models: dense text embeddings (BGE, MiniLM, mpnet, E5, GTE, Nomic, Jina, Snowflake Arctic Embed, Qwen3 Embedding, EmbeddingGemma, and more), sparse text embeddings (SPLADE++, BGE-M3, OpenSearch neural sparse), image embeddings (CLIP-style vision models, ResNet), and cross-encoder reranking models (BGE reranker, Jina reranker). Many models ship in quantized ONNX variants for lower memory use and faster inference, and some newer models route through a candle backend behind optional Cargo features. Typical usage loads a model by enum variant and calls an embed function on a batch of strings to get fixed-size float vectors, which downstream code can index in a vector database or compare directly for similarity search. The project is the Rust counterpart to the original Python fastembed library, with sibling ports in Go and JavaScript, so its output format and default models are compatible across languages. It targets application and infrastructure developers who need to generate embeddings inside a Rust codebase, for retrieval-augmented generation, semantic search, or reranking pipelines, without shipping a Python inference server or depending on a hosted embedding API. It does not store or index vectors itself; that is left to a separate vector database.

https://github.com/anush008/fastembed-rs

embeddingsonnx-runtimererankingrustlocal-inference

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...