vllm-project/vllm
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
Awesome Infra for AI › Model Serving Frameworks
fastembed-rs is a Rust library for producing vector embeddings and reranking scores entirely on-device, without calling an external API. It wraps the ort ONNX Runtime bindings for model inference and the Hugging Face tokenizers library for fast text encoding, and exposes a synchronous API with no dependency on an async runtime like Tokio, making it easy to embed inside CLIs, servers, or edge applications. It supports four kinds of models: dense text embeddings (BGE, MiniLM, mpnet, E5, GTE, Nomic, Jina, Snowflake Arctic Embed, Qwen3 Embedding, EmbeddingGemma, and more), sparse text embeddings (SPLADE++, BGE-M3, OpenSearch neural sparse), image embeddings (CLIP-style vision models, ResNet), and cross-encoder reranking models (BGE reranker, Jina reranker). Many models ship in quantized ONNX variants for lower memory use and faster inference, and some newer models route through a candle backend behind optional Cargo features. Typical usage loads a model by enum variant and calls an embed function on a batch of strings to get fixed-size float vectors, which downstream code can index in a vector database or compare directly for similarity search. The project is the Rust counterpart to the original Python fastembed library, with sibling ports in Go and JavaScript, so its output format and default models are compatible across languages. It targets application and infrastructure developers who need to generate embeddings inside a Rust codebase, for retrieval-augmented generation, semantic search, or reranking pipelines, without shipping a Python inference server or depending on a hosted embedding API. It does not store or index vectors itself; that is left to a separate vector database.
https://github.com/anush008/fastembed-rs
vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.
oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.
A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.
OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.
BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.
vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.
KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.
Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...