Awesome Infra for AI › Model Serving Frameworks

Bessouat40/RAGLight

⭐ 672 Python repository created 2024-12-12

RAGLight is a lightweight and modular Python library that streamlines the implementation of Retrieval-Augmented Generation (RAG). Its primary purpose is to enhance Large Language Models (LLMs) by integrating document retrieval with natural language inference, making it easier to build context-aware AI solutions. The framework emphasizes simplicity and flexibility, providing modular components to seamlessly integrate diverse LLMs (e.g., Ollama, OpenAI, Mistral, Gemini, Bedrock), embedding models (e.g., HuggingFace all-MiniLM-L6-v2), and vector stores (e.g., Chroma, Qdrant). Key features include a robust RAG pipeline, an advanced agentic RAG pipeline for improved performance, and integral support for external tool capabilities via MCP (Microservice Communication Protocol) servers, enabling connections to various data sources and services. It supports flexible document ingestion for different file types and offers an extensible architecture to swap out components as needed. For enhanced retrieval, RAGLight incorporates hybrid search (BM25 + Semantic + RRF) and query reformulation to optimize accuracy in multi-turn conversations through automatic re-writing of follow-up questions. The library also provides full multi-turn conversation history across all supported providers, streaming output for real-time interaction, and observability features through Langfuse for end-to-end tracing of RAG calls. For ease of use, it includes a command-line interface (CLI) for quick setup and interaction, as well as options to deploy as a REST API and use with Docker. RAGLight focuses on the operational and serving aspects of RAG, rather than model training, making it a tool for deploying and managing AI inference workflows.

https://github.com/Bessouat40/RAGLight

agentic-aiagentic-ragagentic-workflowframeworkragretrieval-augmented-generationllm-servingvector-storesobservabilityprompt-managementai-deployment

Also in Model Serving Frameworks

vllm-project/vllm

vLLM is a high-throughput and memory-efficient serving and inference engine for large language models, featuring PagedAttention, continuous batching, and extensive hardware and model support.

jundot/omlx

oMLX is an LLM inference server optimized for Apple Silicon, offering continuous batching, tiered KV caching, and a macOS menu bar interface for managing models locally.

jina-ai/serve

A framework for building and deploying cloud-native AI services, with native support for ML frameworks, high-performance serving, LLM streaming, and Kubernetes/Docker Compose deployment.

bentoml/OpenLLM

OpenLLM allows developers to self-host and run any open-source or custom LLMs as OpenAI-compatible API endpoints in the cloud, streamlining deployment and serving.

bentoml/BentoML

BentoML is an open-source framework for building, shipping, and scaling AI applications, providing tools to serve AI/ML models as production-ready API endpoints.

vllm-project/vllm-omni

vLLM-Omni is an extension of vLLM designed for efficient serving and inference of omni-modality AI models, encompassing text, image, video, and audio data processing.

kserve/kserve

KServe is a standardized, distributed platform for serving generative and predictive AI models on Kubernetes, providing scalable deployment and management.

Michael-A-Kuykendall/shimmy

Shimmy is a pure-Rust WebGPU inference engine providing OpenAI-compatible endpoints for local GGUF models, featuring Airframe engine and TurboShimmy INT4 KV cache compression for efficient GPU util...