Awesome Infra for AI › Vector Databases & Retrieval Infrastructure

LibreChat-AI/rag-api

⭐ 912 Python added to this list on 2026-09-28 repository created 2024-03-17

ID-based RAG FastAPI is a retrieval-augmented-generation backend that combines LangChain with FastAPI to provide asynchronous document indexing and retrieval backed by PostgreSQL with the pgvector extension. Documents are organized into embeddings keyed by a caller-supplied file_id, which lets consumers run targeted queries scoped to a specific file when that id is combined with metadata stored elsewhere, rather than searching an entire undifferentiated vector store. The API exposes endpoints to add, retrieve, and delete documents, and its primary integration target is the LibreChat chat application, though the authors note it can serve any ID-based retrieval use case. A notable part of the current README documents a security hardening pass: earlier versions authorized retrieval and deletion routes by caller-supplied file_id alone, or authorized an entire result set from only the top-ranked hit, which allowed a caller who could reach a multi-document query endpoint or list file ids to read or delete content belonging to other users. The fix introduces an explicit owner-scope check, derived from the verified caller token in one shared module, applied to the vector-store query before ranking, and scopes ingestion rollback to the specific ingestion attempt rather than the whole file id. The project plans to keep evolving its querying, re-ranking, embedding-model, and vector-store choices over time. It is aimed at teams building chat or document-QA applications who need a lightweight, self-hostable RAG retrieval layer with per-file, per-owner data isolation on top of pgvector.

https://github.com/LibreChat-AI/rag-api

ragretrievalpgvectorlangchainfastapiembeddings

Also in Vector Databases & Retrieval Infrastructure

pathwaycom/llm-app

Ready-to-run cloud templates for building real-time RAG, AI pipelines, and enterprise search applications that synchronize with various live data sources.

run-llama/llama_index

LlamaIndex is an open-source data framework for building LLM applications by connecting custom data sources to large language models, focusing on data ingestion, indexing, and retrieval augmented g...

milvus-io/milvus

Milvus is a high-performance, cloud-native vector database designed for scalable vector Approximate Nearest Neighbor (ANN) search, efficiently organizing and searching vast amounts of unstructured ...

VectifyAI/PageIndex

PageIndex is a vectorless, reasoning-based RAG system that builds hierarchical tree indexes from documents and uses LLMs to reason over them for context-aware retrieval.

qdrant/qdrant

Qdrant is an open-source, high-performance vector similarity search engine and vector database designed specifically for AI applications, enabling fast storage, search, and management of vectors wi...

Tencent/WeKnora

WeKnora is an open-source, LLM-powered knowledge framework for enterprise document understanding, semantic retrieval, and autonomous reasoning, featuring RAG, ReAct agents, and an auto-maintaining ...

topoteretes/cognee

Cognee is an open-source AI memory platform that provides AI agents with persistent long-term memory through a self-hosted knowledge graph, combining vector embeddings and graph reasoning.

RyanCodrai/turbovec

TurboVec is a Rust-based approximate nearest neighbor (ANN) vector index with Python bindings, built on Google Research's TurboQuant algorithm for efficient, memory-optimized vector similarity search.