pathwaycom/llm-app
Ready-to-run cloud templates for building real-time RAG, AI pipelines, and enterprise search applications that synchronize with various live data sources.
Awesome Infra for AI › Vector Databases & Retrieval Infrastructure
Ready-to-run cloud templates for building real-time RAG, AI pipelines, and enterprise search applications that synchronize with various live data sources.
LlamaIndex is an open-source data framework for building LLM applications by connecting custom data sources to large language models, focusing on data ingestion, indexing, and retrieval augmented g...
Milvus is a high-performance, cloud-native vector database designed for scalable vector Approximate Nearest Neighbor (ANN) search, efficiently organizing and searching vast amounts of unstructured ...
PageIndex is a vectorless, reasoning-based RAG system that builds hierarchical tree indexes from documents and uses LLMs to reason over them for context-aware retrieval.
Qdrant is an open-source, high-performance vector similarity search engine and vector database designed specifically for AI applications, enabling fast storage, search, and management of vectors wi...
WeKnora is an open-source, LLM-powered knowledge framework for enterprise document understanding, semantic retrieval, and autonomous reasoning, featuring RAG, ReAct agents, and an auto-maintaining ...
Cognee is an open-source AI memory platform that provides AI agents with persistent long-term memory through a self-hosted knowledge graph, combining vector embeddings and graph reasoning.
TurboVec is a Rust-based approximate nearest neighbor (ANN) vector index with Python bindings, built on Google Research's TurboQuant algorithm for efficient, memory-optimized vector similarity search.
Weaviate is an open-source, cloud-native vector database for semantic search, combining vector similarity search with keyword filtering, RAG, and reranking capabilities.
Memvid is a portable, single-file memory layer for AI agents, offering instant retrieval and long-term memory without needing complex RAG pipelines or server-based vector databases.
Zvec is an open-source, lightweight, and lightning-fast in-process vector database designed for embedding directly into applications, providing low-latency and scalable similarity search.
LEANN is an innovative, lightweight, and private vector database designed for personal AI, enabling RAG applications with significantly reduced storage requirements by recomputing embeddings on-dem...
txtai is an all-in-one AI framework providing an embeddings database for semantic search and LLM orchestration capabilities like RAG and agentic workflows.
Claude Context provides semantic code search for AI coding agents, allowing them to access the entire codebase as context in a cost-effective manner using a vector database.
LanceDB is an open-source, embedded, and cloud-native vector database designed for fast, scalable, and production-ready multimodal vector search, built on the Lance columnar format.
Orama is a JavaScript search engine providing full-text, vector, and hybrid search capabilities, designed for use in browsers, servers, or edge networks, and supporting RAG pipelines.
Deep Lake is an AI Data Runtime that provides a serverless multimodal datalake with integrated vector search, optimizing data storage and streaming for AI/ML applications and agentic RAG.
DeepSearcher is an open-source tool that combines LLMs and vector databases to perform deep research, evaluation, and reasoning on private data, generating accurate answers and comprehensive reports.
Vespa is an AI search platform for serving and organizing vectors, tensors, text, and structured data, enabling real-time inference and retrieval at any scale.
PostgresML is a PostgreSQL extension that integrates machine learning and AI capabilities directly into the database, enabling in-database inference, RAG pipelines, and vector search with GPU accel...
Airweave is an open-source context retrieval layer for AI agents and RAG systems, unifying data from various sources into an LLM-friendly search interface.
HelixDB is a graph-vector database built in Rust, designed to unify data types like graph, vector, KV, document, and relational data for AI applications and knowledge graphs.
SPTAG is a distributed approximate nearest neighbor (ANN) search library by Microsoft for large-scale vector search scenarios, offering high-quality vector index building, searching, and distribute...
Infinity is an AI-native database designed for LLM applications, offering incredibly fast hybrid search across dense vectors, sparse vectors, tensors, and full-text.
LongMemory (formerly OpenMemory) is a self-hosted cognitive memory engine that gives LLMs and agents durable long-term recall over SQLite or Postgres, with Python and Node SDKs.
Graph-based retrieval engine for LLM memory that scores knowledge units by the strongest evidence path through a four-layer graph instead of ranking chunks by vector similarity alone.
USearch is a high-performance, compact, and broadly compatible single-file similarity search and clustering engine for vectors and texts, primarily focused on user-defined metrics with minimal depe...
An ingestion pipeline that turns messy unstructured documents into persistent, navigable memory for AI agents — parsing, hierarchy extraction, multimodal structuring and graph construction in one pass.
SocratiCode is an open-source codebase context engine that provides deep semantic understanding of entire codebases for AI assistants, enabling hybrid search, dependency graphs, and impact analysis.
Attu is an AI-native GUI for managing Milvus vector databases, offering multi-cluster management, data exploration, vector search, an AI assistant, and monitoring tools.
Trieve is an all-in-one platform for semantic search, recommendations, and Retrieval-Augmented Generation (RAG) offered via API, featuring self-hosting, hybrid search, and bring-your-own-model capa...
Hora is an efficient, Rust-based library offering a collection of approximate nearest neighbor search algorithms for high-performance similarity search.
SAG is an out-of-the-box document retrieval workbench based on the SAG RAG technique, offering a conversational interface, knowledge graph visualization, and advanced retrieval functionalities for ...
Vearch is a cloud-native distributed vector database designed for efficient similarity search of embedding vectors in AI applications, offering hybrid search, performance, scalability, and reliabil...
pgvecto.rs is a PostgreSQL extension written in Rust, purpose-built for scalable, low-latency, and hybrid-enabled vector similarity search directly within Postgres.
Self-hosted persistent memory backend for AI agents that stores decisions and context in a semantic store with a typed knowledge graph, served over MCP, REST, CLI and a dashboard.
SeekStorm is a high-performance, Rust-native search engine offering sub-millisecond vector and lexical search capabilities as an in-process library and multi-tenancy server.
grepai is a privacy-first CLI for semantic code search, enabling AI agents and developers to find relevant code by intent using vector embeddings, drastically reducing token usage.
VectorChord is a PostgreSQL extension designed for scalable, high-performance, and cost-effective vector search, enabling efficient storage and retrieval of billions of vectors.
JVector is an advanced, embedded, graph-based approximate nearest neighbor (ANN) vector search engine for Java, optimized for large-scale, high-dimensional data retrieval.
DingoDB is a distributed multi-modal vector database offering unified SQL (MySQL-compatible) for structured and unstructured data, ensuring high concurrency and ultra-low latency.
Pixeltable is a unified multimodal backend that integrates data storage, model execution, embedding indexing, and serving for AI data applications.
Python SDK for Milvus, an open-source vector database designed for AI applications, enabling seamless interaction for vector storage and similarity search.
NGT is a high-speed approximate nearest neighbor search library for high-dimensional vector data, featuring graph and tree-based methods, quantization techniques, and multi-language support.
A 7-layer memory operating system for LLM agents like Hermes, providing persistent, context-aware memory using Qdrant for vector storage and various other mechanisms for structured facts, session r...
Python client library for the Qdrant vector search engine, facilitating interaction with Qdrant instances for vector storage, search, and remote inference capabilities.
SeaGOAT is a local-first semantic code search engine that uses vector embeddings to enable natural language queries and regular expressions across your codebase without external API calls.
Endee is a high-performance open-source vector database designed for AI search and retrieval workloads, supporting RAG, semantic search, and hybrid retrieval with optimized indexing and execution.
VectorDBBench is a comprehensive benchmark tool for evaluating the performance and cost-effectiveness of various vector databases and cloud services across diverse scenarios.
JamAI Base is an open-source RAG backend platform with an intuitive spreadsheet-like UI, offering built-in LLM, vector embeddings, and reranker orchestration for AI application development.
Voy is a WASM-based vector similarity search engine implemented in Rust, optimized for fast, tiny, and tree-shakable nearest neighbor search on edge servers and in web applications.
Chromem-go is an embeddable vector database for Go, offering a Chroma-like interface with zero third-party dependencies, designed for in-memory operation with optional persistence.
MyScaleDB is a SQL vector database built on ClickHouse, designed for high-performance vector search, filtered search, and full-text search in scalable AI applications.
Motorhead is an LLM memory and information retrieval server that provides API endpoints for managing conversational memory, summarization, and retrieval-augmented generation (RAG) through vector si...
FastAPI service that indexes documents into pgvector embeddings by file ID and serves scoped retrieval queries for RAG applications, built primarily for LibreChat.
A fully local memory engine for AI coding assistants, exposed over MCP: verbatim recall with encrypted local storage, benchmarked retrieval quality and injected memory packs that cut search tokens.
NornicDB is a distributed graph and vector database with temporal MVCC, offering Neo4j Bolt/Cypher and Qdrant gRPC compatibility, designed for AI-native workloads like Graph-RAG and agent memory.
Epsilla is a high-performance, open-source vector database management system focused on scalable and cost-effective similarity search for embedding vectors.
Neum AI is a data platform for managing large-scale vector embedding creation and synchronization to provide context for LLMs through Retrieval Augmented Generation (RAG).
cuVS is an NVIDIA library providing GPU-accelerated algorithms for vector similarity search and clustering, designed to simplify high-performance vector operations.
AutoMem is a graph-vector memory service providing durable, relational, and context-aware long-term memory for AI assistants using a dual-storage layer of FalkorDB and Qdrant.
Reindexer is an embeddable, in-memory, document-oriented database offering high-performance full-text search, k-nearest neighbors (KNN) search, and hybrid search capabilities.
Wax is a Swift-native memory engine for AI agents, offering on-device, single-file storage for documents, embeddings, and structured knowledge with blazing-fast RAG on Apple Silicon.
NucliaDB is an AI search database for unstructured data, built for Retrieval Augmented Generation (RAG), offering hybrid search with vector, full-text, and graph indexes.
A Pythonic vector database for efficient storage and retrieval of embeddings, leveraging DocArray and Jina for scalable solutions locally or in the cloud.
UStore is a multi-modal transactional database designed for AI and semantic search, featuring vector-search integration and APIs for various data types.
Agentic retrieval-augmented generation library for local document collections, combining hybrid vector and full-text search, reranking and multimodal retrieval on an embedded LanceDB store.
EmbedJs is a Node.js RAG framework for building personalized LLM applications by segmenting data, generating embeddings, and integrating with vector databases for optimized retrieval.
A Go library for embedded vector search and semantic embeddings, using llama.cpp and GGUF BERT models, suitable for small to medium-scale applications with GPU acceleration.
NextPlaid is a local-first multi-vector database and indexing engine that powers ColGREP, a semantic code search tool for terminals and coding agents.
Shared, governed memory for fleets of AI agents: agents write plain text, Caura turns it into searchable multi-tenant memory with scoping, trust tiers and cross-agent outcome propagation.
GraphRAG-rs is a high-performance Rust implementation of GraphRAG for building queryable knowledge graphs from documents, offering server, WASM-only, and hybrid deployment architectures.
Embedbase is an AI backend-as-a-service that provides a dead-simple API for LLM interaction and semantic search through hosted embeddings, supporting various LLM providers.
A multithreaded web crawler that converts web pages into markdown files, specifically designed to preprocess data for LLM RAG applications.
Memlayer provides a plug-and-play, persistent memory layer for LLMs and AI agents, enabling intelligent context recall and knowledge extraction through hybrid vector and graph storage.
Ask Astro is an open-source reference implementation of an LLM application architecture, providing a Q&A interface for Airflow and Astronomer, utilizing RAG, prompt orchestration, and feedback loops.
TemporalStore is a Rust-native, time-aware store for LLM agent memory that ingests events and returns a ranked, token-budgeted ContextPack instead of replaying a growing transcript on every turn.
An agent-agnostic memory sidecar that provides persistent memory, layered recall, and knowledge graphing for AI coding agents, integrating with existing systems without modifying agent internals.
Corpus OS provides a wire-first, vendor-neutral protocol suite and SDK for standardizing LLM, Embedding, Vector, and Graph infrastructure for AI frameworks.
Structured context layer for AI agents that builds a searchable, continuously refreshed graph of domain entities and relationships from schemas and connectors.
redevops-rag is a compact library and CLI tool providing a hybrid RAG pipeline with DuckDB for vector search, BM25, Reciprocal Rank Fusion, recency/keyword priors, and optional cross-encoder rerank...