Awesome Infra for AI › Vector Databases & Retrieval Infrastructure

myscale/MyScaleDB

⭐ 1045 C++ repository created 2024-03-14

MyScaleDB is a specialized SQL vector database for AI applications, forking ClickHouse to integrate high-performance vector search and full-text search capabilities. It allows developers to manage and process large volumes of data for Generative AI applications using familiar SQL syntax, including vector-related functions. Key features include fully SQL compatibility, production readiness for AI applications by unifying structured data, text, vector, JSON, geospatial, and time-series data. It emphasizes improved RAG accuracy through combined vector and rich metadata filtering, offering high-precision and efficient filtered searches. MyScaleDB leverages ClickHouse's columnar storage architecture and advanced vector algorithms to deliver unmatched performance and scalability. It positions itself as a unified system that integrates SQL databases, data warehouses, vector databases, and full-text search engines, aiming to reduce infrastructure costs and enable joint data queries and analytics. The project highlights its ability to perform millisecond searches on billion-scale vectors and offers powerful text/vector hybrid search functions. MyScaleDB demonstrates its distinct advantage over general-purpose databases with vector extensions and specialized vector databases by optimizing for joint queries and resource efficiency.

https://github.com/myscale/MyScaleDB

annbig-dataembeddingimage-searchllmmyscaledbragsearch-enginesimilarity-searchsqlsql-vectorunstructured-analyticsvector-searchvectordb

Also in Vector Databases & Retrieval Infrastructure

pathwaycom/llm-app

Ready-to-run cloud templates for building real-time RAG, AI pipelines, and enterprise search applications that synchronize with various live data sources.

run-llama/llama_index

LlamaIndex is an open-source data framework for building LLM applications by connecting custom data sources to large language models, focusing on data ingestion, indexing, and retrieval augmented g...

milvus-io/milvus

Milvus is a high-performance, cloud-native vector database designed for scalable vector Approximate Nearest Neighbor (ANN) search, efficiently organizing and searching vast amounts of unstructured ...

VectifyAI/PageIndex

PageIndex is a vectorless, reasoning-based RAG system that builds hierarchical tree indexes from documents and uses LLMs to reason over them for context-aware retrieval.

qdrant/qdrant

Qdrant is an open-source, high-performance vector similarity search engine and vector database designed specifically for AI applications, enabling fast storage, search, and management of vectors wi...

Tencent/WeKnora

WeKnora is an open-source, LLM-powered knowledge framework for enterprise document understanding, semantic retrieval, and autonomous reasoning, featuring RAG, ReAct agents, and an auto-maintaining ...

topoteretes/cognee

Cognee is an open-source AI memory platform that provides AI agents with persistent long-term memory through a self-hosted knowledge graph, combining vector embeddings and graph reasoning.

RyanCodrai/turbovec

TurboVec is a Rust-based approximate nearest neighbor (ANN) vector index with Python bindings, built on Google Research's TurboQuant algorithm for efficient, memory-optimized vector similarity search.