Awesome Infra for AI › Vector Databases & Retrieval Infrastructure

dingodb/dingo

⭐ 1703 Java repository created 2021-10-14

DingoDB is an open-source, distributed multi-modal vector database designed for high-performance AI applications. It uniquely integrates real-time strong consistency, relational semantics, and vector semantics into a unified platform, offering a comprehensive solution for managing diverse data types. The database provides exceptional horizontal scalability and elastic scaling capabilities, meeting enterprise-grade high availability requirements. Key features include comprehensive access interfaces supporting SQL, SDK, and API, with 'Table' and 'Vector' as first-class citizen data models. It boasts built-in data high availability, eliminating the need for external components, and supports fully automatic elastic data sharding for efficient resource allocation and business expansion. DingoDB excels in scalar-vector hybrid retrieval, combining traditional database index types with various vector index types, and supports distributed transaction processing. It offers real-time index optimization and 'cold-hot' tiered retrieval for massive datasets, minimizing memory consumption through disk-based vector search. The project is sponsored by DataCanvas and is Apache License Version 2.0.

https://github.com/dingodb/dingo

embedding-searchembedding-storehybrid-searchmysql-compatibilityreal-time-semantic-searchservingstructured-dataunified-sqlunstructured-datavector-databasevector-oceandistributed-storagesimilarity-search

Also in Vector Databases & Retrieval Infrastructure

pathwaycom/llm-app

Ready-to-run cloud templates for building real-time RAG, AI pipelines, and enterprise search applications that synchronize with various live data sources.

run-llama/llama_index

LlamaIndex is an open-source data framework for building LLM applications by connecting custom data sources to large language models, focusing on data ingestion, indexing, and retrieval augmented g...

milvus-io/milvus

Milvus is a high-performance, cloud-native vector database designed for scalable vector Approximate Nearest Neighbor (ANN) search, efficiently organizing and searching vast amounts of unstructured ...

VectifyAI/PageIndex

PageIndex is a vectorless, reasoning-based RAG system that builds hierarchical tree indexes from documents and uses LLMs to reason over them for context-aware retrieval.

qdrant/qdrant

Qdrant is an open-source, high-performance vector similarity search engine and vector database designed specifically for AI applications, enabling fast storage, search, and management of vectors wi...

Tencent/WeKnora

WeKnora is an open-source, LLM-powered knowledge framework for enterprise document understanding, semantic retrieval, and autonomous reasoning, featuring RAG, ReAct agents, and an auto-maintaining ...

topoteretes/cognee

Cognee is an open-source AI memory platform that provides AI agents with persistent long-term memory through a self-hosted knowledge graph, combining vector embeddings and graph reasoning.

RyanCodrai/turbovec

TurboVec is a Rust-based approximate nearest neighbor (ANN) vector index with Python bindings, built on Google Research's TurboQuant algorithm for efficient, memory-optimized vector similarity search.