Awesome Infra for AI › Vector Databases & Retrieval Infrastructure

oramasearch/orama

⭐ 10570 TypeScript repository created 2022-05-10

Orama is a versatile JavaScript search engine that offers a comprehensive suite of search functionalities, including full-text, vector, and hybrid search. It is engineered to operate efficiently across various environments, from web browsers to server-side applications and edge networks, with a minimal footprint (under 2KB). The engine supports advanced features such as search filters, geosearch, result pinning, facets, field boosting, typo tolerance, and exact matching. It also integrates BM25 ranking and multi-language stemming and tokenization for 30 languages. A notable aspect is its built-in support for vector and hybrid search, enabling the integration of text embeddings for semantic search. Orama can generate embeddings at insertion time via a plugin, facilitating the creation of powerful Retrieval-Augmented Generation (RAG) pipelines. While it functions as a general-purpose search engine, its primary utility in an AI/ML context lies in its vector search capabilities, which are crucial for retrieving context for LLMs and other AI agents. The project emphasizes ease of use, with straightforward installation and API for creating, inserting, and searching data, accommodating various data types including vectors.

https://github.com/oramasearch/orama

vector-databasevector-searchsearch-enginefull-text-searchhybrid-searchragjavascripttypescriptedgeserverbrowserembeddingreal-time

Also in Vector Databases & Retrieval Infrastructure

pathwaycom/llm-app

Ready-to-run cloud templates for building real-time RAG, AI pipelines, and enterprise search applications that synchronize with various live data sources.

run-llama/llama_index

LlamaIndex is an open-source data framework for building LLM applications by connecting custom data sources to large language models, focusing on data ingestion, indexing, and retrieval augmented g...

milvus-io/milvus

Milvus is a high-performance, cloud-native vector database designed for scalable vector Approximate Nearest Neighbor (ANN) search, efficiently organizing and searching vast amounts of unstructured ...

VectifyAI/PageIndex

PageIndex is a vectorless, reasoning-based RAG system that builds hierarchical tree indexes from documents and uses LLMs to reason over them for context-aware retrieval.

qdrant/qdrant

Qdrant is an open-source, high-performance vector similarity search engine and vector database designed specifically for AI applications, enabling fast storage, search, and management of vectors wi...

Tencent/WeKnora

WeKnora is an open-source, LLM-powered knowledge framework for enterprise document understanding, semantic retrieval, and autonomous reasoning, featuring RAG, ReAct agents, and an auto-maintaining ...

topoteretes/cognee

Cognee is an open-source AI memory platform that provides AI agents with persistent long-term memory through a self-hosted knowledge graph, combining vector embeddings and graph reasoning.

RyanCodrai/turbovec

TurboVec is a Rust-based approximate nearest neighbor (ANN) vector index with Python bindings, built on Google Research's TurboQuant algorithm for efficient, memory-optimized vector similarity search.