pathwaycom/llm-app
Ready-to-run cloud templates for building real-time RAG, AI pipelines, and enterprise search applications that synchronize with various live data sources.
Awesome Infra for AI › Vector Databases & Retrieval Infrastructure
haiku.rag is a Python library and command line tool that builds retrieval-augmented question answering over a local document collection without requiring a separate database server. Documents are parsed by Docling and stored as full structured documents, which lets the retriever expand context along headings and sections rather than returning isolated chunks. Search is hybrid: dense vector similarity and full-text matching are combined with reciprocal rank fusion, and results can be reranked by a local cross-encoder or by a hosted reranking service. Multimodal ingestion captures embedded figures, and multimodal embedders place picture vectors in the same space as text, so a text query can retrieve a figure and an image can be used as the query. Answers carry citations down to page numbers and section headings, and an optional citation policy requires every answer to declare what grounds it, including declaring that nothing does. An analysis capability executes sandboxed Python for aggregation and computation across several documents, and an evidence compaction capability replaces earlier search results in a conversation with the evidence they cited so long sessions stop resending everything they retrieved. Storage is embedded LanceDB, so a deployment is a directory rather than a service. Embedding and generation providers are pluggable across Ollama, OpenAI, VoyageAI, Cohere, LM Studio and vLLM, with question answering handled through Pydantic AI. Interfaces include a chat terminal application, a web application, and an MCP server that exposes the collection to agent hosts. The project fits self-hosted and privacy-sensitive deployments where the retrieval layer must run entirely on local infrastructure.
https://github.com/ggozad/haiku.rag
Ready-to-run cloud templates for building real-time RAG, AI pipelines, and enterprise search applications that synchronize with various live data sources.
LlamaIndex is an open-source data framework for building LLM applications by connecting custom data sources to large language models, focusing on data ingestion, indexing, and retrieval augmented g...
Milvus is a high-performance, cloud-native vector database designed for scalable vector Approximate Nearest Neighbor (ANN) search, efficiently organizing and searching vast amounts of unstructured ...
PageIndex is a vectorless, reasoning-based RAG system that builds hierarchical tree indexes from documents and uses LLMs to reason over them for context-aware retrieval.
Qdrant is an open-source, high-performance vector similarity search engine and vector database designed specifically for AI applications, enabling fast storage, search, and management of vectors wi...
WeKnora is an open-source, LLM-powered knowledge framework for enterprise document understanding, semantic retrieval, and autonomous reasoning, featuring RAG, ReAct agents, and an auto-maintaining ...
Cognee is an open-source AI memory platform that provides AI agents with persistent long-term memory through a self-hosted knowledge graph, combining vector embeddings and graph reasoning.
TurboVec is a Rust-based approximate nearest neighbor (ANN) vector index with Python bindings, built on Google Research's TurboQuant algorithm for efficient, memory-optimized vector similarity search.