Awesome Infra for AI › LLM Evaluation & Testing

Raudaschl/rag-fusion

⭐ 958 Python repository created 2023-09-25

RAG-Fusion is a methodology designed to improve Retrieval Augmented Generation (RAG) by addressing the challenges of terminology mismatch between user queries and indexed text. It functions by first using a Large Language Model (LLM) to generate multiple diverse query variations from a single user query. These varied queries are then used to perform vector-based searches, casting a wider net across the document corpus. The results from these multiple searches are then combined and re-ranked using Reciprocal Rank Fusion (RRF), a technique that boosts documents appearing consistently across different query perspectives. This process aims to surface relevant material that a single query phrasing might miss. The project includes a robust quantitative evaluation harness, utilizing datasets like NFCorpus from the BEIR benchmark, to compare its effectiveness against baseline retrieval strategies, providing detailed empirical results, including confidence intervals and analysis of various fusion variants. RAG-Fusion is particularly effective when terminology mismatches occur, recall is prioritized over precision, and the downstream consumer (e.g., an LLM for synthesis or a UI presenting multiple candidates) can handle topically broad contexts. It's well-suited for applications such as academic research, legal discovery, and exploratory search, but less so for latency-critical or precision-dominated tasks. The project structure includes core pipeline components, evaluation scripts, and experimental write-ups, providing a comprehensive toolkit for implementing and analyzing advanced RAG techniques.

https://github.com/Raudaschl/rag-fusion

chromadbinformation-retrievalopenaipythonragrag-fusionreciprocal-rank-fusionretrieval-augmented-generationvector-search

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.