promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
EnterpriseRAG-Bench provides a comprehensive dataset comprising over 500,000 synthetic company internal documents and 500 questions designed for benchmarking Retrieval Augmented Generation (RAG) systems. The dataset simulates a realistic enterprise environment with data pulled from various sources like Slack, Gmail, Linear, Google Drive, Hubspot, Fireflies, GitHub, Jira, and Confluence, ensuring broad coverage of business activities. The questions are categorized into ten types, ranging from basic semantic queries to complex intra-document reasoning and questions involving conflicting information, completeness, and 'information not found' scenarios. An additional set of metadata-dependent questions is also available for evaluating metadata-aware RAG systems. The project emphasizes key considerations for realistic data generation: cross-document coherence, realistic volume distribution, inclusion of noise, internal terminology, and generality across enterprise settings. It addresses a gap in existing RAG/IR datasets, which typically focus on publicly accessible data, by providing a specialized resource for internal knowledge bases. The project also features a leaderboard for tracking and comparing the performance of different RAG systems, encouraging reproducible submissions. This framework serves as a valuable resource for teams looking to benchmark their RAG solutions or fine-tune AI agents on enterprise-specific data.
https://github.com/onyx-dot-app/EnterpriseRAG-Bench
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.