promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
Ragas is an open-source evaluation framework specifically designed for building and testing Large Language Model (LLM) applications. It helps developers to scientifically evaluate their LLM systems with objective metrics, moving beyond subjective assessments. The framework includes functionalities for intelligent test data generation, allowing users to automatically create comprehensive test datasets that cover a wide range of scenarios, even without pre-existing evaluation sets. Ragas integrates seamlessly with popular LLM frameworks like LangChain and various observability tools, enabling developers to incorporate evaluation directly into their development workflows. A core feature of Ragas is its ability to facilitate feedback loops, leveraging production data to continually refine and optimize LLM applications. It offers pre-built metrics for common evaluation tasks and supports the creation of custom aspect evaluators. This focus on systematic evaluation, test generation, and continuous improvement makes Ragas a crucial tool for ensuring the quality, reliability, and performance of LLM-powered applications in production environments.
https://github.com/vibrantlabsai/ragas
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.
Cozeloop is an open-source platform offering full-lifecycle management for AI agents, encompassing development, debugging, evaluation, and monitoring.