promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
Evidently is an open-source Python library designed for the evaluation, testing, and monitoring of both traditional Machine Learning operations and Large Language Model (LLM) powered systems. It functions across the entire lifecycle, from experimentation and development to production deployment. The framework supports a wide array of data types, including tabular and text data, and offers over 100 built-in metrics. These metrics cover essential aspects such as data drift detection, data quality checks (e.g., missing values, duplicates, correlations), and performance evaluations for predictive and generative tasks, including classification, regression, and RAG (Retrieval Augmented Generation) evaluations for LLMs. Users can also define custom metrics through its Python interface. Evidently's modular design allows for both one-off evaluations and continuous live monitoring. Key features include 'Reports' for summarizing data, ML, and LLM quality evaluations, which can be viewed interactively or exported, and 'Test Suites' for defining pass/fail conditions for regression testing and CI/CD integration. Additionally, it offers a monitoring dashboard for visualizing metrics and test results over time, which can be self-hosted or accessed via Evidently Cloud. The tool helps ensure the reliability and performance of AI/ML systems by providing robust monitoring and observability capabilities.
https://github.com/evidentlyai/evidently
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.
Cozeloop is an open-source platform offering full-lifecycle management for AI agents, encompassing development, debugging, evaluation, and monitoring.