promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
OpenCompass is an open-source evaluation platform for large language models, developed by the OpenCompass / Shanghai AI Lab team. It provides a large collection of preconfigured datasets and benchmarks spanning language understanding, reasoning, coding, safety, and multimodal tasks, and lets users combine any model with any dataset/prompt template through Python configuration files (the AMOTIC config system). The framework supports both open-weight models (loaded locally via HuggingFace/vLLM-style backends) and closed API models (OpenAI, Anthropic, Gemini, LiteLLM AI Gateway, and others), with inference and evaluation stages that can run independently or concurrently, including task monitoring and heartbeat-based coordination for large-scale runs. Recent releases add multi-round inference for multi-turn instruction-following evaluation, integration with VLMEvalKit for native multimodal evaluation, a CascadeEvaluator that chains multiple evaluators for complex assessment pipelines, a GenericLLMEvaluator for LLM-as-judge scoring, and a MATHVerifyEvaluator for verifying mathematical reasoning outputs. Results feed the public OpenCompass Leaderboard (CompassRank) and CompassHub, where models can be submitted for community-wide ranking, and academic leaderboard results can be reproduced with documented instructions. Repeat-analysis tools detect repetitive or looping model outputs in evaluation runs, useful for diagnosing degenerate model behavior at scale. OpenCompass targets LLM researchers, model developers, and evaluation engineers who need a reproducible, extensible way to benchmark model quality across many dimensions rather than a single dataset, and to compare API-served and self-hosted models under identical evaluation conditions. It is a Python project distributed as a pip package and documented online.
https://github.com/open-compass/opencompass
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.
Cozeloop is an open-source platform offering full-lifecycle management for AI agents, encompassing development, debugging, evaluation, and monitoring.