Awesome Infra for AI › LLM Evaluation & Testing

open-compass/opencompass

⭐ 7492 Python added to this list on 2026-09-21 repository created 2023-06-15

OpenCompass is an open-source evaluation platform for large language models, developed by the OpenCompass / Shanghai AI Lab team. It provides a large collection of preconfigured datasets and benchmarks spanning language understanding, reasoning, coding, safety, and multimodal tasks, and lets users combine any model with any dataset/prompt template through Python configuration files (the AMOTIC config system). The framework supports both open-weight models (loaded locally via HuggingFace/vLLM-style backends) and closed API models (OpenAI, Anthropic, Gemini, LiteLLM AI Gateway, and others), with inference and evaluation stages that can run independently or concurrently, including task monitoring and heartbeat-based coordination for large-scale runs. Recent releases add multi-round inference for multi-turn instruction-following evaluation, integration with VLMEvalKit for native multimodal evaluation, a CascadeEvaluator that chains multiple evaluators for complex assessment pipelines, a GenericLLMEvaluator for LLM-as-judge scoring, and a MATHVerifyEvaluator for verifying mathematical reasoning outputs. Results feed the public OpenCompass Leaderboard (CompassRank) and CompassHub, where models can be submitted for community-wide ranking, and academic leaderboard results can be reproduced with documented instructions. Repeat-analysis tools detect repetitive or looping model outputs in evaluation runs, useful for diagnosing degenerate model behavior at scale. OpenCompass targets LLM researchers, model developers, and evaluation engineers who need a reproducible, extensible way to benchmark model quality across many dimensions rather than a single dataset, and to compare API-served and self-hosted models under identical evaluation conditions. It is a Python project distributed as a pip package and documented online.

https://github.com/open-compass/opencompass

llm-evaluationbenchmarkingleaderboardllm-as-judgemultimodal-evaluation

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.

coze-dev/coze-loop

Cozeloop is an open-source platform offering full-lifecycle management for AI agents, encompassing development, debugging, evaluation, and monitoring.