Awesome Infra for AI › LLM Evaluation & Testing

LLM Evaluation & Testing

56 projects

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

⭐ 25711 TypeScript

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

⭐ 20439 Python

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

⭐ 18632 Python

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

⭐ 15929 Python

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

⭐ 9431 Python

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

⭐ 7971 Jupyter Notebook

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

⭐ 7492 Python added 2026-09-21

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.

⭐ 5861 Python

coze-dev/coze-loop

Cozeloop is an open-source platform offering full-lifecycle management for AI agents, encompassing development, debugging, evaluation, and monitoring.

⭐ 5759 Go

Marker-Inc-Korea/AutoRAG

AutoRAG is an open-source framework designed to automate the evaluation and optimization of Retrieval-Augmented Generation (RAG) pipelines using an AutoML-style approach for specific datasets.

⭐ 5111 Python

langwatch/langwatch

LangWatch is a platform for end-to-end LLM evaluations, AI agent testing, and observability, offering tools for simulations, performance monitoring, prompt optimization, and an AI gateway for gover...

⭐ 4913 TypeScript

EvolvingLMMs-Lab/lmms-eval

LMMs-Eval is a unified, reproducible, and efficient evaluation toolkit for multimodal large language models (LMMs) across diverse tasks like text, image, video, and audio.

⭐ 4443 Python

truera/trulens

TruLens is an open-source framework for systematically evaluating and tracking LLM applications and AI agents, providing fine-grained instrumentation and comprehensive feedback functions.

⭐ 3589 Python

hegelai/prompttools

PromptTools provides open-source utilities for experimenting with, testing, and evaluating prompts, LLMs, and vector databases through code, notebooks, and a local playground.

⭐ 3057 Python

ianarawjo/ChainForge

ChainForge is an open-source visual programming environment designed for battle-testing, comparing, and evaluating prompts and LLM responses across different models and settings.

⭐ 3034 TypeScript

uptrain-ai/uptrain

UpTrain is an open-source platform providing evaluation and monitoring for Generative AI applications, offering preconfigured checks, root cause analysis, and production monitoring for LLMs.

⭐ 2367 Python

future-agi/future-agi

Future AGI is an open-source, end-to-end platform for evaluating, observing, simulating, and protecting LLM and AI agent applications, offering tracing, evals, guardrails, and a performant gateway.

⭐ 2109 Python

msoedov/agentic_security

Agentic Security is an open-source vulnerability scanner and AI red teaming kit designed to test Large Language Models (LLMs) and agent workflows against jailbreaks, fuzzing, and multimodal attacks.

⭐ 2017 Python

ShenSeanChen/waku-agent

Waku Agent is a local-first, personal AI assistant emphasizing a transparent architecture for its harness, loop, memory, and evaluation, designed for clarity and customizability.

⭐ 1904 Python added 2026-07-27

cyberark/FuzzyAI

FuzzyAI is an automated LLM fuzzing tool designed to identify and mitigate potential jailbreaks and security vulnerabilities in LLM APIs.

⭐ 1582 Jupyter Notebook

langwatch/better-agents

Better Agents is a CLI tool and set of standards for building, testing, and collaborating on AI agents, integrating with various frameworks and coding assistants for production readiness.

⭐ 1562 TypeScript

Jwuthri/Tracely-ai

Trace-native CI/CD for AI agents that grades production traces, clusters failures, freezes bad runs into hermetic replayable regression cases and blocks the pull request that would ship them again.

⭐ 1470 Python added 2026-08-24

plurai-ai/intellagent

IntellAgent evaluates and optimizes conversational AI agents through simulated, realistic synthetic interactions to uncover failure points and improve performance.

⭐ 1259 Python

superlinear-ai/raglite

RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) that provides configurable components for LLMs, vector databases, and rerankers, with optimized strategies for chunking, retriev...

⭐ 1210 Python

cvs-health/uqlm

UQLM is a Python library for detecting and mitigating hallucination in Large Language Model (LLM) outputs using uncertainty quantification techniques.

⭐ 1207 Python

Ricky-7-Yan/intelligent-audit-system

AuditPilot is an enterprise AI agent workbench designed for auditable, evidence-grounded workflows, featuring governed tools, evaluation harnesses, human review, and remediation delivery for audit ...

⭐ 1172 Python added 2026-07-27

prometheus-eval/prometheus-eval

Prometheus-Eval is a framework and a collection of open-source LLM judges designed for evaluating the quality of LLM responses in generation tasks, supporting both absolute grading and pairwise ran...

⭐ 1118 Python

JudgmentLabs/judgeval

Judgeval is an open-source Python SDK enabling continuous improvement for AI agents through OpenTelemetry-based tracing, agent-judge evaluations, and online monitoring of LLM-powered applications.

⭐ 1063 Python

juanjuandog/FinSight-AI

FinSight AI is an open-source AI equity research agent that develops evidence-grounded reports with resilient workflow orchestration, RAG evaluation, and comprehensive backend infrastructure.

⭐ 1063 Java

ZhangJinHaHaHa/AgentLens

AgentLens is a decentralized marketplace and infrastructure for AI Agents, providing verifiable proof of capabilities, security, and track record using on-chain auditing, TEE attestation, and ZK pr...

⭐ 1028 TypeScript

Raudaschl/rag-fusion

RaG-Fusion enhances RAG via multi-query generation and Reciprocal Rank Fusion to improve retrieval, especially for term mismatches, including an evaluation harness with NFCorpus/BEIR.

⭐ 958 Python

TIGER-AI-Lab/ClawBench

ClawBench is an open-source benchmark for evaluating AI browser agents on a diverse set of everyday online tasks across live websites, measuring end-to-end task success.

⭐ 942 Python

darkrishabh/agent-skills-eval

A test runner for Agent Skills that evaluates the effectiveness of AI agent skills by comparing model performance with and without a skill, using a judge model for grading.

⭐ 797 TypeScript

HowieHwong/TrustLLM

TrustLLM is an open research toolkit and CLI for benchmarking the trustworthiness of large language models across six dimensions: truthfulness, safety, fairness, robustness, privacy and ethics.

⭐ 633 Python

onyx-dot-app/EnterpriseRAG-Bench

EnterpriseRAG-Bench offers a benchmark dataset and evaluation framework for RAG systems using realistic company internal documents and a diverse set of questions.

⭐ 576

PacificAI/langtest

LangTest is an open-source library for testing and evaluating Large Language Models and NLP models for various quality aspects like robustness, bias, fairness, and accuracy.

⭐ 561 Python

xinxuxin/keystone-bench

Evaluation benchmark that tests whether chat assistants change clinical advice correctly when a decisive fact is added, removed, or contradicted, and whether physician-written grading rubrics still apply after the edit.

⭐ 530 Python added 2026-09-21

agencyenterprise/PromptInject

PromptInject is a framework for quantitatively analyzing the robustness of LLMs against adversarial prompt attacks through modular prompt assembly and evaluation.

⭐ 526 Python added 2026-06-29

relari-ai/continuous-eval

continuous-eval is an open-source framework for data-driven, modular evaluation of LLM-powered applications, offering a comprehensive metric library and probabilistic evaluation capabilities.

⭐ 518 Python

vectara/open-rag-eval

An open-source Python toolkit for evaluating Retrieval-Augmented Generation (RAG) pipelines, offering flexible metrics and connectors without requiring golden answers.

⭐ 412 Python

rhesis-ai/rhesis

Rhesis is an open-source collaborative testing platform for LLM and agentic applications, providing AI-powered test generation, conversation simulation, adversarial testing, and comprehensive evalu...

⭐ 396 Python

ukanwat/aaabench

An open-ended benchmark harness that hands a coding agent a live Unreal Engine 5 editor over MCP and measures whether it can build a whole open-world game unaided.

⭐ 387 Shell added 2026-08-17

alphadl/AdaRubrics

AdaRubric offers task-adaptive rubrics and dense reward signals for evaluating LLM agent trajectories, enhancing evaluation reliability and reward learning.

⭐ 364 Python

OpenBMB/UltraEval-Audio

UltraEval-Audio is a unified open-source framework for comprehensive and reproducible evaluation of audio foundation models across speech understanding and speech generation tasks.

⭐ 326 Python added 2026-07-27

aaron-for-value/VeriRun

Distributed runtime for reproducible, isolated evaluation of AI-generated code and agent benchmarks, with sandboxed execution, provenance tracking, and a reward path for post-training loops.

⭐ 323 Python added 2026-09-07

athina-ai/athina-evals

A Python SDK offering 50+ preset and custom evaluations for LLM-generated responses, integrating with the Athina IDE for experimentation and dataset comparison.

⭐ 303 Python

iMeanAI/WebCanvas

WebCanvas is an open-source framework for building, training, and evaluating LLM-based web agents in dynamic, real-time online environments.

⭐ 281 Python

cvs-health/langfair

LangFair is a Python library for conducting use-case level bias and fairness assessments of large language models (LLMs) by allowing users to bring their own prompts for evaluation.

⭐ 262 Python

JinjieNi/MixEval

MixEval is a dynamic, ground-truth-based benchmark and evaluation suite for large language and multimodal models.

⭐ 254 Python

LeoYeAI/myclaw-bench

MyClaw Bench provides a comprehensive benchmark for evaluating AI agents on OpenClaw, featuring 45 tasks across four difficulty tiers with a focus on real-world outcomes and complex reasoning.

⭐ 223 Python

LLAMATOR-Core/llamator

LLAMATOR is a Python framework for red teaming and security testing of chatbots, Generative AI systems, LLMs, RAGs, Agents, and Vision Language Models (VLMs) against various attacks and vulnerabili...

⭐ 223 Python

MigoXLab/LMeterX

LMeterX is a professional, full-lifecycle platform for load testing and performance benchmarking of large language models and other AI inference services, offering real-time monitoring and AI-power...

⭐ 211 Python added 2026-08-03

jsdhwfmax/EvalForge

Portable evaluation-evidence and policy-gate tool that turns Ragas, promptfoo and DeepEval results into versioned, CI-enforceable pass/fail reports for RAG systems.

⭐ 210 Python added 2026-09-14

GiovanniPasq/chunky

Chunky is an open-source toolkit for preparing documents for Retrieval Augmented Generation (RAG) pipelines, offering PDF-to-Markdown conversion, cleaning, chunk inspection, and chunking strategy c...

⭐ 185 Python

hermes-labs-ai/lintlang

Local, deterministic static linter that scans agent instructions, MCP tool schemas and prompt files for ambiguity, conflicts and missing bounds before an agent runs.

⭐ 135 Python added 2026-09-28

huangyiminghappy/ai-eval-platform

An open-source, self-hosted AI evaluation platform offering a web UI for assessing RAG, AI Agents, and multi-turn conversations through dataset management, scenario presets, metrics, LLM-as-a-Judge...

⭐ 12 Python added 2026-08-03