promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.
Cozeloop is an open-source platform offering full-lifecycle management for AI agents, encompassing development, debugging, evaluation, and monitoring.
AutoRAG is an open-source framework designed to automate the evaluation and optimization of Retrieval-Augmented Generation (RAG) pipelines using an AutoML-style approach for specific datasets.
LangWatch is a platform for end-to-end LLM evaluations, AI agent testing, and observability, offering tools for simulations, performance monitoring, prompt optimization, and an AI gateway for gover...
LMMs-Eval is a unified, reproducible, and efficient evaluation toolkit for multimodal large language models (LMMs) across diverse tasks like text, image, video, and audio.
TruLens is an open-source framework for systematically evaluating and tracking LLM applications and AI agents, providing fine-grained instrumentation and comprehensive feedback functions.
PromptTools provides open-source utilities for experimenting with, testing, and evaluating prompts, LLMs, and vector databases through code, notebooks, and a local playground.
ChainForge is an open-source visual programming environment designed for battle-testing, comparing, and evaluating prompts and LLM responses across different models and settings.
UpTrain is an open-source platform providing evaluation and monitoring for Generative AI applications, offering preconfigured checks, root cause analysis, and production monitoring for LLMs.
Future AGI is an open-source, end-to-end platform for evaluating, observing, simulating, and protecting LLM and AI agent applications, offering tracing, evals, guardrails, and a performant gateway.
Agentic Security is an open-source vulnerability scanner and AI red teaming kit designed to test Large Language Models (LLMs) and agent workflows against jailbreaks, fuzzing, and multimodal attacks.
Waku Agent is a local-first, personal AI assistant emphasizing a transparent architecture for its harness, loop, memory, and evaluation, designed for clarity and customizability.
FuzzyAI is an automated LLM fuzzing tool designed to identify and mitigate potential jailbreaks and security vulnerabilities in LLM APIs.
Better Agents is a CLI tool and set of standards for building, testing, and collaborating on AI agents, integrating with various frameworks and coding assistants for production readiness.
Trace-native CI/CD for AI agents that grades production traces, clusters failures, freezes bad runs into hermetic replayable regression cases and blocks the pull request that would ship them again.
IntellAgent evaluates and optimizes conversational AI agents through simulated, realistic synthetic interactions to uncover failure points and improve performance.
RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) that provides configurable components for LLMs, vector databases, and rerankers, with optimized strategies for chunking, retriev...
UQLM is a Python library for detecting and mitigating hallucination in Large Language Model (LLM) outputs using uncertainty quantification techniques.
AuditPilot is an enterprise AI agent workbench designed for auditable, evidence-grounded workflows, featuring governed tools, evaluation harnesses, human review, and remediation delivery for audit ...
Prometheus-Eval is a framework and a collection of open-source LLM judges designed for evaluating the quality of LLM responses in generation tasks, supporting both absolute grading and pairwise ran...
Judgeval is an open-source Python SDK enabling continuous improvement for AI agents through OpenTelemetry-based tracing, agent-judge evaluations, and online monitoring of LLM-powered applications.
FinSight AI is an open-source AI equity research agent that develops evidence-grounded reports with resilient workflow orchestration, RAG evaluation, and comprehensive backend infrastructure.
AgentLens is a decentralized marketplace and infrastructure for AI Agents, providing verifiable proof of capabilities, security, and track record using on-chain auditing, TEE attestation, and ZK pr...
RaG-Fusion enhances RAG via multi-query generation and Reciprocal Rank Fusion to improve retrieval, especially for term mismatches, including an evaluation harness with NFCorpus/BEIR.
ClawBench is an open-source benchmark for evaluating AI browser agents on a diverse set of everyday online tasks across live websites, measuring end-to-end task success.
A test runner for Agent Skills that evaluates the effectiveness of AI agent skills by comparing model performance with and without a skill, using a judge model for grading.
TrustLLM is an open research toolkit and CLI for benchmarking the trustworthiness of large language models across six dimensions: truthfulness, safety, fairness, robustness, privacy and ethics.
EnterpriseRAG-Bench offers a benchmark dataset and evaluation framework for RAG systems using realistic company internal documents and a diverse set of questions.
LangTest is an open-source library for testing and evaluating Large Language Models and NLP models for various quality aspects like robustness, bias, fairness, and accuracy.
Evaluation benchmark that tests whether chat assistants change clinical advice correctly when a decisive fact is added, removed, or contradicted, and whether physician-written grading rubrics still apply after the edit.
PromptInject is a framework for quantitatively analyzing the robustness of LLMs against adversarial prompt attacks through modular prompt assembly and evaluation.
continuous-eval is an open-source framework for data-driven, modular evaluation of LLM-powered applications, offering a comprehensive metric library and probabilistic evaluation capabilities.
An open-source Python toolkit for evaluating Retrieval-Augmented Generation (RAG) pipelines, offering flexible metrics and connectors without requiring golden answers.
Rhesis is an open-source collaborative testing platform for LLM and agentic applications, providing AI-powered test generation, conversation simulation, adversarial testing, and comprehensive evalu...
An open-ended benchmark harness that hands a coding agent a live Unreal Engine 5 editor over MCP and measures whether it can build a whole open-world game unaided.
AdaRubric offers task-adaptive rubrics and dense reward signals for evaluating LLM agent trajectories, enhancing evaluation reliability and reward learning.
UltraEval-Audio is a unified open-source framework for comprehensive and reproducible evaluation of audio foundation models across speech understanding and speech generation tasks.
Distributed runtime for reproducible, isolated evaluation of AI-generated code and agent benchmarks, with sandboxed execution, provenance tracking, and a reward path for post-training loops.
A Python SDK offering 50+ preset and custom evaluations for LLM-generated responses, integrating with the Athina IDE for experimentation and dataset comparison.
WebCanvas is an open-source framework for building, training, and evaluating LLM-based web agents in dynamic, real-time online environments.
LangFair is a Python library for conducting use-case level bias and fairness assessments of large language models (LLMs) by allowing users to bring their own prompts for evaluation.
MixEval is a dynamic, ground-truth-based benchmark and evaluation suite for large language and multimodal models.
MyClaw Bench provides a comprehensive benchmark for evaluating AI agents on OpenClaw, featuring 45 tasks across four difficulty tiers with a focus on real-world outcomes and complex reasoning.
LLAMATOR is a Python framework for red teaming and security testing of chatbots, Generative AI systems, LLMs, RAGs, Agents, and Vision Language Models (VLMs) against various attacks and vulnerabili...
LMeterX is a professional, full-lifecycle platform for load testing and performance benchmarking of large language models and other AI inference services, offering real-time monitoring and AI-power...
Portable evaluation-evidence and policy-gate tool that turns Ragas, promptfoo and DeepEval results into versioned, CI-enforceable pass/fail reports for RAG systems.
Chunky is an open-source toolkit for preparing documents for Retrieval Augmented Generation (RAG) pipelines, offering PDF-to-Markdown conversion, cleaning, chunk inspection, and chunking strategy c...
Local, deterministic static linter that scans agent instructions, MCP tool schemas and prompt files for ambiguity, conflicts and missing bounds before an agent runs.
An open-source, self-hosted AI evaluation platform offering a web UI for assessing RAG, AI Agents, and multi-turn conversations through dataset management, scenario presets, metrics, LLM-as-a-Judge...