promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
Keystone is a benchmark and Python toolkit for evaluating clinical chat assistants on decision-evidence sensitivity rather than static answer quality. Starting from HealthBench source conversations, each is annotated with a decision frame (what is being decided, the action the evidence supports, and the facts it depends on) and paired with twins produced by changing exactly one fact: removing a decisive detail, adding a contradiction, changing a demographic attribute, or inserting an irrelevant distractor, plus a paraphrase-only control that changes nothing. Every twin is labeled with the resulting evidence state and the actions a clinician would accept or forbid, and a separate judge model maps each assistant reply to that annotation so the metric is about whether the action changed appropriately with the evidence, not just whether the reply matches a rubric written for the original question. The release ships thousands of twins across eight edit families, evidence-state annotations reviewed by an independent vendor, reference results for five assistants, and reproducibility tooling: a CLI that rebuilds the dataset locally from the public HealthBench source and label files, with published file hashes to verify the rebuild. It also ships validity checks, including a grader-accuracy check, a degenerate-strategy check confirming no fixed policy can win, and an audit for whether an edit leaves a detectable fingerprint separate from the evidence itself. It is aimed at teams building or evaluating clinical or health-adjacent chat assistants, and at benchmark designers interested in the underlying decision-evidence methodology for testing rubric applicability rather than only reply quality.
https://github.com/xinxuxin/keystone-bench
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.