Awesome Infra for AI › LLM Evaluation & Testing

xinxuxin/keystone-bench

⭐ 530 Python added to this list on 2026-09-21 repository created 2026-09-07

Keystone is a benchmark and Python toolkit for evaluating clinical chat assistants on decision-evidence sensitivity rather than static answer quality. Starting from HealthBench source conversations, each is annotated with a decision frame (what is being decided, the action the evidence supports, and the facts it depends on) and paired with twins produced by changing exactly one fact: removing a decisive detail, adding a contradiction, changing a demographic attribute, or inserting an irrelevant distractor, plus a paraphrase-only control that changes nothing. Every twin is labeled with the resulting evidence state and the actions a clinician would accept or forbid, and a separate judge model maps each assistant reply to that annotation so the metric is about whether the action changed appropriately with the evidence, not just whether the reply matches a rubric written for the original question. The release ships thousands of twins across eight edit families, evidence-state annotations reviewed by an independent vendor, reference results for five assistants, and reproducibility tooling: a CLI that rebuilds the dataset locally from the public HealthBench source and label files, with published file hashes to verify the rebuild. It also ships validity checks, including a grader-accuracy check, a degenerate-strategy check confirming no fixed policy can win, and an audit for whether an edit leaves a detectable fingerprint separate from the evidence itself. It is aimed at teams building or evaluating clinical or health-adjacent chat assistants, and at benchmark designers interested in the underlying decision-evidence methodology for testing rubric applicability rather than only reply quality.

https://github.com/xinxuxin/keystone-bench

llm-evaluationhealthcare-aibenchmarkevaluation-methodologyllm-as-judge

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.