Awesome Infra for AI › LLM Evaluation & Testing

HowieHwong/TrustLLM

⭐ 633 Python repository created 2023-12-23

TrustLLM is an evaluation toolkit built around the ICML 2024 TrustLLM benchmark, packaged as a Python library (`trustllm`) installable directly from GitHub with optional extras for local-model inference and scoring dependencies. It provides a three-stage workflow exposed both as CLI commands and as a Python API: downloading the benchmark dataset and task definitions (`trustllm download`, `trustllm tasks`), generating model responses (`trustllm generate`) against either an OpenAI-compatible API endpoint or local Hugging Face weights (with device selection including CPU, CUDA and MPS), and evaluating the generated responses into scores (`trustllm evaluate`). Evaluation covers six dimensions - truthfulness (misinformation, hallucination, sycophancy), safety (jailbreaks, misuse), fairness (stereotypes, preferences), robustness (adversarial and out-of-domain inputs), privacy (awareness and leakage), and ethics (moral judgment) - using a mix of rule-based scoring, a downloaded classifier, embeddings, and an LLM-as-judge call configurable via an environment variable. Generation supports bounded retries, per-sample checkpoints, and a resumable `--resume` flag, and produces HTML reports plus saved dataset hashes and settings for reproducibility. A `--limit` flag allows quick smoke tests before a full run. The project includes a documented guide for orchestrating evaluations from an AI agent via CLI or Python, with JSON-structured results, and a machine-readable documentation index. The original 0.3 generation engine is preserved as a legacy module for users needing exact reproduction of the original paper's results. It is MIT licensed, with dataset terms following the original upstream sources, and documentation translated into several languages.

https://github.com/HowieHwong/TrustLLM

llm-evaluationbenchmarktrustworthinesssafety-testingfairnessrobustnessprivacyresearch-toolkitllm-as-judge

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.