Awesome Infra for AI › LLM Evaluation & Testing

cvs-health/uqlm

⭐ 1207 Python repository created 2025-04-17

UQLM (Uncertainty Quantification for Language Models) is a Python package designed to detect and mitigate hallucinations in LLM responses. It offers a suite of response-level scorers that quantify the uncertainty of LLM outputs, providing a confidence score (0 to 1) where higher scores indicate a lower likelihood of errors or hallucinations. The library categorizes its scorers into several types: Black-Box Scorers (consistency-based), which assess uncertainty by comparing multiple responses generated from the same prompt and are compatible with any LLM; White-Box Scorers (token-probability-based), which leverage token probabilities for faster and cheaper single-generation scoring but require access to the LLM's internal probabilities; LLM-as-a-Judge Scorers, which use other LLMs to evaluate responses; Ensemble Scorers, which combine various scoring methods; and Long-Text Scorers, for claim-level analysis. UQLM integrates with LangChain Chat Models, making it versatile for various LLM deployments. It also provides functionality to select uncertainty-minimized responses, effectively mitigating potential hallucinations. The project emphasizes practical application with code examples and detailed explanations of each scorer type, citing relevant research papers.

https://github.com/cvs-health/uqlm

LLMhallucination detectionuncertainty quantificationAI safetyLLM evaluationconfidence scoringLangChain

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.