Awesome Infra for AI › LLM Evaluation & Testing

confident-ai/deepeval

⭐ 18632 Python repository created 2023-08-10

DeepEval is a robust, open-source framework designed for the unit testing and evaluation of Large Language Model (LLM) systems. Operating much like Pytest but specialized for LLM applications, it incorporates the latest research to provide a comprehensive suite of evaluation metrics. These metrics, such as G-Eval, Task Completion, Answer Relevancy, and Faithfulness, leverage LLM-as-a-judge paradigms and other NLP models that can run locally. The framework is versatile, suitable for diverse applications including AI agents, Retrieval-Augmented Generation (RAG) pipelines, and chatbots, whether built with tools like LangChain or OpenAI. Key capabilities include helping developers determine optimal models, prompts, and architectures to enhance AI quality, prevent prompt drifting, and confidently transition between different LLMs (e.g., OpenAI to Claude). It offers a wide array of ready-to-use metrics, categorized for custom, agentic, RAG, and multi-turn use cases. DeepEval supports evaluation based on custom criteria, agent goal accomplishment, tool correctness, factual alignment, and conversational consistency, empowering developers to rigorously test and improve their LLM applications.

https://github.com/confident-ai/deepeval

LLM evaluationAI evaluationRAG evaluationagent evaluationchatbot evaluationLLM unit testingprompt engineeringG-EvalPythonopen-source

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.

coze-dev/coze-loop

Cozeloop is an open-source platform offering full-lifecycle management for AI agents, encompassing development, debugging, evaluation, and monitoring.