promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
Garak (Generative AI Red-teaming & Assessment Kit) is a cybersecurity tool specifically engineered to identify and categorize vulnerabilities within Large Language Models (LLMs) and conversational AI systems. Its primary function is to act as an adversarial testing framework, probing LLMs to expose undesirable behaviors or security flaws. These include, but are not limited to, hallucination, inadvertent data leakage, prompt injection attacks, generation of misinformation, toxic content production, and various jailbreaking techniques. Inspired by traditional cybersecurity tools like nmap and Metasploit, Garak provides a structured approach to assessing LLM resilience. It employs a combination of static, dynamic, and adaptive probes to explore the model's failure modes. The tool natively supports a wide range of LLM providers and platforms, including Hugging Face models (both local and API-based), OpenAI API, AWS Bedrock, Replicate, and generally any LLM accessible via a REST API. It also handles GGUF models (e.g., llama.cpp) and integrations via litellm. Users can specify target models and select from a variety of built-in probes or develop custom ones to simulate different attack vectors. After execution, Garak generates detailed logs and provides an evaluation of the LLM's susceptibility to the tested vulnerabilities, indicating which probes led to desired outcomes or failures. This capability is crucial for developers and organizations deploying LLMs to ensure their safety, reliability, and compliance, offering a dedicated solution for AI safety, guardrails, and robust evaluation.
https://github.com/NVIDIA/garak
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.
Cozeloop is an open-source platform offering full-lifecycle management for AI agents, encompassing development, debugging, evaluation, and monitoring.