promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
LLAMATOR is an open-source Python framework designed for the crucial task of red teaming and security testing of Large Language Models (LLMs) and other Generative AI systems. It provides a robust set of tools and methodologies to identify and mitigate vulnerabilities such as prompt injection, jailbreaks, system prompt leakage, misinformation generation, and unbounded consumption. The framework supports testing various AI clients, including those built with LangChain, OpenAI-like APIs, and custom classes for platforms like Telegram, WhatsApp, and Selenium. LLAMATOR offers a large selection of pre-defined attacks and allows users to define custom attacks and datasets, enhancing its adaptability to specific security requirements. It also provides functionalities for configuring chat clients, recording attack requests and responses, and generating detailed test reports in formats like Excel, CSV, and DOCX. By aligning with OWASP Top 10 for LLM Application Security Risks, LLAMATOR helps developers and security professionals evaluate the robustness and safety of their AI applications against malicious exploitation and unexpected behavior, ensuring more secure and reliable deployments.
https://github.com/LLAMATOR-Core/llamator
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.