promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
ChainForge is a visual, data-flow based prompt engineering environment that facilitates the analysis and evaluation of Large Language Model (LLM) responses. Its primary purpose is to enable rapid comparison of prompts, models, and response quality, moving beyond ad-hoc interactions with individual LLMs. Users can query multiple LLMs simultaneously to quickly test prompt ideas and variations. It allows for the comparison of response quality across various prompt permutations, different models (e.g., OpenAI, Anthropic, Google Gemini, Ollama), and diverse model settings to identify the optimal combination for specific use cases. The platform supports setting up custom evaluation metrics, often via Python scripts, and immediately visualizes results across prompts, parameters, and models. Key features include combinatorial prompting, where ChainForge takes the cross product of inputs to prompt templates, generating every possible combination for robust testing. It supports ground-truth evaluations using tabular data nodes, allowing comparison of LLM answers against expected outcomes. Users can interactively inspect and compare responses across models and prompt variables. Additionally, ChainForge provides AI-powered features to streamline evaluation, such as generating synthetic tables and input examples. The tool is built with ReactFlow and Flask, offering both local installation and a web-based version for convenient access. It is an essential tool for developers and researchers focused on refining and validating their LLM applications by systematically comparing model behavior and response quality.
https://github.com/ianarawjo/ChainForge
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.