promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
LMeterX is a comprehensive platform designed for performance testing and benchmarking of large language models (LLMs) and general business HTTP interfaces. It supports a wide range of inference frameworks such as vLLM, LiteLLM, and TensorRT-LLM, as well as cloud AI services from Azure OpenAI, AWS Bedrock, and Google Vertex AI. The platform provides an intuitive web-based user interface to create, manage, and monitor test tasks in real-time, offering detailed performance analysis reports crucial for model deployment and optimization. Key features include broad framework compatibility, support for full modality and scenarios (text, multimodal, streaming), and hybrid protocol testing for both standard Chat APIs and business HTTP interfaces. LMeterX offers multi-mode and high-concurrency load testing, allowing for fixed or stepped concurrency strategies to identify performance inflection points and system capacity limits. It includes built-in dual-mode datasets for ease of use, an automated warm-up mechanism to eliminate cold-start effects, and multi-dimensional visualization of indicators like TTFT, RPS, TPS, and throughput. The platform also provides engine resource monitoring for CPU, memory, and network bandwidth, helping to pinpoint resource bottlenecks. A standout feature is its AI-driven data insights, which generate analysis reports with multi-model comparisons to guide optimization efforts. It integrates with AI Agents for natural language-driven task generation, supports web parsing for automatic API discovery, and offers enterprise-grade security and scaling with distributed elastic deployment and LDAP/AD integration. LMeterX utilizes a microservices architecture based on FastAPI, Locust, React, and MySQL, and is easily deployable via Docker/Kubernetes.
https://github.com/MigoXLab/LMeterX
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.