promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
AuditPilot is an enterprise AI agent workbench specifically tailored for audit delivery scenarios. It focuses on integrating evidence retrieval, tool execution, harness-based evaluation, human review, and remediation into a single traceable workflow. Unlike general-purpose agent projects, AuditPilot addresses the 'harder' enterprise requirements of auditing, including the need for strong evidence, control, risk management, review mechanisms, reporting, and a closed-loop remediation process, emphasizing traceability and interpretability. The platform offers a comprehensive set of modules: an Audit Workspace for managing audit projects, an Agent Runtime with a bounded Plan/Execute/Reflect cycle and human review gateways, Agentic RAG for evidence grounding and quality gates, Skills/MCP-style tools with schema and access control, various Memory types, and a crucial Evaluation Harness for multi-layered assessment of task outcomes, execution traces, tool calls, evidence, safety, and robustness. It also incorporates an Evidence Graph to link tasks, steps, tools, and artifacts, ensuring source coverage and dependency integrity. Governed Improvement ensures that failures are treated as experience candidates, reusable only after regression testing and human approval. Observability features provide privacy-friendly task traces, standardized fields, and comprehensive summaries. AuditPilot's architecture includes a hybrid intent router, multiple specialized agents (Planner, Evidence, Control, Risk, Compliance, Remediation), an agentic RAG system, an evidence graph, and skills. Critical design boundaries ensure that LLMs are confined to understanding and summarization, while auditable aspects like evidence gaps, quality gates, permissions, and risk signals remain transparent and controllable. High-risk or low-confidence outputs are systematically routed for human review and evidence completion. The project supports Python 3.10+ and allows for execution without LLM API keys via a deterministic fallback mode for testing workflow, RAG, evaluation, and UI components.
https://github.com/Ricky-7-Yan/intelligent-audit-system
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.