Awesome Infra for AI › LLM Evaluation & Testing

Ricky-7-Yan/intelligent-audit-system

⭐ 1172 Python added to this list on 2026-07-27 repository created 2025-10-28

AuditPilot is an enterprise AI agent workbench specifically tailored for audit delivery scenarios. It focuses on integrating evidence retrieval, tool execution, harness-based evaluation, human review, and remediation into a single traceable workflow. Unlike general-purpose agent projects, AuditPilot addresses the 'harder' enterprise requirements of auditing, including the need for strong evidence, control, risk management, review mechanisms, reporting, and a closed-loop remediation process, emphasizing traceability and interpretability. The platform offers a comprehensive set of modules: an Audit Workspace for managing audit projects, an Agent Runtime with a bounded Plan/Execute/Reflect cycle and human review gateways, Agentic RAG for evidence grounding and quality gates, Skills/MCP-style tools with schema and access control, various Memory types, and a crucial Evaluation Harness for multi-layered assessment of task outcomes, execution traces, tool calls, evidence, safety, and robustness. It also incorporates an Evidence Graph to link tasks, steps, tools, and artifacts, ensuring source coverage and dependency integrity. Governed Improvement ensures that failures are treated as experience candidates, reusable only after regression testing and human approval. Observability features provide privacy-friendly task traces, standardized fields, and comprehensive summaries. AuditPilot's architecture includes a hybrid intent router, multiple specialized agents (Planner, Evidence, Control, Risk, Compliance, Remediation), an agentic RAG system, an evidence graph, and skills. Critical design boundaries ensure that LLMs are confined to understanding and summarization, while auditable aspects like evidence gaps, quality gates, permissions, and risk signals remain transparent and controllable. High-risk or low-confidence outputs are systematically routed for human review and evidence completion. The project supports Python 3.10+ and allows for execution without LLM API keys via a deterministic fallback mode for testing workflow, RAG, evaluation, and UI components.

https://github.com/Ricky-7-Yan/intelligent-audit-system

agent-evaluationagent-runtimeagentic-ragai-agentauditevaluation-harnessfastapihuman-in-the-loopknowledge-graphllmopsmcpmulti-agentpythonrag

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.