promptfoo/promptfoo
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
Awesome Infra for AI › LLM Evaluation & Testing
VeriRun is a pre-alpha Python runtime for running code- and agent-evaluation workloads (benchmarks such as EvalPlus, LiveCodeBench, Harbor, and Terminal-Bench) in a way that is reproducible, recoverable, attributable, and appropriately sandboxed for untrusted, model-generated code. It positions itself as the runtime underneath executable evaluation and reward computation, not another leaderboard. Its design rests on six invariants: immutable provenance for every result (benchmark, prompt, candidate, tests, verifier image, model revision, sampling config, runtime policy); replay of frozen candidates without re-calling the model; structured failure semantics that keep compile errors, test failures, timeouts, OOMs, policy violations, and infrastructure failures distinct; isolation backed by an explicit threat model and attack-regression suite (Kubernetes + gVisor); at-least-once execution with idempotent, effectively-once result commits; and separate reporting of model capability versus infrastructure reliability. The target architecture routes CLI/API/CI requests through an eval control plane, a run/task/attempt store, admission and budget controls, a Ray/KubeRay orchestrator, a bounded async OpenAI-compatible model gateway, and a sandbox manager that executes Kubernetes Jobs under gVisor, with results committed idempotently and exported via OpenTelemetry. Delivery is incremental: released milestones cover immutable manifests and deterministic replay (v0.1), a bounded async model gateway with classified retries (v0.2), and a restricted local Kubernetes/gVisor sandbox (v0.3); planned work adds Ray/KubeRay scaling, full observability, and integration with the veRL asynchronous reward loop for post-training. It is aimed at teams building custom code- and agent-evaluation or RL-reward pipelines who need reproducibility and isolation guarantees beyond ad hoc benchmark scripts.
https://github.com/aaron-for-value/VeriRun
Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.
iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...
DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.
Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.
Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.
Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.
Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.
Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.