Awesome Infra for AI › LLM Evaluation & Testing

aaron-for-value/VeriRun

⭐ 323 Python added to this list on 2026-09-07 repository created 2026-07-22

VeriRun is a pre-alpha Python runtime for running code- and agent-evaluation workloads (benchmarks such as EvalPlus, LiveCodeBench, Harbor, and Terminal-Bench) in a way that is reproducible, recoverable, attributable, and appropriately sandboxed for untrusted, model-generated code. It positions itself as the runtime underneath executable evaluation and reward computation, not another leaderboard. Its design rests on six invariants: immutable provenance for every result (benchmark, prompt, candidate, tests, verifier image, model revision, sampling config, runtime policy); replay of frozen candidates without re-calling the model; structured failure semantics that keep compile errors, test failures, timeouts, OOMs, policy violations, and infrastructure failures distinct; isolation backed by an explicit threat model and attack-regression suite (Kubernetes + gVisor); at-least-once execution with idempotent, effectively-once result commits; and separate reporting of model capability versus infrastructure reliability. The target architecture routes CLI/API/CI requests through an eval control plane, a run/task/attempt store, admission and budget controls, a Ray/KubeRay orchestrator, a bounded async OpenAI-compatible model gateway, and a sandbox manager that executes Kubernetes Jobs under gVisor, with results committed idempotently and exported via OpenTelemetry. Delivery is incremental: released milestones cover immutable manifests and deterministic replay (v0.1), a bounded async model gateway with classified retries (v0.2), and a restricted local Kubernetes/gVisor sandbox (v0.3); planned work adds Ray/KubeRay scaling, full observability, and integration with the veRL asynchronous reward loop for post-training. It is aimed at teams building custom code- and agent-evaluation or RL-reward pipelines who need reproducibility and isolation guarantees beyond ad hoc benchmark scripts.

https://github.com/aaron-for-value/VeriRun

agent evaluationcode execution sandboxbenchmarkingreproducibilitygVisorkubernetesreward computation

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.