Awesome Infra for AI › LLM Evaluation & Testing

ukanwat/aaabench

⭐ 387 Shell added to this list on 2026-08-17 repository created 2026-07-31

AAABench is an evaluation harness for coding agents built around a task that cannot be pattern-matched: build an open-world game in Unreal Engine 5. The harness boots a real editor, exposes it to the agent live over MCP, hands over a demanding brief and a shelf of production knowledge, and then stays out of the way. Scripts in the repository set up the required plugins, renderer features and Python libraries, copy a project skeleton, and keep the agent running across sessions. The governing rule of the harness is that the operator supplies conditions, resources and the demand but never a diagnosis, a fix or an answer — whether the model notices its own mistakes is part of what is being measured, so any hint invalidates the result. What it probes is deliberately broad: real-world understanding, since a city built without knowledge of ports, rail, money and sunlight looks wrong to any observer; causal reasoning, because the brief forbids placing anything merely because it looked good; long-horizon execution across cold-started sessions where the only continuity is what the agent chose to write down; writing, since the world needs a story bible, characters, missions, signage and street names; and self-verification, because the agent has viewport capture, play-in-editor and its own screenshots to judge its work against reference photography. The repository ships the harness, the brief and the rules rather than a leaderboard — the author invites others to run it on a different model and report what broke. It is MIT-licensed, driven by shell scripts, and aimed at people who want a long-running, unfakeable measurement of agent capability rather than a short question-and-answer benchmark.

https://github.com/ukanwat/aaabench

agent-evaluationbenchmarkharnessmcplong-horizontesting

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.