Awesome Infra for AI › LLM Evaluation & Testing

jsdhwfmax/EvalForge

⭐ 210 Python added to this list on 2026-09-14 repository created 2026-08-24

EvalForge is an evaluation interoperability and release-gating layer for RAG applications and AI assistants. Rather than replacing existing evaluation frameworks such as Ragas, DeepEval or promptfoo, it imports their differently shaped results into one versioned, evaluator-neutral JSON artifact, compares it against a baseline, applies an explicit versioned policy, and emits reports in formats CI systems already understand: JSON, JUnit XML, SARIF 2.1.0 and Markdown job summaries. The workflow is to import documents and golden questions with labeled relevant sources, run a baseline and a candidate configuration, compare multi-dimensional deltas on the same dataset fingerprint, apply quality, security, latency and cost thresholds, and return a non-zero exit code on regression so a CI pipeline can block a release. A built-in RAG evaluator supports BM25, deterministic vector and hybrid retrieval, and can score answers offline with an extractive baseline or against OpenAI-compatible endpoints, covering token-F1 correctness, citation support, hallucination proxies, latency, token counts and configured cost, plus basic security checks like prompt-injection and canary-exfiltration probes. Every artifact carries deterministic SHA-256 digests of inputs and policy plus producer and source-revision metadata for audit purposes, and an opt-in comparison mode blocks the gate when dataset fingerprints, evaluator versions or metric definitions do not match between baseline and candidate. It ships as a PyPI package with a CLI, a FastAPI/OpenAPI service, a Streamlit comparison dashboard, and a reusable GitHub Action, targeting teams that need machine-readable, auditable evidence that an LLM or RAG change did not regress quality before it reaches production.

https://github.com/jsdhwfmax/EvalForge

LLM evaluationRAG evaluationCI gatepromptfooRagasDeepEval

Also in LLM Evaluation & Testing

promptfoo/promptfoo

Promptfoo is a CLI and library for evaluating LLM applications, offering automated testing, red teaming, and vulnerability scanning for prompts, models, agents, and RAGs.

ifixai-ai/iFixAi

iFixAi is a diagnostic tool that evaluates AI models and agents for operational misalignment, including fabrication, manipulation, deception, unpredictability, and opacity, by running up to 45 insp...

confident-ai/deepeval

DeepEval is an open-source LLM evaluation framework, offering a variety of metrics and tools for assessing the performance of AI agents, RAG pipelines, and chatbots through unit testing.

vibrantlabsai/ragas

Ragas is an evaluation framework for LLM applications that provides objective metrics, test data generation, and feedback loops for continuous improvement.

NVIDIA/garak

Garak is an open-source LLM vulnerability scanner designed to red-team and assess generative AI models for weaknesses like hallucination, data leakage, prompt injection, and toxicity.

evidentlyai/evidently

Evidently is an open-source Python framework for evaluating, testing, and monitoring ML and LLM systems, providing comprehensive data and model quality checks from experiments to production.

open-compass/opencompass

Open-source LLM evaluation platform for running standardized benchmarks across models, with configurable datasets, prompt templates, and an official public leaderboard.

Giskard-AI/giskard-oss

Giskard is an open-source Python library for testing and evaluating agentic systems and LLM applications, offering tools for scenario-based testing, red teaming, and vulnerability scanning.