Awesome Infra for AI › AI Safety & Guardrails

dodoguardai/dodoguard

⭐ 202 TypeScript repository created 2026-08-25

DodoGuard is a self-hosted platform for evaluating and red-teaming LLM-based agents before and after release. It targets five risk categories: prompt injection and jailbreaks that bypass system prompts and guardrails; data and tool abuse such as sensitive-data leaks or unauthorized tool/MCP calls; quality regressions after a model or prompt change; configuration and attack-surface gaps on agent platforms (Dify, Coze, FastGPT, n8n); and the lack of durable, private evidence behind a release decision. Its workflow follows a lifecycle of intake, detect-and-release, and observe, built around the principle that acknowledging a finding is not the same as releasing with it unresolved. The product registers agents as assets in a ledger, runs versioned detection tasks - LLM evaluation (prompt, provider, and variable-case runs with multiple assertion types), red-team testing (plugins and adaptive probes against live apps, not just base models), and platform-hardening scans - and surfaces findings with severity and evidence through a dashboard. Reports can be exported for release sign-off, and a remediation loop lets a team fix an issue, retest it, and then decide whether to go live. DodoGuard works against OpenAI, Anthropic, Azure, Bedrock, Ollama, and other OpenAI-compatible providers. It ships as a CLI (Node.js 22.12 or newer), a web console started with Docker Compose, and an Electron desktop app, and is designed to run entirely on the user's own infrastructure so prompts, traces, and findings never leave it. A quick CLI run against a bundled example evaluates factuality in about a minute. The project is MIT-licensed and documents a private vulnerability-disclosure process.

https://github.com/dodoguardai/dodoguard

red teamingllm evaluationai safetyguardrailsprompt injectionagent securityself-hosted

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.