Awesome Infra for AI › AI Safety & Guardrails

AI Safety & Guardrails

30 projects

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

⭐ 11157 Python added 2026-06-29

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

⭐ 7246 Python

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

⭐ 6766 TypeScript

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

⭐ 6748 Python

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

⭐ 6392 Python

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

⭐ 5247 MDX added 2026-08-24

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

⭐ 3213 Python

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.

⭐ 2480 JavaScript added 2026-08-24

kenryu42/cc-safety-net

A PreToolUse hook that blocks destructive commands and secret access before AI coding agents run them, parsing command semantics so shell wrappers and flag reordering cannot bypass it.

⭐ 1576 TypeScript added 2026-08-17

archestra-ai/OpenAPPA

OpenAPPA is a deterministic policy engine (APPA) placed between an AI agent and its tools that checks every tool call against a declarative data-sensitivity policy before it runs, blocking unauthorized data flows.

⭐ 1405 Rust

wuyoscar/Internal-Safety-Collapse

ISC-Bench is a research project and benchmark for evaluating the "Internal Safety Collapse" vulnerability in LLMs, where models bypass safety controls when completing complex tasks.

⭐ 1198 Python

cuga-project/cuga-agent

CUGA is an open-source generalist agent harness for enterprises, supporting complex task execution on web and APIs, OpenAPI/MCP integrations, composable architecture, reasoning modes, and policy-aw...

⭐ 888 Python

toby-bridges/api-relay-audit

A local security audit tool for AI API relays and LLM proxies, designed to detect prompt injection, model substitution, tool-call rewriting, and other tampering.

⭐ 865 Python

hoophq/hoop

Open-source sidecar that sits in front of databases and other resources to mask sensitive data and block destructive queries before AI agents can run them.

⭐ 830 Go added 2026-09-28

decionis/agent-safe-pipeline

Reference architecture and TypeScript library that routes AI agent actions through an independent authorization boundary, turning proposals into policy verdicts, human approvals and single-use grants.

⭐ 589 TypeScript added 2026-08-24

cordum-io/cordum

Source-available control plane that gates AI-agent actions behind approval workflows, a policy safety kernel, and a compliance firewall before they execute.

⭐ 510 Go added 2026-09-07

deadbits/vigil-llm

Vigil is a security scanner for LLM prompts and responses, designed to detect prompt injections, jailbreaks, and other adversarial attacks using various scanning methods like vector databases, YARA...

⭐ 498 Python

Justin0504/Aegis

Aegis is a pre-execution firewall for AI agents, providing runtime policy enforcement, cryptographic audit trails, human-in-the-loop approvals, and a kill switch without code changes.

⭐ 486 TypeScript

ZaxbyHub/opencode-swarm

OpenCode Swarm is an architect-centric agentic plugin for OpenCode that orchestrates specialized AI agents to generate, review, and test code, enforcing rigorous gated pipelines and security guardr...

⭐ 485 TypeScript added 2026-06-22

SponsioLabs/Sponsio

Sponsio provides deterministic runtime safety solutions for AI agents, enforcing contracts on agent procedures in milliseconds with zero LLM cost.

⭐ 440 Python

PrismorSec/prismor

Prismor provides runtime security hooks for AI coding agents, blocking dangerous commands, preventing secret leaks, and defending against prompt injection.

⭐ 401 Python added 2026-07-06

pegasi-ai/reins

Reins provides security controls for AI agents by enforcing deterministic policies, scanning for vulnerabilities, tracking drift with an immutable audit trail, and intervening on risky actions.

⭐ 392 Python

agentcontrol/agent-control

Agent Control provides a centralized control plane for enforcing runtime guardrails and safety policies for AI agents, blocking prompt injections, PII leakage, and other risks.

⭐ 320 Python

ucsandman/DashClaw

DashClaw is an AI agent governance runtime that intercepts actions, enforces guard policies, manages approvals, and produces audit-ready decision trails for AI agents interacting with real systems.

⭐ 310 TypeScript

DobermanCore/Doberman-Core

Local, open-source runtime guardrail that sits between AI coding agents and their tools, giving every action a PASS/AUTH/BLOCK verdict to stop destructive or exfiltrating commands before they run.

⭐ 277 Python added 2026-09-21

praetorian-inc/julius

A Go tool that fingerprints LLM serving infrastructure on network endpoints, identifying which of 60+ AI services — Ollama, vLLM, LiteLLM, TGI and others — is running behind a port.

⭐ 243 Go added 2026-08-17

nizos/probity

Hooks into Claude Code, Codex, and GitHub Copilot CLI to enforce coding rules, such as strict TDD or a ban on destructive commands, by checking every file write and shell command before it runs.

⭐ 219 TypeScript added 2026-09-21

privacera/paig

PAIG (Privacera AI Guardrails) is an open-source framework designed to protect Generative AI applications by ensuring security, safety, and observability for responsible AI deployment.

⭐ 210 CSS

dodoguardai/dodoguard

DodoGuard is a self-hosted platform (CLI, web console, desktop app) that evaluates LLM prompt quality, red-teams agents for jailbreaks and tool or data abuse, audits agent platforms like Dify and n8n, and produces private release-evidence reports.

⭐ 202 TypeScript

openguardrails/openafw

Local proxy that sits in front of Anthropic, OpenAI, and Gemini API calls from coding agents, replacing secrets in outgoing requests with placeholders and restoring them in responses so raw keys never reach the model or provider.

⭐ 106 Rust added 2026-09-21