Awesome Infra for AI › AI Safety & Guardrails

microsoft/agent-governance-toolkit

⭐ 6392 Python repository created 2026-03-02

The Microsoft Agent Governance Toolkit (AGT) is an open-source framework designed to ensure the safe and reliable operation of autonomous AI agents in production environments. AGT addresses critical challenges such as preventing unauthorized actions, attributing actions to specific agents in multi-agent systems, and providing tamper-evident records for auditing and compliance. Unlike prompt-based safety mechanisms, AGT intercepts and evaluates every tool call, message send, and delegation in deterministic application code *before* the model's intent executes, making denied actions structurally impossible rather than merely unlikely. The toolkit offers a robust policy engine that allows developers to define YAML-based policies for governing agent behavior, including blocking destructive operations, requiring human approval for sensitive actions, and enforcing identity and access controls. It supports integration with various agent frameworks and programming languages (Python, TypeScript, .NET), providing a `govern` wrapper for tools and a `PolicyEvaluator` API for programmatic control. AGT is built to align with security best practices, explicitly covering all 10 items of the OWASP Agentic Top 10. By preventing agents from misbehaving at the application layer, AGT helps organizations deploy AI agents with greater confidence and maintain regulatory compliance.

https://github.com/microsoft/agent-governance-toolkit

agent-frameworkai-agentsai-safetycompliancegovernancemicrosoftowasppolicy-enginepythonsecuritytrustzero-trust

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.

kenryu42/cc-safety-net

A PreToolUse hook that blocks destructive commands and secret access before AI coding agents run them, parsing command semantics so shell wrappers and flag reordering cannot bypass it.