Awesome Infra for AI › AI Safety & Guardrails

ucsandman/DashClaw

⭐ 310 TypeScript repository created 2026-02-08

DashClaw provides a critical governance layer for AI agents, specifically designed for those that interact with external systems. It functions by intercepting potentially risky actions initiated by agents, evaluating these actions against defined policies, and, when necessary, routing them for human approval. The platform ensures verifiable evidence is recorded for every decision and tracks terminal outcomes to prevent issues like silent double-execution during agent retries. Key functionalities include intercepting actions with policies for blocking, warning, or approval; verifying agent identity using OIDC bearer tokens with replay protection; enforcing declarative policies such as risk thresholds and access rules; and managing approvals through a centralized dashboard or integrated communication channels like Telegram and Discord. Every action processed by DashClaw generates a replayable decision record, capturing the agent's intent, reasoning, risk score, and applied policies. It also finalizes terminal outcomes durable, ensuring lost confirmations are surfaced and retries don't lead to duplicate executions. DashClaw can govern external systems by wrapping HTTP APIs with a capability registry, allowing for per-agent access rules and rate limits, and composing these into governed workflows. The platform integrates with various agent frameworks like LangChain, CrewAI, and AutoGen, and provides a control plane with a decisions ledger, mission control for fleet posture, and analytics for cost, volume, and policy enforcement.

https://github.com/ucsandman/DashClaw

agent-governanceai-agentsai-governancepolicy-enforcementaudit-trailagent-runtimeai-safetyguardrailsllm-opsai-opsdecision-engine

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.