Awesome Infra for AI › AI Safety & Guardrails

nizos/probity

⭐ 219 TypeScript added to this list on 2026-09-21 repository created 2026-04-18

Probity is a rule-enforcement layer for AI coding agents, running before every file write and shell command an agent issues rather than only at commit time. Rules are TypeScript functions that inspect the action and recent session activity and either allow it or block it with a message explaining the violation and the next legal step, which the agent then reads and corrects course from. Rules can be purely deterministic, string or regex matches on a command or a file's content, or AI-validated, using each provider's official SDK to judge an action in context; deterministic rules add no extra tokens or latency, while AI-validated rules add a model call per check. Built-in rules include one that requires a failing test before production code and blocks skipping straight to implementation, ones for blocking dangerous shell commands or unwanted code patterns, one for enforcing that a command must have run first, such as tests before a commit, and one for enforcing filename casing. Because it reads the agent's own session transcript directly rather than instrumenting a specific test framework, it works with any language and test runner the agent already supports, and one configuration file works across supported agents. It positions itself as a successor to an earlier, similarly scoped project, with claimed improvements in handling refactors, multi-step edits, and concurrent agent sessions. It is aimed at teams and individuals who want an AI coding agent to reliably follow a development discipline, most commonly test-driven development, or a safety policy without manually re-prompting or reviewing every step.

https://github.com/nizos/probity

ai-coding-agentstdd-enforcementguardrailsclaude-codepolicy-enforcement

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.