Awesome Infra for AI › AI Safety & Guardrails

archestra-ai/OpenAPPA

⭐ 1405 Rust repository created 2026-08-18

OpenAPPA implements APPA (Agentic Permissions Policy Algebra), a deterministic alternative to probabilistic classifiers and PII detectors for controlling what data an AI agent is allowed to send to which tool. It tracks the sensitivity and trust level of everything an agent has read during a session and, before each tool call executes, checks that call against the tracked data using a policy written in declarative TOML. The decision engine works purely from the event log, makes no network or file calls, and therefore returns the same decision for the same log every time. It can run in-process inside an agent or as a separate sidecar process that intercepts tool calls. The project reports benchmark results on two suites: Bench-Corp (20 multi-step enterprise workflows) and AgentThreatBench (OWASP Top 10 for Agentic Applications), under both standard and adversarial prompts, claiming zero successful attacks across 1,320 evaluations while completing 88-90% of legitimate tasks, compared with Microsoft FIDES (28-35% of attacks succeeding) and Claude Code auto mode (10 attacks succeeding). A Claude Code plugin lets users try the policy engine interactively via an install script and a guided setup skill. The APPA runtime can also be embedded directly in a custom agent from any language, or applied at the LLM proxy level through the separate Archestra product, which wires OpenAPPA into Claude Code, Claude Desktop, Cursor, Codex, OpenCode, Copilot CLI and n8n. A CLI (`appa describe --check`, `appa replay`) validates policy configuration and replays scripted tool calls against expected decisions, suited to CI gating. The project is described by its authors as a preview and RFC, with a companion paper accepted at a NeurIPS 2026 workshop; it is MIT licensed.

https://github.com/archestra-ai/OpenAPPA

ai-safetyguardrailsagent-securitypolicy-engineaccess-controlinformation-flow-controlllm-agentsclaude-code-plugin

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.