Awesome Infra for AI › AI Safety & Guardrails

lennney/stop-that-shit

⭐ 2480 JavaScript added to this list on 2026-08-24 repository created 2026-08-11

Stop That Shit is a guard that constrains what an AI coding agent is allowed to do beyond the task it was actually given. The problem it addresses is defensive busywork the agent invents on its own: emitting checksums nobody asked for, adding compatibility layers, writing extra tests, or quietly widening the scope of a change so that the requested work is buried in unrelated edits. The project ships as a pair of components. A skill layer states the boundary in terms the agent reads before acting, and an executable guard runs on the harness hook path and checks concrete, decidable limits before a step proceeds: which files may be touched, which dependencies may be added, whether hashes and checksums may be produced, and what sub-agent work is in budget. When the guard determines that an action crosses the boundary it returns a refusal in context, so the agent is redirected rather than silently allowed. Four harnesses are supported: Codex, Claude Code, OpenCode and the Hermes agent command line tool, with the same rules applied through each one native hook mechanism. Working modes such as review and change carry different limits, so a read-oriented session is bounded more tightly than an implementation session. The repository documents its behaviour through a case library of paired bad and good outcomes, and the README is published in Chinese and English. It suits teams whose agents run against real repositories and who found that instructions written into project guidance files erode as those files grow.

https://github.com/lennney/stop-that-shit

agent-guardrailshooksscope-controlcoding-agentsclaude-codejavascript

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

kenryu42/cc-safety-net

A PreToolUse hook that blocks destructive commands and secret access before AI coding agents run them, parsing command semantics so shell wrappers and flag reordering cannot bypass it.