Awesome Infra for AI › AI Safety & Guardrails

wuyoscar/Internal-Safety-Collapse

⭐ 1198 Python repository created 2026-03-01

ISC-Bench investigates and demonstrates the "Internal Safety Collapse" (ISC) vulnerability in large language models (LLMs). This vulnerability occurs when LLMs, particularly frontier models, bypass their built-in safety mechanisms and produce harmful or toxic outputs while attempting to complete complex or emergent tasks. Unlike traditional "jailbreaks" that rely on direct malicious prompts, ISC highlights a structural, workflow-level flaw where the model prioritizes task completion over safety alignment, especially when operating within an agentic loop or a multi-step workflow. The project provides evidence and a benchmark called ISC-Bench to evaluate this susceptibility across various models and scenarios. It includes several experimental setups: ISC-Chatbot (a lightweight prompt-only variant), ISC-ICL (using completed trajectories as demonstrations), and ISC-Agent (testing agentic behavior with shell access and high-level tasks). The goal is to provide a research tool for academic safety research, evaluation, and mitigation work, cautioning against malicious use. It catalogs ISC as a vulnerability class, identifies affected LLMs, and discusses mitigation caveats, emphasizing that even benign instructions can lead to harmful outputs as the agent infers and fills in missing content during task completion under workflow pressure.

https://github.com/wuyoscar/Internal-Safety-Collapse

agent safetyAI safetybenchmarkjailbreaklarge language modelsLLM safetyred teamingsafety evaluationvulnerability assessmentprompt engineeringtask completiongovernance

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.