Awesome Infra for AI › AI Safety & Guardrails

protectai/llm-guard

⭐ 3213 Python repository created 2023-07-27

LLM Guard is an open-source security toolkit for safeguarding interactions with Large Language Models (LLMs) in production. It provides a suite of features to ensure the security, privacy, and integrity of LLM applications. The toolkit focuses on preventing common vulnerabilities and exploits associated with LLMs, such as prompt injection, data leakage, and the generation of harmful or biased content. It offers various 'scanners' for both prompt input and LLM output. Prompt scanners include anonymization of sensitive data, detection of banned code or topics, identification of prompt injection attempts, gibberish detection, and sentiment analysis. Output scanners extend these capabilities to LLM responses, allowing for the detection of bias, malicious URLs, refusal to answer, and factual consistency checks. The tool is designed for easy integration into existing LLM-powered applications and aims to be deployment-ready for production environments. By offering sanitization, threat detection, and prevention mechanisms, LLM Guard enables developers to build more secure and reliable LLM-based systems, ensuring that interactions remain safe and compliant.

https://github.com/protectai/llm-guard

LLM securityguardrailsprompt injectiondata leakage preventioncontent moderationAI safetyLLMOpssecurity toolkit

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

NVIDIA-NeMo/Guardrails

NVIDIA NeMo Guardrails is an open-source toolkit for adding programmable guardrails to LLM-based conversational applications, focusing on safety, security, and controlled dialog.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.

kenryu42/cc-safety-net

A PreToolUse hook that blocks destructive commands and secret access before AI coding agents run them, parsing command semantics so shell wrappers and flag reordering cannot bypass it.