Awesome Infra for AI › AI Safety & Guardrails

NVIDIA-NeMo/Guardrails

⭐ 7246 Python repository created 2023-04-18

NVIDIA NeMo Guardrails is an open-source library designed to integrate programmable guardrails into large language model (LLM) based conversational applications. Its primary goal is to enhance the safety, trustworthiness, and controlled behavior of AI assistants by enabling developers to define specific rules and constraints for LLM outputs and interactions. The toolkit helps prevent LLMs from generating undesirable content, engaging in off-topic discussions, or deviating from predefined conversational paths. It introduces concepts like input rails, dialog rails, output rails, retrieval rails, and fact-checking rails to provide comprehensive control over the LLM's responses and interactions with external tools and services. NeMo Guardrails allows for the definition of guardrails to address common LLM vulnerabilities such as jailbreaks and prompt injections. It supports various use cases, including question-answering systems over documents (Retrieval Augmented Generation), domain-specific assistants, securing LLM endpoints, and integrating with frameworks like LangChain. The library is designed to be easily integrated into existing LLM applications with minimal code changes, offering both synchronous and asynchronous APIs. It is compatible with a wide range of LLMs, including OpenAI models like GPT-3.5 and GPT-4, LLaMa-2, Falcon, Vicuna, and Mosaic models, making it a flexible solution for enhancing the operational safety and reliability of LLM deployments.

https://github.com/NVIDIA-NeMo/Guardrails

LLMguardrailsAI safetysecurityconversational AINVIDIAPythonLLM governanceprompt safetyagent control

Also in AI Safety & Guardrails

data-privacy-stack/presidio

Presidio is an open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data, leveraging NLP and customizable pipelines.

superagent-ai/superagent

Superagent is an open-source SDK providing safety features for AI applications, including prompt injection detection, PII redaction, repository scanning for threats, and red teaming capabilities fo...

Tencent/AI-Infra-Guard

AI-Infra-Guard is a full-stack AI red teaming platform providing comprehensive security analysis, vulnerability scanning, and jailbreak evaluation for AI ecosystems and LLMs.

microsoft/agent-governance-toolkit

AI Agent Governance Toolkit (AGT) provides policy enforcement, identity management, execution sandboxing, and reliability engineering to secure autonomous AI agents in production.

FailproofAI/failproofai

Observability and policy enforcement for AI agent harnesses, hooking twelve coding and chat harnesses to record every run and block dangerous tool calls before they execute.

protectai/llm-guard

LLM Guard is a comprehensive open-source security toolkit designed to fortify Large Language Model (LLM) interactions by providing robust sanitization, malicious content detection, data leakage pre...

lennney/stop-that-shit

Multi-platform hook and skill guard for AI coding agents that blocks unrequested work such as generated hashes, checksums and task-scope expansion at the harness hook boundary.

kenryu42/cc-safety-net

A PreToolUse hook that blocks destructive commands and secret access before AI coding agents run them, parsing command semantics so shell wrappers and flag reordering cannot bypass it.