# AI Guardrails AI guardrails are constraints applied to AI systems to prevent harmful, unintended, or low-quality outputs. They operate at multiple levels: training-time alignment, runtime filtering, and system-level restrictions. ## Types **Input guardrails** filter what reaches the model: - [[Prompt injection]] detection and blocking - Content moderation on user inputs - Input validation and sanitization **Output guardrails** filter what leaves the model: - Toxicity and bias detection - Hallucination detection (cross-referencing claims against known sources) - Format validation (ensuring structured outputs match expected schemas) - Confidence thresholds (flagging low-confidence responses) **Action guardrails** constrain what agents can do: - Permission systems (which tools an agent can use, which files it can modify) - Human-in-the-loop approval for high-risk actions (sending messages, deleting data, deploying code) - Rate limiting and cost caps - Sandbox environments for code execution In [[Agentic Engineering]], action guardrails are critical because the [[Agentic loops|agentic loop]] gives the model autonomous power. [[Claude Code]] implements this through its permission system: some tool calls require explicit user approval. The [[Lethal Trifecta for AI Agents]] describes what happens when guardrails fail in agentic systems. Guardrails complement but don't replace [[AI Alignment]] (training the model to want the right things) and [[Responsible AI]] practices (organizational policies for safe deployment). ## Guardrails as a classifier battery There are three common ways to put rules around an LLM, and each has a weak spot. Rules in the **system prompt** sit exactly where a jailbreak talks its way past them. A **second LLM as a guard** adds a full LLM call of latency and cost to every turn, and it can be jailbroken too. **The lab's built-in refusals** are drawn differently by each lab and move with every model version, so they aren't *your* policy. The alternative I find most convincing is a battery of small classifiers on both sides of the LLM call, with the actual rules written in code. It's an old idea (moderation classifiers have been around for years), and [[Decision Models (DMs)|decision models]] like [[Jev]] make it cheap enough to run on every message going in AND every reply coming out. TypeSafe's guardrails cookbook does it with one request per message: four Noul questions (does it try to override the assistant's instructions? does it ask for help with harm or a crime? does it ask for a diagnosis or a dosage? does it signal self-harm?) plus one severity Score from "none" to "serious physical harm". The model returns probabilities; code decides: - Each hazard has a **review threshold** and an **action threshold**. The action can be block, or **support** (a crisis path for self-harm instead of hanging up on someone) - The severity Score can turn a review into a block - A **policy** is just those numbers under a name. "Strict": review at 0.35, action at 0.70, block when severity ≥ 2.0. "Permissive": action at 0.85 Results on 10 prompts and 5 replies (the jailbreaks are real, taken from a public in-the-wild collection): a DAN jailbreak scored 0.98 and was blocked; a lockpicking-for-burglary request scored 0.95 on harm and was blocked; a self-harm message scored 0.96 and went to support; a melatonin dosage question scored 0.55 and went to review; a novelist asking how a detective describes poisoning *passed* (0.05), because fiction framing isn't intent. On the output side, a reply that declined to help passed (0.07) and a jailbroken reply scored 0.94 and was blocked. The same assessment (jailbreak 0.74, severity 0.51) gets blocked under "strict" and sent to review under "permissive": the probabilities don't move, only the product decision does. "Ignore your instructions" gets *scored* as a jailbreak instead of *working* as one. That's the part I like: the guard never follows the text it reads. Two caveats: TypeSafe's own docs say the state isn't treated as hostile by default and adversarial content can still move answers, and the thresholds should come from labeled examples of your own traffic. See [[Confidence-Gated Routing]] for the threshold logic. ### RAG passages are an input too Retrieved passages flow into the prompt just like user input, so they deserve the same treatment. TypeSafe's RAG passage cookbook asks four Nouls per retrieved passage (`is_relevant`, `contains_answer_evidence`, `contradicts_query_premise`, `contains_prompt_injection`) and routes in code, first match wins: 1. injection > 0.70 → exclude (security goes first) 2. contradicts the premise > 0.70 → put it in a separate *conflict* block (before the evidence test, because a passage that denies the premise also contains usable text) 3. relevant < 0.45 → exclude 4. evidence > 0.55 → include 5. otherwise → exclude On 81 Supabase auth doc passages with one planted forum injection, embedding similarity ranked the injection **first** (cosine 0.584, in a range of only 0.455 to 0.584). The Noul scored it 0.99 and dropped it. A passage refuting a false premise, which relevance (0.49) and evidence (0.51) alone would have dropped, scored 0.92 on contradiction and went to the conflict block. The generator then refused to invent a 30-day setting that doesn't exist and flagged the conflict. That's [[Prompt injection]] defense and [[AI Hallucination]] prevention in the same pass. The cookbook is explicit that the injection Noul is a filter, not a security boundary: the prompt still has to treat every passage as untrusted. See also [[Reranking]] and [[RAG Pipelines]]. ## References - [Guardrails for LLMs cookbook (TypeSafe docs)](https://docs.typesafe.ai/cookbooks/llm_guardrails) - [Classifying RAG passages cookbook (TypeSafe docs)](https://docs.typesafe.ai/cookbooks/classifying_rag_passages) - [In-the-wild jailbreak prompts dataset (Hugging Face)](https://huggingface.co/datasets/TrustAIRLab/in-the-wild-jailbreak-prompts) ## Related - [[AI Safety]] - [[AI Alignment]] - [[AI Hallucination]] - [[Prompt injection]] - [[Responsible AI]] - [[Agentic Engineering]] - [[Agentic loops]] - [[AI Agent Harness]] - [[Claude Code]] - [[Lethal Trifecta for AI Agents]] - [[EU AI Act]] - [[Confidence-Gated Routing]] - [[Decision Models (DMs)]] - [[Jev]] - [[Reranking]] - [[RAG Pipelines]] - [[Human-in-the-Loop]]