Guardrails
Also known as: AI guardrails, safety filters
Checks placed around an AI system that block unsafe requests going in and harmful or leaking answers coming out.
Draft - this entry has not been reviewed yet.
Formal
Rules, filters and small checking models that run before and after a large language model, screening input and output against a policy - blocking jailbreak attempts, removing personal data, forcing an allowed format - independently of how the model was trained.
In plain English
Like the barriers along a mountain road - they do not steer the car, but they stop it from going over the edge when the driver makes a mistake.
In practice
A webshop's customer service manager has every chat message pass a filter that flags jailbreak wording, and every reply pass a second one that hides card numbers before the customer sees it.
Why it matters
Training never makes a model fully safe, so outside checks give a second, testable line of defence that the owner can change quickly without training the model again.
How to put it into practice
The usual steps, in order. Adapt them to your organisation.
- Write a short policy for the application stating what it must never accept or say, such as jailbreak attempts, personal data like CPR and card numbers, secrets, off-topic or harmful content, and which output format is allowed.
- Map the points where checks can sit, namely the user's message, retrieved documents, tool results, tool calls and the final answer, and decide which rule applies where.
- Start with deterministic checks, such as regular expressions with a Luhn check for card numbers, allowlists for URLs and domains, length limits and JSON Schema validation of structured output.
- Add classifier checks for what rules cannot catch, such as a prompt-injection detector on user input and retrieved content, a safety classifier on answers and a PII detector, and use an LLM judge only where nothing simpler works.
- Decide for each check whether it blocks, rewrites, masks, asks the user again or hands the case to a person, and log every hit without storing more personal data than needed.
- Build a labelled test set with normal requests and attacks, including Danish, encoded and multi-turn variants, measure false positives and false negatives, and tune the thresholds per use case.
- For streamed answers, check the text in chunks with the option to retract it, or buffer high-risk answers until the output check has seen them in full.
- Rerun the test set whenever the model, prompt or guardrail changes, add new attacks from incidents and red teaming, and review block rates and user complaints every month.
Common pitfalls
- Checking only the user's message, while indirect prompt injection arrives through retrieved documents, web pages and tool results.
- Treating a guardrail as a wall instead of a classifier with error rates, and giving the agent broad permissions because the filter should catch misuse.
- Tuning only on English examples, so Danish, encoded or split-up attacks pass straight through.
- Blocking so much that users give up and move to unmanaged AI tools.
Good guides
- OWASP Top 10 for LLM Applications 2025(opens in a new tab) · OWASP
- LLM Prompt Injection Prevention Cheat Sheet(opens in a new tab) · OWASP
- Prompt Shields in Azure AI Content Safety(opens in a new tab) · Microsoft
- NIST AI 600-1 - Artificial Intelligence Risk Management Framework, Generative Artificial Intelligence Profile(opens in a new tab) · NIST
Technical deep dive
Architecturally, guardrails are policy enforcement points in the request path of an LLM application, analogous to a WAF or DLP gateway. NVIDIA's NeMo Guardrails makes the stages explicit: input rails on the user message, retrieval rails on RAG chunks before they enter the context, dialog rails that steer conversation flow (defined in its Colang language), execution rails around tool calls, and output rails on the generated response. Each rail can reject, rewrite, redact, ask a clarifying question, or escalate to a human. Placing a check on retrieved content and tool results - not only on the user's message - is what makes guardrails relevant to indirect prompt injection.
Implementations fall into three classes. Deterministic checks: regular expressions and validators for card numbers (with a Luhn check), CPR numbers, API-key formats, URL and domain allow-lists, length limits, and JSON Schema validation of structured output. Classifier models: purpose-trained safety classifiers such as Meta's Llama Guard family (an LLM fine-tuned to label prompts and responses against a hazard taxonomy, from Llama Guard 3 aligned with the MLCommons taxonomy), prompt-injection detectors such as Azure AI Content Safety Prompt Shields, and PII detectors based on named-entity recognition. LLM-as-judge: a second model prompted with a policy that grades the output. Constrained decoding - grammar- or schema-guided generation - is a related technique that prevents malformed output at generation time instead of filtering it afterwards.
Every guardrail is a classifier with a false-positive and a false-negative rate, and both matter: over-blocking pushes users to unmanaged tools, under-blocking lets attacks through. Guardrails should be evaluated on labelled test sets, including multilingual and encoded variants, with thresholds tuned per use case. Known weaknesses include obfuscation (Base64, homoglyphs, splitting a payload across turns), low-resource languages the classifier was not trained on, the classifier itself being susceptible to prompt injection when it is an LLM, and streaming, where tokens reach the user before an output check has seen the full response; mitigations are chunked checking with the ability to retract, or buffering high-risk responses. Each extra model call also adds latency and cost.
Guardrails complement rather than replace alignment: alignment changes what the model tends to produce, guardrails constrain what the application accepts and emits, and they can be updated in hours without retraining. Neither replaces least privilege on tools, since a filter can be bypassed but a permission the agent does not hold cannot be abused. OWASP's LLM Top 10 recommends input and output filtering as a layer for LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure and LLM05 Improper Output Handling, and NIST AI 600-1 treats content filtering as one of several risk-management actions for generative AI.
What to learn first
Everything this builds on, foundations first.
- Token
- →Transformer
- →Large language model (LLM)
- →Guardrails
Relationships
- Requires
- Large language model (LLM)
- Don't confuse with
- AI alignment
Sources & further reading
Standards & official texts
- NIST AI 600-1 - Artificial Intelligence Risk Management Framework, Generative AI Profile · NIST
Reference works
Where this data comes from
This entry was drafted by an AI from the sources above and has not yet been checked by a person. Treat it as a starting point, and check anything important against the sources.
See the review queueSuggest a correction on GitHubThis term as JSON
Check yourself
Loading…