Berk Bayri

Guardrails

Controls placed around an AI system — input and output checks, permission limits, review steps and escalation rules — that keep its behaviour within acceptable bounds.

Guardrails are the controls that keep an AI system's behaviour within acceptable limits. They can sit before the model (screening inputs), after it (checking outputs for policy, accuracy or sensitive data), or around its actions (restricting which tools, data and operations it can use, and when a human must approve).

The word is often used as reassurance: "we have guardrails". A more useful habit is to make them specific and testable: what exactly is blocked, what is reviewed, who is alerted, and how often does each control fire or fail?

What good guardrails include

  • Permission boundaries that reflect that capability is not authority
  • Review and approval steps where a person's judgment is genuinely needed (human in the loop)
  • Checks against known failure modes such as hallucination
  • A route for uncertain cases to be handed back
  • A plan for failure: how it is detected, contained and reversed, which is the substance of a recovery contract

Guardrails reduce risk; they do not remove it. Ask what happens when one fails.

Read more in OpenAI Decisions API turns probability into policy.