Guardrails
AI & ML: lesson 22 of 32
Checks around the model, because the model is not the boundary.
Lesson 22 of 32 · 5 min
Guardrails
Step 1 of 10
The model sits in the middle of three checks, and its own instructions are the weakest of the four layers.
The Idea
A guardrail is a check outside the model. Input filters catch prompt injection and unsupported requests. Tool permissions decide what the loop can actually reach. Output checks scan for secrets and unsafe content. The model's own instructions are the weakest layer, because text can argue with text.
Real-World Example
A bank teller is polite and well trained, and still cannot move a large sum alone. The limit lives in the system, not in their judgment, precisely so that a convincing story is not enough.
The Tradeoff
Every layer adds latency, cost and false positives, and an over-eager filter that blocks real work gets switched off by the first frustrated team. Put the hard limits where they cannot be talked around — scopes, quotas, allow-lists — and keep the soft checks for things a human will review anyway.
Your turn
Put the steps in the right order.
- Run the model, which may propose a tool call
- Screen the incoming request and the retrieved context
- Scan the finished answer before it reaches the user
- Check the call against the credentials and scopes it is allowed
Mini quiz
1 / 3