Skip to content
BytePatterns

Guardrails

AI & ML: lesson 22 of 32

Checks around the model, because the model is not the boundary.

Lesson 22 of 32 · 5 min

Guardrails

Step 1 of 10

The model sits in the middle of three checks, and its own instructions are the weakest of the four layers.

The Idea

A guardrail is a check outside the model. Input filters catch prompt injection and unsupported requests. Tool permissions decide what the loop can actually reach. Output checks scan for secrets and unsafe content. The model's own instructions are the weakest layer, because text can argue with text.

Real-World Example

A bank teller is polite and well trained, and still cannot move a large sum alone. The limit lives in the system, not in their judgment, precisely so that a convincing story is not enough.

The Tradeoff

Every layer adds latency, cost and false positives, and an over-eager filter that blocks real work gets switched off by the first frustrated team. Put the hard limits where they cannot be talked around — scopes, quotas, allow-lists — and keep the soft checks for things a human will review anyway.

Your turn

Put the steps in the right order.

  1. Run the model, which may propose a tool call
  2. Screen the incoming request and the retrieved context
  3. Scan the finished answer before it reaches the user
  4. Check the call against the credentials and scopes it is allowed

Mini quiz

1 / 3

Instructions in retrieved documents should be treated as:

New lessons land every few weeks

Leave an address and we will tell you when the next one is up. That is the only reason we will use it.

One address, stored so we can email you. Nothing else, ever.