Skip to content
BytePatterns

Model Routing and Fallbacks

AI & ML: lesson 26 of 32

Send each request to the cheapest model that answers it well.

Lesson 26 of 32 · 5 min

Model Routing and Fallbacks

Step 1 of 9

Without a router every request goes to the large model: 1,000 of them cost $15.00 and wait 1.8 s each.

The Idea

Most traffic is easy: summaries, classifications, short lookups. A router scores each request: easy ones go to a small, cheap, fast model, hard ones to a large one. A fallback catches the misses: when the small model's confidence is low, or a call fails, the request moves up a tier instead of retrying the same one.

Real-World Example

A print shop sends business cards to the office printer and a thousand-page catalogue to the industrial press. A job the office printer chokes on goes to the press, not back into the same tray.

The Tradeoff

Every escalation pays for both models, and the router itself has to cost less than it saves. Track quality and escalation rate per tier, not just the bill. On Amazon Bedrock, intelligent prompt routing does this between two models of one family, with a fallback model as the anchor.

Your turn

Put the steps in the right order.

  1. Escalate low-confidence answers to the larger model
  2. Score how hard the incoming request is
  3. Send it to the cheapest tier that can handle that score
  4. Review cost, quality and escalation rate for each tier

Mini quiz

1 / 3

A router sends 700 of 1,000 requests to a small model. For the router to pay off, each routing decision must cost:

New lessons land every few weeks

Leave an address and we will tell you when the next one is up. That is the only reason we will use it.

One address, stored so we can email you. Nothing else, ever.