Model Routing and Fallbacks
AI & ML: lesson 26 of 32
Send each request to the cheapest model that answers it well.
Lesson 26 of 32 · 5 min
Model Routing and Fallbacks
Step 1 of 9
Without a router every request goes to the large model: 1,000 of them cost $15.00 and wait 1.8 s each.
The Idea
Most traffic is easy: summaries, classifications, short lookups. A router scores each request: easy ones go to a small, cheap, fast model, hard ones to a large one. A fallback catches the misses: when the small model's confidence is low, or a call fails, the request moves up a tier instead of retrying the same one.
Real-World Example
A print shop sends business cards to the office printer and a thousand-page catalogue to the industrial press. A job the office printer chokes on goes to the press, not back into the same tray.
The Tradeoff
Every escalation pays for both models, and the router itself has to cost less than it saves. Track quality and escalation rate per tier, not just the bill. On Amazon Bedrock, intelligent prompt routing does this between two models of one family, with a fallback model as the anchor.
Your turn
Put the steps in the right order.
- Escalate low-confidence answers to the larger model
- Score how hard the incoming request is
- Send it to the cheapest tier that can handle that score
- Review cost, quality and escalation rate for each tier
Mini quiz
1 / 3