Skip to content
BytePatterns

Trim a Chat to Fit the Context

MediumAI & ML#token-budget#newest-first~20m

Problem

A chat app must fit each request into the model's context window of budget tokens. The request always carries the system prompt (system tokens) and must leave reserve tokens free for the reply. turns lists the conversation's token counts oldest first; it alternates user and assistant and ends with the new user message, so its length is odd. Keep the newest user message, then keep older (user, assistant) pairs from newest to oldest while they fit, never dropping half of a pair and never leaving a gap in the history. Return how many of the oldest turns are dropped, or None if even the system prompt, the new message and the reserve do not fit.

Examples

Input:  budget = 100, system = 20, reserve = 25, turns = [10, 30, 10, 20, 15]
Output: 2
Why:    55 tokens are free: 15 for the new message and 30 for the newest pair, and the oldest pair of 40 does not fit in the 10 left
Input:  budget = 200, system = 20, reserve = 25, turns = [10, 30, 10, 20, 15]
Output: 0
Why:    the whole conversation fits
Input:  budget = 50, system = 20, reserve = 10, turns = [10, 30, 40]
Output: None
Why:    edge case, the new message alone is larger than the 20 tokens left

Hints

0 / 3

Stuck on the idea rather than the code? Context Windows covers it.