Skip to content
BytePatterns

The Tool-Use Loop: How Function Calling Works

8 min readBytePatterns

How function calling works as a loop: the model proposes a call, your code validates and runs it, errors return as observations and a step budget ends the run.

"Function calling" sounds like the model calls a function. It never does. The model writes a request: a tool name and some arguments. Your program decides whether that request is valid, runs it, and writes the result back into the conversation, and then the model is asked again. That round trip, repeated until the model answers or a budget runs out, is the tool-use loop. Below, it is built message by message.

The problem it solves

Many tasks need several steps whose inputs depend on earlier outputs: look up the customer, fetch their orders, compare two totals. The loop turns one-shot generation into a sequence the model can steer, while every side effect stays in code you control.

What agents are, why permissions belong in code, and how injected instructions in tool output are handled are covered in LLM agents explained. Here the subject is the mechanics of the loop itself, the part you write by hand when you build one.

The intuition

Think in turns. Each model turn produces one assistant message, and that message either contains text, which ends the run, or one or more tool calls. As of October 2026, the major model APIs describe each tool's parameters with a JSON Schema and return each call as a structured object with its own id (from memory). Then, for every call in the turn:

  1. Validate the name and arguments against the schema. The model can name a tool that does not exist, omit a required field, or send "1" where an integer belongs.
  2. Run it only if validation passed.
  3. Append exactly one result, tagged with that call's id, whether it is data or an error message.

Three rules follow from that:

  • An error is an observation. The text "limit must be int" is the only evidence the model gets that its call was wrong. Feed it back, do not throw it away and do not crash.
  • Every call gets a result. A turn can contain several calls; leave one unanswered and the transcript no longer matches what the model asked for.
  • The budget counts model turns. Each turn is a full model call: latency, tokens and money. A loop that waits for the model to declare itself done can run forever, so cap it.

Watch it run

An agent is a loop, not a model: the goal goes in once, and the loop is what does the work, with five steps of budget. The model emits a tool name and arguments, search(q), and has not run anything; four steps are left. Your code checks the arguments against a schema before anything executes, and the badge says ok. The tool runs, and its result, three rows, is appended to the transcript as an observation. The model reads the whole transcript again and decides the next move, and three steps are left.

This call is malformed, so it is refused, and the refusal is itself an observation: the transcript now holds two. Because the model can see why, the next attempt is different rather than identical, and that retry runs. With enough evidence, it stops calling tools and answers, with one step to spare. The last frame shows the other ending: a model that keeps retrying the same failing call simply never stops on its own. The budget is what halts it, at zero, with five observations that taught it nothing.

The Tool-Use Loop

Step 1 of 9

An agent is a loop, not a model. The goal goes in once; the loop is what does the work.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy model: the "model" is a scripted Python function that replays a fixed plan, so the loop can be run and tested exactly. The plan mirrors the animation: a good call, a malformed one, a corrected retry, then an answer.

import json
import random

ORDERS = {"ada": [("A1", 30), ("A2", 12), ("A3", 55)], "lin": [("L1", 8)]}
ran = []                                            # every tool that really executed

def search_orders(customer, limit=10):
    ran.append(("search_orders", customer, limit))
    return {"rows": ORDERS.get(customer, [])[:limit]}

TOOLS = {"search_orders": (search_orders, {"customer": str}, {"limit": int})}

def validate(name, args):
    """Your code, not the model: an error message, or None if the call may run."""
    if name not in TOOLS:
        return "unknown tool " + name
    _, required, optional = TOOLS[name]
    for key in required:
        if key not in args:
            return "missing argument " + key
    for key, value in args.items():
        want = required.get(key) or optional.get(key)
        if want is None:
            return "unexpected argument " + key
        if type(value) is not want:                 # type(), so True is not an int
            return "%s must be %s, got %r" % (key, want.__name__, value)
    return None

def run(model, goal, budget=5):
    transcript = [{"role": "user", "content": goal}]
    for turn in range(1, budget + 1):               # one model turn = one step
        reply = model(transcript)
        transcript.append(reply)
        if not reply.get("calls"):
            return reply["text"], turn, transcript
        for call in reply["calls"]:                 # every call gets exactly one result
            error = validate(call["name"], call["args"])
            result = {"error": error} if error else TOOLS[call["name"]][0](**call["args"])
            transcript.append({"role": "tool", "id": call["id"], "content": json.dumps(result)})
    return "budget spent", budget, transcript

def scripted(plan):
    """A toy model: replays a fixed list of turns, then answers."""
    def model(transcript):
        turn = sum(m["role"] == "assistant" for m in transcript)
        if turn < len(plan):
            return {"role": "assistant", "calls": plan[turn]}
        return {"role": "assistant", "text": "done after %d tool turns" % turn}
    return model

plan = [
    [{"id": "c1", "name": "search_orders", "args": {"customer": "ada", "limit": 3}}],
    [{"id": "c2", "name": "search_orders", "args": {"customer": "lin", "limit": "1"}}],
    [{"id": "c3", "name": "search_orders", "args": {"customer": "lin", "limit": 1}}],
]
text, turns, transcript = run(scripted(plan), "compare ada and lin")
print(text, turns)                       # done after 3 tool turns 4
for m in transcript:
    if m["role"] == "tool":
        print(m["id"], m["content"])
# c1 {"rows": [["A1", 30], ["A2", 12], ["A3", 55]]}
# c2 {"error": "limit must be int, got '1'"}
# c3 {"rows": [["L1", 8]]}
print(len(ran))                          # 2

Three calls, two executions: c2 never reached the tool, but it still got a result with its id. Four turns spent out of five. Next, the stubborn model, and a cheap guard that ends the run the moment a call that already failed validation is proposed again:

def stubborn(transcript):
    n = sum(m["role"] == "assistant" for m in transcript)
    call = {"id": "s%d" % n, "name": "search_orders", "args": {"customer": "ada", "limit": "3"}}
    return {"role": "assistant", "calls": [call]}

print(run(stubborn, "orders for ada")[:2])       # ('budget spent', 5)

def stop_on_repeat(model):
    """Wrap a model: the same failing call twice ends the run early."""
    failed = set()
    def watched(transcript):
        reply = model(transcript)
        for call in reply.get("calls", []):
            key = (call["name"], json.dumps(call["args"], sort_keys=True))
            if key in failed:
                return {"role": "assistant", "text": "stopped: repeated a failing call"}
            if validate(call["name"], call["args"]):
                failed.add(key)
        return reply
    return watched

print(run(stop_on_repeat(stubborn), "orders for ada")[:2])
# ('stopped: repeated a failing call', 2)

The seeded check: 2,000 random plans of one to three calls per turn, mixing good calls, wrong types (including True for an integer), missing and unexpected arguments and an unknown tool, under random budgets. Turns, answers, result ids and executions are compared against the plan read directly, with the validation rules written out a second time:

POOL = [("search_orders", {"customer": "ada"}), ("search_orders", {"customer": "lin", "limit": 1}),
        ("search_orders", {"customer": "ada", "limit": True}), ("search_orders", {"limit": 2}),
        ("search_orders", {"customer": 7}), ("search_orders", {"customer": "ada", "page": 2}),
        ("delete_orders", {"customer": "ada"})]

def valid_by_hand(name, args):
    """Brute force: the rules written out again, independently of validate()."""
    return (name == "search_orders" and isinstance(args.get("customer"), str)
            and set(args) <= {"customer", "limit"}
            and ("limit" not in args or type(args["limit"]) is int))

rng = random.Random(39)
ok = True
for _ in range(2000):
    plan = [[dict(zip(("name", "args"), rng.choice(POOL)), id="t%dc%d" % (t, c))
             for c in range(rng.randint(1, 3))] for t in range(rng.randint(0, 7))]
    budget = rng.randint(1, 6)
    ran.clear()
    text, turns, transcript = run(scripted(plan), "goal", budget)
    used = plan[:budget]
    ok &= turns == min(len(plan) + 1, budget)
    ok &= text.startswith("done") == (len(plan) < budget)
    ok &= [m["id"] for m in transcript if m["role"] == "tool"] == [c["id"] for t in used for c in t]
    ok &= len(ran) == sum(valid_by_hand(c["name"], c["args"]) for t in used for c in t)
print(ok)                                         # True

The complexity

  • Model calls: at most budget. Tool executions: at most the number of valid calls across those turns.
  • Tokens: each turn resends the whole transcript, so total input tokens grow roughly with the square of the number of turns unless old results are trimmed: the context window problem again.

Where it goes wrong

  • Dropping errors. Swallow a validation failure and the model retries blind, often with the identical call.
  • Raising instead of returning. An exception in your loop ends the run; an error result lets the model fix itself.
  • An unanswered call. Run the first of two parallel calls and forget the second, and the transcript is malformed.
  • isinstance for integers. True is an int in Python, so a boolean slips through; compare type() or use a real schema validator.
  • Counting tool calls instead of turns. A turn with three calls is one model call; budget the expensive thing.
  • No repeat guard. A budget eventually stops a stuck model; a repeat check stops it at the second identical failure.

When it shows up in interviews

In AI engineering and system design rounds: "walk me through a function call", "how do you stop an agent looping", "what if the model passes bad arguments". It also appears as a coding exercise: implement the loop with a fake model, as above.

How to say it in an interview

"The model never runs anything. Each turn it returns either text, which ends the loop, or tool calls with ids. My code validates each call against its schema, runs the valid ones, and appends exactly one result per call id, using the error message as the result when validation fails, so the model can correct itself. I cap the number of model turns, because that is where cost and latency live, and I stop early if it repeats a call that already failed."