Skip to content
BytePatterns

What Is an LLM? Large Language Models Explained Simply

8 min readBytePatterns

What an LLM is: a model trained to score every possible next token, run in a loop to write text. Training, generation, why it sounds sure when wrong, with code.

A large language model answers questions, writes code and translates text, so it is tempting to picture a search engine or a database of facts behind it. There is neither. An LLM is a function with one job: given the text so far, split into tokens, assign a probability to every token that could come next. Everything else it appears to do is that one step, repeated.

The problem it solves

Human language is too irregular for hand-written rules, yet it is predictable enough that a reader can often guess the next word. A language model turns that predictability into a training signal that needs no labels: hide the next token of any sentence and ask the model to guess it. Every page of text becomes millions of free exercises.

"Large" refers to scale on both sides. The model is a neural network, almost always a transformer, with billions of adjustable numbers called parameters, trained on a corpus measured in trillions of tokens. As of October 2026, openly published models range from roughly one billion to several hundred billion parameters; those figures are approximate and move quickly.

The intuition

Think of it as three separate pieces:

  1. Tokens in. Text is cut into tokens, usually subword pieces. The vocabulary is fixed, typically tens of thousands to a few hundred thousand entries.
  2. Scores out. For the current context, the model outputs one score per vocabulary entry, turned into probabilities that sum to 1. That distribution is the model's entire output.
  3. A loop around it. Picking one token from the distribution is a separate decision: always the top one (greedy) or a random draw shaped by temperature and top-p. The chosen token is appended, and the model runs again on the longer context. A 500-token answer is 500 passes.

Training happens in stages. Pretraining adjusts the parameters so that the probability given to the actual next token, across the whole corpus, goes up. That alone produces a model that continues documents rather than answering requests. Post-training then fine-tunes it on examples of instructions and good responses, and on human or automated preferences between candidate answers, so that the most likely continuation of a question is a helpful answer. Neither stage adds a fact checker: the objective rewards text that is likely given the training data.

That is why a model can hallucinate. A plausible citation, a believable API name or a confident wrong date is exactly what a next-token objective can produce when the true continuation was rare or absent in training. The fix is outside the model: give it the right documents at question time with retrieval, or let it call tools in an agent loop.

Watch it run

The animation opens on the definition: a large language model has one job, given the tokens so far, score every possible next token. Its vocabulary has five tokens; a real one has tens of thousands. This toy model learns by counting what follows what across one nine-word text. After "the" it saw cat, mat, cat: three observations, two distinct tokens. The counts become a distribution, cat two thirds and mat one third, and every other token scores zero. The bars sum to 1, and they are the raw output, not yet an answer. Choosing from them is a separate decision; take the highest, cat. Append that choice and run exactly the same step again: after "cat" the text had sat once and ran once, an even split with no favourite. Pick one, append it, predict again; a long answer is just this loop. A real model conditions on thousands of previous tokens rather than one, but the step is identical. The last two frames are the warning: nothing in that loop checks the world. It is trained for plausible continuations, which is why a wrong answer can arrive sounding every bit as confident as a right one.

What Is an LLM

Step 1 of 12

A large language model has one job: given the tokens so far, score every possible next token.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy model: a bigram counter in place of a neural network, so every probability is an exact fraction. The interface is the real one, a context in and a score for every vocabulary entry out, with generation as a loop:

from collections import Counter, defaultdict
from fractions import Fraction

def train(words):
    """'Training' for this toy: count which token follows which."""
    counts = defaultdict(Counter)
    for a, b in zip(words, words[1:]):
        counts[a][b] += 1
    return counts, sorted(set(words))

def next_token_probs(model, context):
    """Score every vocabulary entry. This toy only looks at the last token."""
    counts, vocab = model
    seen = counts[context[-1]]
    total = sum(seen.values())
    return {tok: Fraction(seen[tok], total) if total else Fraction(0) for tok in vocab}

model = train("the cat sat on the mat the cat ran".split())
probs = next_token_probs(model, ["the"])
print({t: str(p) for t, p in probs.items() if p})   # {'cat': '2/3', 'mat': '1/3'}
print(sum(probs.values()), len(probs))              # 1 6

def generate(model, context, steps):
    out = list(context)
    for _ in range(steps):
        p = next_token_probs(model, out)
        if not any(p.values()):
            break                                   # "ran" was never followed by anything
        out.append(max(p, key=p.get))               # greedy: always the top score
    return " ".join(out)

print(generate(model, ["the"], 6))                  # the cat ran

Greedy decoding hit a tie after "cat" and took the first entry in vocabulary order, "ran", which then had no successor. Sampling instead of taking the top token reproduces the distribution, and the chain rule gives the probability of a whole continuation:

import random

def sample(model, context, rng):
    p = next_token_probs(model, context)
    live = [t for t in p if p[t]]
    return rng.choices(live, weights=[float(p[t]) for t in live])[0]

rng = random.Random(34)
draws = Counter(sample(model, ["the"], rng) for _ in range(30000))
print(round(draws["cat"] / 30000, 3), round(draws["mat"] / 30000, 3))   # 0.666 0.334

def sequence_prob(model, tokens):
    """Chain rule: a continuation's probability is a product of next-token scores."""
    p = Fraction(1)
    for i in range(1, len(tokens)):
        p *= next_token_probs(model, tokens[:i])[tokens[i]]
    return p

print(sequence_prob(model, "the cat sat on the mat".split()))   # 1/9
print(sequence_prob(model, "the mat sat".split()))              # 0

"the mat sat" gets zero because the counter never saw "mat sat". A neural model shares what it learns across similar contexts, so it gives unseen but sensible text a small nonzero probability; that generalisation is what the billions of parameters buy.

Finally, the model is checked on 400 seeded random corpora. Each distribution must equal a brute-force rescan of every adjacent pair, and all three-token continuations together must sum to exactly 1:

from itertools import product

rng = random.Random(34)
ok = True
for _ in range(400):
    body = [rng.choice("abcde") for _ in range(rng.randint(2, 30))]
    corpus = body + [body[0]]                       # every token now has a successor
    m = train(corpus)
    start = rng.choice(corpus)
    # brute force 1: rescan every adjacent pair for what followed `start`
    after = [corpus[i + 1] for i in range(len(corpus) - 1) if corpus[i] == start]
    for tok, p in next_token_probs(m, [start]).items():
        ok &= p == Fraction(after.count(tok), len(after))
    # brute force 2: all three-token continuations together sum to exactly 1
    total = sum(sequence_prob(m, [start, *rest]) for rest in product(m[1], repeat=3))
    ok &= total == 1
print(ok)                                           # True

The complexity

  • One step: one forward pass over the context, then a distribution over the whole vocabulary.
  • An answer of k tokens: k sequential steps; the KV cache keeps each one from recomputing the whole prefix.
  • Memory and context: the parameters must fit in accelerator memory, and the prompt plus the answer must fit in the context window.

Where it goes wrong

  • Treating fluency as evidence. Confidence in the wording says nothing about whether the claim is true.
  • Asking about events after training. The parameters stop changing when training ends; anything newer must come in through the prompt.
  • Counting characters. The model sees tokens, not letters, so spelling and character counts are harder than they look.
  • Expecting the same answer twice. With sampling, two runs on the same prompt can differ by design.

When it shows up in interviews

As "what is an LLM?" or "how does an LLM generate text?" in machine learning, product and general engineering rounds. The follow-ups are why models hallucinate, what temperature changes, why long prompts cost more, and when you would add retrieval instead of fine-tuning.

How to say it in an interview

"An LLM is a neural network, usually a transformer, trained to predict the next token. Given the tokens so far, it outputs a probability for every token in its vocabulary. Generation is a loop: pick a token, greedily or by sampling, append it, and run again. Pretraining on a huge text corpus teaches it to continue text; post-training on instructions and preferences makes it answer helpfully. Nothing in the objective checks facts, so it can be fluent and wrong, which is why we ground it with retrieval or tools."