What Is Machine Learning? Rules Learned From Examples
8 min readBytePatterns
What machine learning is: instead of writing a rule, you fit one from labelled examples. Training, held-out data, overfitting and drift, with runnable code.
Most programs state a rule and let the computer apply it: if the total is over 50, shipping is free. Machine learning turns that around. You hand over examples paired with the right answers, and training finds a rule that reproduces them. The rule comes out as numbers, not code, and that one change explains both what machine learning is good at and how it fails.
The problem it solves
Some rules are short and knowable: a tax rate, a date format. Write those as code; they are exact and easy to audit. Other rules exist but nobody can write them down. What does a handwritten seven look like, across every hand on earth? Which card payments are fraud? People can label examples far more easily than they can state the rule.
That is the trade machine learning offers: labels instead of logic. Collect inputs, record the correct output for each, and let training search for parameter values that map one to the other. This setting, examples with known answers, is called supervised learning, and it is what most production systems and most interview questions mean. Two other families are worth naming: unsupervised learning finds structure in data without labels (clustering, anomaly detection), and reinforcement learning learns from rewards for actions rather than from correct answers.
The intuition
Four ideas carry the whole field:
- A model is a function with adjustable numbers. A threshold, a line, or a neural network with millions of weights: the shape is fixed, the numbers are not.
- Training is a search. Make a prediction, compare it with the label, nudge the numbers to reduce the error, repeat. The measure of error is called the loss.
- Held-out data is the only honest score. Before training, set aside a slice the model never sees. Accuracy on training data proves memory, not skill. Teams often split three ways: training, a validation set for tuning choices, and a final test set opened once.
- A fitted rule only works near its data. When live inputs drift from what the model was trained on, accuracy falls, often without any error message.
Two failure modes sit on either side. Overfitting is a model that learned the training set, noise included, and scores worse on new data. Underfitting is a model too simple to capture the pattern. Compare the training score with the held-out score to tell them apart.
Watch it run
The animation puts the two approaches side by side. Ordinary code states the rule: input goes in, the rule runs, the answer comes out, readable and exact. Then it asks for the rule for a hand-drawn seven, which nobody can write. So it flips: collect examples and label each one, 1.4 million envelopes in this illustration. A 20% slice is held back that training never sees, the only honest score. Training feeds the other 80% through and compares against the labels; when wrong, it nudges the parameters and tries again, millions of times, and that loop is the learning. The finished rule lives in 9.1 million numbers, not in lines anybody can read. Scored on the held-out slice, data it was never trained on, it reaches 94.2%. It is deployed and watched on live envelopes, at 94.0%, until a new pen style arrives and it degrades quietly to 78.5%, because a fitted rule only works near its training data. The final frame is the counterweight: when the rule is short and knowable, like a fixed tax rate, plain code stays cheaper, exact and auditable.
What Is Machine Learning
Step 1 of 12
Ordinary code states the rule, and the machine applies it to whatever comes in.
The same interactive animation as the lesson — step through it with the controls.
The code
A toy model with one parameter, so every step is visible. The world's hidden rule is "label 1 when x is above 62", with 5% of labels flipped to imitate human mistakes. Training never reads the 62; it only sees examples:
import random
CUT = 62 # the world's hidden rule; training never reads it
def make_data(rng, n, cut=CUT, noise=0.05):
"""Labelled examples: x is a measurement, y the answer a human wrote down."""
data = []
for _ in range(n):
x = round(rng.uniform(0, 100), 1)
y = int(x > cut)
if rng.random() < noise:
y = 1 - y # real labels are never perfect
data.append((x, y))
return data
def fit_threshold(train):
"""Training: one sorted sweep picks the cut with the fewest training errors."""
pts = sorted(train)
errors = sum(1 for _, y in pts if y == 0) # cut below everything: all predicted 1
best_err, best_cut = errors, pts[0][0] - 1
for i, (x, y) in enumerate(pts):
errors += 1 if y == 1 else -1 # x now falls below the cut: predicted 0
if i + 1 < len(pts) and pts[i + 1][0] == x:
continue # cannot cut between equal values
cut = x if i + 1 == len(pts) else (x + pts[i + 1][0]) / 2
if errors < best_err:
best_err, best_cut = errors, cut
return best_cut, best_err
def accuracy(cut, data):
return sum((x > cut) == y for x, y in data) / len(data)
rng = random.Random(35)
data = make_data(rng, 1000)
train, held_out = data[:800], data[800:] # hold back 20% before training starts
cut, _ = fit_threshold(train)
print(round(cut, 2)) # 62.1
print(round(accuracy(cut, train), 3), round(accuracy(cut, held_out), 3)) # 0.954 0.955
Training recovered 62.1 from examples alone, and the held-out score matches the training score: the model generalises. Neither reaches 100%, because 5% of the labels are wrong. Compare a "model" that simply memorises its training set, then let the world drift:
memo = dict(train) # a "model" that memorises its training set
recall = lambda x: memo.get(x, 0)
print(sum(recall(x) == y for x, y in train) / len(train)) # 0.97625
print(round(sum(recall(x) == y for x, y in held_out) / len(held_out), 3)) # 0.785
live = make_data(random.Random(7), 1000, cut=70) # the world drifts: the rule is now 70
print(round(accuracy(cut, live), 3)) # 0.861
new_cut, _ = fit_threshold(live[:800]) # retrain on fresh labels
print(round(new_cut, 1), round(accuracy(new_cut, live[800:]), 3)) # 70.1 0.93
The memoriser looks better on training data (it falls short of 1.0 only because a few repeated x values carry conflicting noisy labels) and collapses to 0.785 on data it has not seen: overfitting in its purest form. When the hidden rule moves to 70, the old model drops to 0.861 without raising any error, and only fresh labels fix it.
Finally, the sweep is checked against brute force on 300 seeded datasets with many ties: the best cut found in one pass must make exactly as few errors as the best of every possible cut:
ok = True
for seed in range(300):
r = random.Random(seed)
pts = [(r.randint(0, 20), r.randint(0, 1)) for _ in range(r.randint(1, 25))]
got_cut, got_err = fit_threshold(pts)
errs = lambda c: sum((x > c) != y for x, y in pts)
candidates = [min(x for x, _ in pts) - 1] + [x + 0.5 for x, _ in pts]
ok &= got_err == min(errs(c) for c in candidates) == errs(got_cut)
print(ok) # True
The complexity
- This toy: training is one sort plus one sweep,
O(n log n); a prediction is one comparison. - Real models: training cost grows with examples × parameters × passes over the data; one prediction stays a fixed amount of arithmetic.
- The hidden cost: labelled data, often slower and dearer to collect than the training itself.
Where it goes wrong
- Scoring on training data. It measures memory. Hold data back first, and never tune against the test set.
- Leakage. A feature that secretly contains the answer, such as a refund flag when predicting fraud, gives brilliant offline scores.
- Silent drift. Inputs change and accuracy falls with no exception thrown. Monitor live accuracy or proxies for it.
- Using it where a rule exists. If the rule fits on one line and is exact, code beats a model on cost, correctness and auditability.
When it shows up in interviews
As an opener in machine learning and AI engineering rounds ("explain machine learning to a new engineer"), and as the first five minutes of any ML system design question. Expect follow-ups on train/validation/test splits, overfitting, how you would know a deployed model is degrading, and when you would not use ML at all. It also underpins large language models, which as of October 2026 are pretrained self-supervised: the label is simply the next token of real text.
How to say it in an interview
"Ordinary code states a rule. Machine learning fits a rule from labelled examples: a model is a function with adjustable parameters, and training repeatedly compares predictions with labels and nudges the parameters to reduce a loss. I hold out data the model never trains on, because only that score tells me it generalises rather than memorised. The fitted rule works near its training distribution, so in production I monitor for drift and retrain on fresh labels. And if the rule is short and knowable, I write the code instead."