Quick recap from Blog 5: Learning Automata could adapt to what was happening right now, but they were hopeless at delayed consequences. A chess move that loses you the game 20 turns later. A word that quietly breaks a sentence three clauses down. They had no way to connect "what I did" with "what happened much later."
This post is about the two people who went after that gap, and the third person who spotted a problem hiding behind it. By the end you'll know what an eligibility trace is, how backpropagation actually works (with a tiny example), and why every modern AI system still runs on ideas from this decade.
First, the problem in plain English
Imagine you're baking cookies for the first time. At step 2 you accidentally dump in a tablespoon of salt instead of a teaspoon. You keep going: mix, chill, shape, bake. At step 50 (taste time) the cookies are awful.
Now ask yourself: which step was the mistake?
You didn't taste anything at step 2. The feedback only showed up at the very end. To learn, you need a way to trace the bad result back to the step that caused it. That's the credit assignment problem (or "blame assignment", if you prefer): when something good or bad happens, which of your earlier decisions deserves the credit or the blame?
By the early 1970s the field had two things that didn't talk to each other:
- Bellman's equations (Blog 3): mathematically perfect and provably optimal, but they need a complete model of the world. Nobody has that.
- Samuel's checkers player and Tsetlin's automata (Blogs 4 and 5): practical and adaptive, but with no theory behind them and no good answer for late-arriving consequences.
The 1970s produced three ideas that started to close that gap. Here's the roadmap so you always know where we are:
- Klopf gives a single neuron a drive and a short memory of what it just did. (Who gets credit, and when?)
- Werbos gives a whole network a direction: a method to push blame backward through many layers. (How far back does the blame go?)
- Grossberg adds a warning: learn too eagerly and you'll erase what you already knew. (How do we keep our old skills?)
Then we'll see how Adaptive Dynamic Programming wired them together, and how it produced the Actor-Critic idea that still runs inside modern AI.
Quick glossary
You'll see these words a lot. Don't memorize them, just glance:
- Neuron (artificial): a tiny calculator. It takes numbers in, multiplies each by a weight, adds them up, and outputs a result.
- Weight: how much the neuron "listens" to one particular input. Learning = adjusting weights.
- Reward: a number telling the system "that went well" (positive) or "that went badly" (negative).
- Expectation / baseline: the reward the system predicted it would get.
- Layer: a group of neurons. A network stacks layers: input -> hidden layers -> output.
- Loss: a single number measuring how wrong the output was. Lower is better.
- Gradient: a list of "if I nudge this weight a little, how much does the loss change?"
Harry Klopf and the Hedonistic Neuron (1972)
Most of the 1970s AI crowd was busy building symbolic systems: expert systems, logic provers, decision trees. Harry Klopf, a researcher at the US Air Force Cambridge Research Laboratories, was looking at a single neuron and asking something stranger.
What if the neuron isn't a passive wire? What if it wants something?
In his 1972 report, Brain Function and Adaptive Systems: A Heterostatic Theory, Klopf proposed that neurons act like tiny hedonists: pleasure-seekers that strengthen the connections that led to good outcomes and weaken the ones that didn't.
That was a direct challenge to the textbook picture of a neuron as a threshold switch: inputs in, weighted sum, fire or don't. In that picture the neuron has no memory of outcomes and no preference about them. Klopf said that couldn't be the whole story.
Heterostasis vs. homeostasis (the thermostat vs. the athlete)
Biology textbooks of the time said living systems seek homeostasis: balance and stability. Think of a thermostat. It wants the room at 21°C. Hit 21 and it relaxes.
Klopf proposed heterostasis: a drive to push beyond the comfortable state toward something better. Think of an athlete. Hitting last month's personal best isn't "done", it's the new starting line.
Why does this matter for AI? Because it reframes learning as forward-looking. A thermostat-style learner stops when things are fine. A heterostat-style learner keeps asking "can I do better than I expected?" And that is exactly what reinforcement learning is.
The timing problem, and the eligibility trace
Back to the cookies. Klopf hit the same wall: rewards arrive after the action that caused them. When the reward finally shows up, how does the neuron know which of its recent activity deserves credit?
His answer: the eligibility trace. Think of it like a glowing footprint. When a connection is active, it leaves a footprint that fades over time. If a reward arrives while the footprint is still glowing, that connection is "eligible" for credit. If the footprint has faded, it isn't.
Here's a simplified teaching version of the idea as an update rule (Klopf's own later formulations got more elaborate, but this captures the heart of it):
Let's read it piece by piece:
-
Δw_i: how much to change the weight on inputi(the "learning step"). -
α(alpha): the learning rate, i.e., how big a step to take. -
r(t): the reward you actually got. -
r̄(t): the reward you expected to get. -
x_i(t − τ): the footprint, i.e., how active inputiwas a moment ago.
The middle piece, r − r̄, is the important one. It's surprise:
- Got more than expected (
r > r̄)? Strengthen whatever was recently active. - Got less than expected? Weaken it.
- Got exactly what you expected? Change nothing. Nothing to learn.
Example: you order your usual coffee and it's good. No surprise, no learning. You try a new cafe, and the coffee is fantastic. Big positive surprise, and your brain strongly reinforces "go to that cafe." Learning is driven by the gap, not by the raw reward.
Hold onto this "surprise" idea. In Blog 8, Richard Sutton turns it into the TD error, which is arguably the most important quantity in all of RL. Klopf had the biological intuition first; the math came later.
And the footprint idea didn't die either: modern algorithms called TD(λ) still use eligibility traces in essentially this form.
Pause and check. Can you say, in one sentence, why a neuron needs both a surprise signal and an eligibility trace? (Surprise tells it how much to change. The trace tells it what to change.)
Where we are: Klopf gave a neuron a drive and a short memory. But a single neuron is tiny. What happens when you have layers of them, and the reward only connects to the last layer? That's the next problem.
Paul Werbos and the Algorithm Nobody Read (1974)
A neuron that wants reward but can't figure out which of its decisions caused it will thrash forever. And this gets brutal in multi-layer networks. If the final answer is wrong, which weight, in which layer, is responsible? In 1974, nobody had a practical answer for neural networks.
Paul Werbos did. Getting anyone to pay attention was the hard part.
In his 1974 Harvard PhD thesis, Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences, Werbos described how to train multi-layer neural networks by propagating error backward through the layers. Today we call this backpropagation (often just "backprop").
Backprop without the scary math: the assembly line
Picture a three-station assembly line making a gadget. Station 1 passes its work to Station 2, then Station 3, then the finished gadget gets inspected. The inspector says: "This is off by 6 millimeters."
Whose fault is that? The inspector can only see the final product. But here's the trick:
- Station 3 asks: "How much did my output contribute to that error?" It can answer directly, because it's the last one.
- Station 3 tells Station 2: "Here's how sensitive the final result is to what you gave me."
- Station 2 uses that to work out its own share of blame, then tells Station 1.
Blame flows backward, station by station, each one using what the next one told it. That's backpropagation.
The chain rule, with gears
The math behind this is the chain rule from calculus, and it's much friendlier than its name. Here's the idea with gears.
Say turning Knob A by 1 unit moves Knob B by 3 units. And turning Knob B by 1 unit moves the final output by 2 units. Then turning Knob A by 1 unit moves the output by 3 × 2 = 6 units.
That's the whole chain rule: when effects are chained together, you multiply the sensitivities.
Translating the symbols:
-
L: the loss (how wrong the final output was). -
a: a neuron's output (its "activation"). -
z: the neuron's input sum before it produces its output. -
w_i: the weight we want to adjust.
Read the equation left to right as a story: "How much does the loss change if I nudge this weight?" = "how much the loss changes with the neuron's output" × "how much the output changes with its input sum" × "how much the input sum changes with the weight." Three gear ratios multiplied together. Do this for every weight in every layer and you get the gradient: a complete to-do list of nudges that reduce the error.
For the first time, a multi-layer network had a principled way to learn: not just the output layer, but every layer, including the ones far from the reward.
Werbos's bigger vision: backprop as a tool for Bellman
Here's the part that matters for this series. Werbos didn't think of backprop as just "neural network training." He saw it as a way to make Bellman's equations usable on real problems. Remember from Blog 3: Bellman's math is perfect but chokes on large problems because you'd need a giant table with an entry for every possible situation. Werbos's idea: don't store the table, approximate it with a neural network and train that network with backprop. He called this framework Adaptive Dynamic Programming (ADP).
That exact idea (a neural network standing in for a value table) is what DeepMind used in DQN in 2013, thirty-nine years after Werbos's thesis (Blog 11). He saw the destination early; the hardware and data took decades to catch up.
Pause and check. Backprop answers "how much is each weight to blame?" ADP asks "what should we be using backprop on?" The answer: an estimate of how good each situation is. That estimate has a name, and it's next.
Adaptive Dynamic Programming: wiring it all together
We need one new idea before this makes sense: the value function.
A value is a number answering "how good is this situation in the long run?", not just "how good does it feel right now?"
Back to cookies: imagine you've already added the salt disaster and the dough is mixed. Nothing bad has happened yet. No reward, no punishment. But that dough is already doomed. A good value function would score it low immediately, before anyone tastes it. That's the power of values: they let you judge a situation before the consequences show up.
- A high-value state is one where, even with no immediate reward, the future looks rich.
- A low-value state is a trap: maybe it feels fine now, but it funnels you toward failure.
The Bellman equation says how values relate to each other:
In words: the value of a situation = the best action's (immediate reward + discounted value of wherever you land next).
-
V(s): value of situation (state)s. -
a: an action you could take.maxmeans "pick the best one." -
R(s,a): the immediate reward. -
γ(gamma): a discount between 0 and 1. A reward tomorrow is worth a bit less than a reward today. -
P(s' | s, a): the probability of landing in the next states', andV(s')is its value.
The catch, as always: you can't write a table for every state of a real problem. This is the curse of dimensionality: the number of states explodes as problems grow. ADP's workaround is to use a neural network V(s; weights) as a stand-in for the table, and train it with backprop. The curse isn't gone, but it stops being a wall and becomes a negotiation.
The Actor-Critic idea
Out of ADP came one of the cleanest architectural ideas in all of RL, and it's still everywhere: Actor-Critic.
- The Actor decides what to do. It tries things, explores, proposes moves. This is Klopf's trial-and-error learner.
- The Critic judges how good the situation is, using a value function trained with backprop. This is Werbos's contribution.
An everyday analogy: a student (Actor) writes essays, and a coach (Critic) reads them. The coach doesn't write anything but learns to predict how good an essay will be. The student improves by listening to the coach's feedback, not just waiting for a final grade at the end of term. Meanwhile the coach also improves by comparing predictions to actual results. Both learn at once.
And notice how the pieces line up: the Critic's prediction is the "expectation" in Klopf's update rule. When reality beats the Critic's prediction, that surprise tells the Actor "do more of that."
A quick date check so nothing here gets overstated. ADP was being developed in the late 1970s (Werbos's early forms of it date to around 1977). The actor-critic architecture as a working, tested system came in 1983, when Barto, Sutton and Anderson used it to learn pole balancing, which is the subject of Blog 7.
The big conceptual change: before this, learners mostly reacted. With a value function they could anticipate. A system balancing a pole can learn that a certain tilt is "bad news" before it falls. That's the difference between reflex and planning.
Stephen Grossberg: the warning label (1976)
We now have a learner with drive, direction, and anticipation. Stephen Grossberg raised a problem that sounds almost boring but turned out to be one of the hardest in the field: learning can destroy what you already know.
Train a network hard on chess, then train it on Go, and it may forget chess, not gradually but catastrophically. New updates overwrite old structure. (The name catastrophic forgetting was attached to this in the late 1980s, but Grossberg was already wrestling with the underlying issue in the 1970s.)
An everyday version: you spend a year learning Spanish, then cram Italian for two months, and suddenly your Spanish is full of Italian words. Your brain used the same "shelf space" for both.
Grossberg's Adaptive Resonance Theory (ART), from 1976, framed this as a tug-of-war between two forces:
- Plasticity: the ability to learn new things.
- Stability: the ability to keep what you already know.
Too much plasticity and you forget everything. Too much stability and you can never learn anything new. ART's answer: when a new input resembles something you already know, the two "resonate" and the memory is updated. When it's very different, don't overwrite, create a new category instead.
This tension never went away. Continual learning, fine-tuning without forgetting, and knowledge distillation are all modern attempts at the problem Grossberg named in the 1970s.
Putting it together: desire, direction, memory
Step back. By 1980 three ideas were on the table:
| Person | Gave learning systems | The key idea in one line |
|---|---|---|
| Klopf | A drive | Learn from the gap between expected and actual reward; use a fading trace to know what to credit |
| Werbos | A direction | Send blame backward through layers with the chain rule; use it to train value estimates |
| Grossberg | A memory | Balance new learning against keeping old knowledge |
Learning Automata had none of the three. And remember our cookie problem? Klopf's trace says which recent actions are eligible. Werbos's backprop says how to push the blame back through the whole recipe. The value function says this dough is doomed before anyone tastes it.
The scale was tiny by modern standards: a handful of neurons, small simulations, and a thesis that sat mostly unread. But the logic was complete.
Why this still matters today
Here's a connection that makes this history feel less like a museum visit. Modern RLHF (reinforcement learning from human feedback), the process used to tune chatbots like the one you may be chatting with, follows the same blueprint:
- A reward signal encodes what people prefer (helpful, correct, safe). That's Klopf's pleasure.
- When an output beats expectations, the pathways behind it get reinforced, and the blame or credit travels back through billions of weights via backprop. That's Werbos.
- The algorithm commonly used for this, PPO, is an actor-critic method: one part proposes, another part estimates value. That's the architecture from this post.
- Keeping the model from forgetting how to do everything else while it's being tuned? That's Grossberg's problem.
The math is more sophisticated and the networks are enormous. But the skeleton hasn't changed. What evolved over five decades wasn't the idea. It was the scale.
What's next
The pieces are on the table, but they're not assembled yet:
- Klopf's biological intuitions need a rigorous mathematical home.
- Backprop needs to be tied explicitly to temporal credit assignment (credit across time).
- Actor-Critic needs someone to build it, test it, and write it down properly.
In 1981, Richard Sutton and Andrew Barto began doing exactly that, starting from a model of classical conditioning (think Pavlov's dog). Within a few years they'd produce the first rigorous treatment of temporal credit assignment in RL, including that pole-balancing actor-critic. That's when the field stopped being a collection of brilliant isolated ideas and started becoming a coherent science.
Next: Blog 7, The Sutton and Barto Foundation: Formalizing Temporal Credit Assignment (1981–1984).













Top comments (0)