Cognitive Connie
What is reward-based learning?
Reward-based learning addresses one of the most fundamental questions in psychology: how do the consequences of our actions shape what we do next? Edward Thorndike's Law of Effect (1898) provided the first systematic empirical answer. Working with cats in puzzle boxes, Thorndike observed that responses followed by a "satisfying state of affairs" were more likely to be repeated, while responses followed by an "annoying state of affairs" became less likely. This deceptively simple principle — that outcomes feed back to strengthen or weaken the behaviours that produced them — became the cornerstone of an entire research tradition in learning psychology.
01. Overview
From predictable to unpredictable reward: how schedule type shapes behaviour.
Skinner and Ferster's (1957) systematic mapping of reinforcement schedules showed that the predictability of reward is as important as its magnitude. Moving from fixed to variable schedules dramatically increases both response rates and resistance to extinction.
Fixed Interval (FI)
Reinforcement is delivered for the first response after a fixed period of time has elapsed (e.g., every 60 seconds). Produces a characteristic 'scallop' pattern: responding drops immediately after reinforcement and accelerates as the interval end approaches. The organism is essentially tracking time rather than responding continuously.
Fixed Ratio (FR)
Reinforcement is delivered after a fixed number of responses (e.g., every 10th press). Produces a high, steady rate of responding with a brief post-reinforcement pause. The ratio requirement creates a predictable relationship between effort and reward, making pause-and-run patterns typical.
Variable Interval (VI)
Reinforcement is delivered for the first response after an unpredictable interval (e.g., on average every 60 seconds, but variable). Because the organism cannot predict when the interval will expire, it maintains a low but steady rate of responding throughout — and is moderately resistant to extinction.
Variable Ratio (VR)
Reinforcement is delivered after an unpredictable number of responses (e.g., on average every 10th press, but variable). Produces the highest response rates of any schedule and is the most resistant to extinction — the organism cannot discriminate periods of non-reinforcement from the variable training history. Slot machines operate on this schedule.
Key figures
Edward Thorndike
1874–1949American psychologist whose puzzle-box experiments with cats produced the Law of Effect (1898, 1911) — the first empirical demonstration that reward strengthens behaviour and the direct conceptual precursor to operant conditioning. Thorndike's connectionist account of learning as the stamping in and stamping out of stimulus-response bonds shaped the behaviourist research tradition that followed.
B.F. Skinner
1904–1990American behaviourist who developed the operant conditioning chamber and, with Ferster (1957), systematically mapped the effects of reinforcement schedules on behaviour. Skinner demonstrated that the pattern of reinforcement — fixed versus variable, ratio versus interval — determines response rate, post-reinforcement pauses, and resistance to extinction. His analysis of schedules remains the most comprehensive empirical account of how reward structure shapes learned behaviour.
Robert Rescorla & Allan Wagner
Rescorla: 1940–2020 / Wagner: 1934–presentAmerican psychologists who jointly proposed the Rescorla-Wagner model (1972) — a formal mathematical account of Pavlovian conditioning in which learning is driven by the discrepancy between expected and obtained outcomes. The model correctly predicted blocking (Kamin, 1969) and overexpectation, and introduced the concept of prediction error as the fundamental learning signal, a formulation that became central to subsequent learning theory.
George Ainslie
1944–presentAmerican psychiatrist and psychologist whose work on impulsive choice and self-control established hyperbolic discounting as the characteristic form of temporal reward devaluation in humans and other animals. His 1975 Psychological Bulletin paper proposed that hyperbolic discounting produces preference reversals — choosing the smaller, sooner reward when it is close, but the larger, later reward when both options are distant — a pattern incompatible with rational exponential discounting and with direct relevance to addiction and self-control research.
Key concepts
Law of Effect
Thorndike's foundational principle (1898): responses followed by a satisfying state of affairs are strengthened — the stimulus-response bond is 'stamped in' — while responses followed by an annoying or neutral state of affairs are weakened and 'stamped out'. Originally derived from animal experiments with puzzle boxes, the Law of Effect established that the consequences of behaviour feed back to alter the probability of that behaviour recurring. It was the direct conceptual precursor to Skinner's operant conditioning.
Reinforcement
Any consequence that increases the future probability of the behaviour it follows. Positive reinforcement adds a desirable stimulus after a response (e.g., praise following correct recall). Negative reinforcement removes an aversive stimulus after a response (e.g., relief of discomfort following avoidance behaviour). Both increase responding — negative reinforcement is frequently confused with punishment, but they are opposite in effect: punishment decreases responding, reinforcement increases it.
Reward prediction error
The discrepancy between the reward an organism expects and the reward it actually receives. The Rescorla-Wagner model (1972) formalised this as the primary driver of learning: the change in associative strength (ΔV) on any trial is proportional to the difference between what was obtained (λ) and what was predicted (V). When reward exceeds prediction, the association strengthens; when reward falls short of prediction, it weakens; when reward exactly matches prediction, no learning occurs. The model correctly predicted blocking — the finding that a fully predicted outcome cannot be used to condition a new stimulus.
Delay discounting
The psychological devaluation of rewards as a function of their temporal distance from the present moment. Psychologically, discounting is hyperbolic rather than exponential: the subjective value of a reward drops steeply for short delays and more gradually for longer ones, creating preference reversals. A person may prefer £100 now over £110 next week, yet prefer £110 in 53 weeks over £100 in 52 weeks — a logically inconsistent pattern explained by hyperbolic discounting. This feature of reward psychology has direct implications for understanding impulsive choice, addiction, and self-control failure.
Partial reinforcement extinction effect (PREE)
The counterintuitive finding that behaviours reinforced on partial schedules are more resistant to extinction than behaviours reinforced on every trial. Jenkins and Stanley (1950) established this as one of the most robust phenomena in learning psychology. The explanation most consistently supported is the discrimination hypothesis: under continuous reinforcement, the organism quickly discriminates the extinction condition (no reward) from training; under partial reinforcement, non-reward during training is indistinguishable from extinction, so the organism persists far longer before extinguishing.
Secondary reinforcement
A stimulus that acquires reinforcing properties not through its intrinsic biological value but through repeated pairing with a primary reinforcer (food, water, warmth, pain relief). Money is the most pervasive secondary reinforcer in human behaviour — valuable only because of its consistent association with primary rewards. In clinical psychology, token economy programmes exploit secondary reinforcement by awarding tokens for target behaviours; the tokens can later be exchanged for primary or otherwise valued reinforcers, allowing contingencies to be applied in settings where direct primary reinforcement is impractical.
Test your knowledge
Frequently asked questions
What is Thorndike's Law of Effect?+
The Law of Effect, proposed by Edward Thorndike (1898), states that behaviours followed by satisfying consequences become more likely to be repeated, while behaviours followed by unsatisfying or aversive consequences become less likely. Thorndike arrived at this principle from experiments with cats placed in puzzle boxes: the cats initially made random movements, but those that accidentally triggered the escape mechanism were rewarded with food; over trials, they learned to perform the effective response more quickly. The Law of Effect was the first systematic empirical account of how reward shapes behaviour and became the conceptual foundation of operant conditioning.
Why do variable reinforcement schedules produce more persistent behaviour than continuous reinforcement?+
The most widely supported explanation is the discrimination hypothesis: under continuous reinforcement, the shift to extinction is immediately detectable — reward was always present and is now absent. Under partial reinforcement, non-reward is part of the normal training experience, so the organism cannot easily discriminate the extinction phase from training. It therefore persists far longer before extinguishing. This is the partial reinforcement extinction effect (PREE), and it helps explain why compulsive behaviours maintained by intermittent reward — gambling, compulsive checking, certain relationship dynamics — are so difficult to extinguish even when the reinforcer is permanently withdrawn.
What is a reward prediction error?+
A reward prediction error is the discrepancy between the reward an organism expected and the reward it actually received. The Rescorla-Wagner model (1972) formalised this as the key driver of associative learning: if you receive more reward than you predicted, the association between the preceding stimuli and the outcome is strengthened; if you receive less than predicted, it weakens; if reward exactly matches prediction, no new learning occurs. The model elegantly explained blocking — the finding that a stimulus added to an already-sufficient predictor of reward fails to acquire any predictive value, because the reward is already fully predicted and no prediction error is generated.
What is delay discounting and how does it affect self-control?+
Delay discounting refers to the psychological tendency to value rewards less as they become more temporally distant. The crucial finding from Ainslie (1975) and subsequent research is that human discounting is hyperbolic, not exponential: the subjective value of a reward drops disproportionately steeply for small delays near the present. This creates preference reversals — at a distance, you prefer the larger, later reward, but as the smaller, sooner reward approaches, it temporarily overtakes the larger reward in subjective value. This is why self-control failures typically occur at the moment of temptation rather than in advance: the impulsive option has an artificially inflated subjective value when it is immediately available. Delay discounting is used as a measure of impulsivity in clinical psychology and is elevated in populations with addiction, ADHD, and conduct disorder.
Sources
Last reviewed August 2025- 1.
Thorndike E.L. (1911). Animal Intelligence: Experimental Studies. Macmillan.
+About this source
Contains the formal statement of the Law of Effect and the puzzle-box experimental series from which it was derived.
- 2.
Skinner B.F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century-Crofts.
+About this source
Introduced the operant conditioning framework and the distinction between respondent and operant behaviour.
- 3.
Ferster C.S. & Skinner B.F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
+About this source
The definitive systematic study of fixed and variable ratio and interval schedules, establishing the characteristic behavioural patterns each produces.
- 4.
Rescorla R.A. & Wagner A.R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A.H. Black & W.F. Prokasy (Eds.), Classical Conditioning II: Current Research and Theory (pp. 64–99). Appleton-Century-Crofts.
+About this source
Introduced the formal prediction error model of associative learning, predicting blocking and overexpectation and establishing prediction error as the fundamental learning signal.
- 5.
Ainslie G. (1975). Specious reward: A behavioral theory of impulsiveness and impulse control. Psychological Bulletin, 82(4), 463–496.
+About this source
Proposed hyperbolic discounting as the form of temporal reward devaluation, explaining preference reversals and their relevance to impulsivity and self-control.
- 6.
Jenkins W.O. & Stanley J.C. Jr. (1950). Partial reinforcement: A review and critique. Psychological Bulletin, 47(3), 193–234.
+About this source
Established and reviewed the partial reinforcement extinction effect (PREE) as one of the most robust findings in the conditioning literature.
- 7.
Madden G.J. & Bickel W.K. (Eds.) (2010). Impulsivity: The Behavioral and Neurological Science of Discounting. American Psychological Association.
+About this source
Comprehensive review of delay and probability discounting research, its measurement, and its relevance to clinical populations.