Cognitive Connie
What is operant conditioning?
Operant conditioning describes learning through the consequences of behaviour. Unlike classical conditioning, which involves associating stimuli, operant conditioning involves associating actions with their outcomes: behaviours that produce rewarding consequences are strengthened and repeated; behaviours that produce aversive consequences — or that fail to produce expected rewards — are weakened and discontinued. The term was coined by B.F. Skinner, who developed the framework from Edward Thorndike's earlier Law of Effect.
01. Overview
The four quadrants of operant conditioning.
Every operant consequence falls into one of four categories, defined by two dimensions: whether a stimulus is added or removed, and whether behaviour increases or decreases as a result.
Positive Reinforcement
Adding a desirable stimulus following a behaviour, increasing the likelihood the behaviour will be repeated. Praise, money, food, and attention are common positive reinforcers. The most effective and ethically preferred form of behaviour change — it works by creating motivation toward a behaviour rather than away from an aversive state.
Negative Reinforcement
Removing an aversive stimulus following a behaviour, increasing the likelihood the behaviour will be repeated. Taking paracetamol removes pain — the relief reinforces the pill-taking behaviour. Buckling a seatbelt removes an annoying alarm — the relief reinforces seatbelt use. Often confused with punishment: negative reinforcement always increases behaviour.
Positive Punishment
Adding an aversive stimulus following a behaviour, decreasing the likelihood the behaviour will be repeated. A speeding fine, a reprimand, or a painful electric shock are positive punishers. Effective at suppressing behaviour quickly but associated with negative side effects: emotional distress, avoidance of the punisher, and failure to teach alternative behaviours.
Negative Punishment
Removing a desirable stimulus following a behaviour, decreasing the likelihood the behaviour will be repeated. Taking away a teenager's phone for breaking a curfew; removing TV privileges after misbehaviour. Generally considered more ethical than positive punishment as it involves withdrawal rather than the imposition of pain.
Key figures
Edward Thorndike
1874–1949American psychologist whose puzzle-box experiments with cats produced the Law of Effect — the precursor to operant conditioning. Thorndike showed that learning is driven by the consequences of behaviour, not by insight or understanding, establishing the core logic that Skinner later formalised and extended.
American psychologist who developed operant conditioning into a comprehensive scientific framework, invented the operant chamber, and explored its applications in education (programmed learning), behaviour modification, and the design of ideal societies (Walden Two). His 1938 book The Behaviour of Organisms laid the empirical foundation; his 1971 Beyond Freedom and Dignity applied the framework to society — controversially arguing that concepts of free will and personal responsibility are incompatible with a scientific account of behaviour.
Key concepts
Schedules of reinforcement
The pattern governing when a behaviour is reinforced. Four basic schedules: fixed ratio (FR — reinforce after every nth response, e.g. piece-rate pay), variable ratio (VR — reinforce after an unpredictable number of responses, e.g. slot machines — produces highest and most persistent responding), fixed interval (FI — reinforce first response after a set time, e.g. weekly salary — produces scalloping pattern), variable interval (VI — reinforce after unpredictable time intervals — produces slow, steady responding). Variable ratio schedules produce the greatest resistance to extinction.
Shaping
Reinforcing successive approximations to a target behaviour that the organism has not yet performed. By reinforcing behaviours that are closer and closer to the desired behaviour, complex behaviours that could never occur spontaneously can be established — a pigeon pecking a specific target, a child learning to write. Shaping is the operant mechanism behind most skill acquisition.
Extinction (operant)
The weakening of an operant behaviour when reinforcement is withheld. If pressing the lever no longer produces food, the rat eventually stops pressing. Extinction is often preceded by an extinction burst — a temporary increase in the rate and intensity of the behaviour before it declines — which explains why parents who eventually give in to a tantrum actually reinforce it more powerfully.
Discriminative stimulus
A stimulus that signals when reinforcement is available. In the presence of a green light, lever-pressing produces food; in the presence of a red light, it does not. The green light becomes a discriminative stimulus (SD) that controls the probability of responding. Discriminative stimuli are everywhere: a phone notification signals that checking your phone will be rewarding; opening hours on a shop signal that entering will produce a desired outcome.
Thorndike's Law of Effect
Edward Thorndike's foundational principle (1898) that behaviours producing satisfying consequences are strengthened (stamped in) and behaviours producing unsatisfying consequences are weakened (stamped out). Derived from puzzle-box experiments with cats, it established the core logic that Skinner formalised into operant conditioning — that learning is fundamentally about the consequences of behaviour.
Test your knowledge
Frequently asked questions
What is operant conditioning?+
Operant conditioning is a form of learning in which behaviour is shaped by its consequences. Behaviours that produce rewarding outcomes are reinforced (made more likely to occur again); behaviours that produce aversive outcomes, or that fail to produce expected rewards, are weakened. Developed by B.F. Skinner from Thorndike's Law of Effect, it is one of the most thoroughly researched frameworks in psychology and underlies much of applied behaviour analysis, behaviour therapy, and everyday practices of reward and discipline.
What is the difference between positive and negative reinforcement?+
Both positive and negative reinforcement increase the likelihood of a behaviour — they differ in how. Positive reinforcement adds a desirable stimulus (you receive a bonus for good work — the bonus increases your motivation to work hard). Negative reinforcement removes an aversive stimulus (you take a painkiller to relieve a headache — the relief reinforces pill-taking). The confusion arises because 'negative' sounds like punishment. Remember: reinforcement always increases behaviour; what's 'negative' is the removal of something, not the outcome.
What is the difference between reinforcement and punishment?+
Reinforcement (positive or negative) increases the likelihood of a behaviour being repeated. Punishment (positive or negative) decreases it. Positive reinforcement adds something pleasant; negative reinforcement removes something unpleasant — both make behaviour more frequent. Positive punishment adds something aversive; negative punishment removes something pleasant — both make behaviour less frequent. The 'positive/negative' labels refer to adding or removing a stimulus, not to whether the outcome is good or bad.
What are the schedules of reinforcement?+
Schedules of reinforcement describe when a behaviour is reinforced. The four basic schedules are: fixed ratio (every nth response is reinforced — like piece-rate pay), variable ratio (reinforcement after an unpredictable number of responses — like a slot machine, produces the most persistent and rapid responding), fixed interval (first response after a set time is reinforced — like a weekly salary, produces a scalloping pattern), and variable interval (reinforcement after an unpredictable time — like randomly checking social media, produces slow, steady responding). Variable ratio schedules are the most resistant to extinction, which is why gambling behaviour is so hard to extinguish.
What is an example of operant conditioning in everyday life?+
Operant conditioning is everywhere once you look for it. Children receive praise (positive reinforcement) for good behaviour; teenagers lose phone privileges (negative punishment) for breaking rules. An employee works harder when promised a bonus (positive reinforcement). Someone buckles their seatbelt to stop an annoying alarm (negative reinforcement). Social media platforms are engineered using variable ratio reinforcement — likes and notifications arrive unpredictably, producing compulsive checking. Even the extinguishing of behaviour matters: if you ignore a child's tantrum consistently (withhold reinforcement), the behaviour should eventually decrease.
What is a variable ratio schedule and why is it so compelling?+
A variable ratio (VR) schedule delivers reinforcement after an unpredictable number of responses — sometimes after 3, sometimes after 20, sometimes after 7. Because the next reward might be the very next response, the organism keeps responding at a high, steady rate and does not pause to 'wait out' a predictable gap. VR schedules also produce behaviour that is highly resistant to extinction: the organism has learned that non-reward sometimes just means the next reward is near. This is why slot machines, social media notifications, and loot boxes exploit VR mechanics — the unpredictability sustains engagement far more effectively than fixed delivery.
What is the difference between operant and classical conditioning?+
Classical conditioning (Pavlov) involves reflexive, involuntary responses — a neutral stimulus becomes associated with an unconditioned stimulus until it elicits the same response on its own. The organism is passive; learning is about prediction of what comes next. Operant conditioning involves voluntary behaviour — the organism acts on its environment, and the consequences of that action determine whether the behaviour is repeated. The organism is active; learning is about control over outcomes. In practice, both processes often operate together: a fear response classically conditioned to a situation may be maintained operantly because avoidance behaviour (negatively reinforced by reduced anxiety) prevents extinction.
What is shaping and how is it used in practice?+
Shaping is a method for establishing behaviours that an organism would never spontaneously produce by reinforcing successive approximations — behaviours that are progressively closer to the desired target. Each time the organism meets the current criterion, the bar is raised slightly. Shaping underlies most professional animal training, clinical behaviour modification in ABA, and many educational approaches. It works because you do not need to wait for the complete target behaviour to appear; you build it incrementally from responses the organism already makes.
Sources
Last reviewed July 2025- 1.
Skinner, B. F. (1938). The Behavior of Organisms: An Experimental Analysis. Appleton-Century-Crofts.
+About this source
Skinner's foundational text introducing the operant conditioning framework and the Skinner box.
- 2.
Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. Psychological Review Monograph Supplements, 2(4), 1–109.
+About this source
Introduced the Law of Effect from puzzle-box experiments — the conceptual precursor to operant conditioning.
- 3.
Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts.
+About this source
The definitive empirical account of how different reinforcement schedules affect behaviour.
- 4.
Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.
+About this source
The standard graduate textbook for applied behaviour analysis, covering both theory and clinical application.