Cognitive Connie
What is experimental design?
Choosing how to assign participants to conditions is one of the most consequential decisions in experimental design. It affects how many participants you need, what threats to validity you face, and what statistical procedures are appropriate. No design is universally superior — each is a set of trade-offs between internal validity, ecological validity, and practical constraints.
Key figures
R. A. Fisher
1890–1962British statistician who formalised the logic of experimental design in The Design of Experiments (1935). Fisher introduced the concepts of randomisation, replication, and control as the three core principles of experimental validity, and developed the analysis of variance (ANOVA) to handle multi-condition and multi-factor designs. His work at Rothamsted Agricultural Research Station provided the foundational framework that all subsequent experimental psychology adopted.
Donald T. Campbell
1916–1996American psychologist who, with Julian Stanley, published Experimental and Quasi-Experimental Designs for Research (1963) — the seminal taxonomy of threats to internal and external validity. Campbell systematised the threats that experimental designs must address (history, maturation, testing effects, instrumentation, selection bias, etc.) and introduced the quasi-experimental designs appropriate when true random assignment is impossible.
Martin Orne
1927–2000American psychiatrist who coined the term "demand characteristics" in 1962, demonstrating that participants in psychological experiments actively seek cues about the study's purpose and modify their behaviour accordingly. His work transformed researchers' understanding of the social nature of the experiment and made demand characteristics a standard consideration in experimental design — motivating blinding procedures and more naturalistic paradigms.
Key concepts
Independent groups design
Different participants are assigned to each condition. Advantages: no order effects (participants only experience one condition). Disadvantages: participant variables may differ between groups, requiring more participants and random allocation to minimise pre-existing differences. Also called between-subjects design.
Repeated measures design
The same participants complete all conditions, serving as their own control. Advantages: participant variables are controlled (differences between conditions cannot be due to pre-existing participant differences); requires fewer participants. Disadvantages: order effects (practice, fatigue, boredom) and heightened demand characteristics (participants more easily guess the study's purpose across conditions). Also called within-subjects design.
Matched pairs design
Participants are paired on key variables likely to affect the DV (e.g. IQ, age, baseline performance), then one member of each pair is assigned to each condition. Combines reduced participant variability (like repeated measures) with absence of order effects (like independent groups). The main limitation is the practical difficulty of finding well-matched pairs, especially on multiple variables simultaneously.
Order effects
In repeated measures designs, performance in a later condition may be affected by experience from an earlier one. Practice effects improve performance with task familiarity; fatigue and boredom effects impair later performance. Order effects do not exist in independent groups designs because each participant experiences only one condition.
Counterbalancing
The standard solution to order effects in repeated measures designs. Rather than all participants completing conditions in the same order (A then B), participants are divided: half complete A then B, half complete B then A. This distributes order effects evenly across conditions rather than consistently benefiting one. Counterbalancing does not eliminate order effects — it balances them. For more than two conditions, Latin square designs ensure each condition appears in each position equally often.
Laboratory experiment
Conducted in an artificial, controlled setting where the researcher manipulates the IV. Maximises control of extraneous variables and therefore internal validity. The trade-off is ecological validity: behaviour in an artificial setting may not reflect everyday behaviour. Associated with high internal validity but sometimes lower ecological validity.
Field experiment
Takes place in participants' natural environment with the IV still actively manipulated by the researcher. Preserves ecological validity but sacrifices some control over extraneous variables. Classic example: Bickman's (1974) uniform study, which manipulated authority by having a confederate in different uniforms give instructions to passers-by in a real street setting.
Natural experiment
The IV varies naturally or through circumstances beyond the researcher's control — policy changes, natural disasters, the introduction of technology to a community. Random allocation is impossible; participants are determined by circumstance. Has high ecological validity but cannot establish causation with the confidence of true experiments. Classic example: Charlton et al.'s study of children's behaviour before and after television was introduced to the island of St Helena.
Demand characteristics
Cues in the experimental situation that allow participants to guess the study's purpose, leading them to alter their behaviour in line with (or deliberately against) perceived expectations. Threaten internal validity because the DV may reflect participants' responses to their inferences about the study rather than to the IV alone. Minimised by single-blind procedures and ethical deception.
Double-blind procedure
Both participants and the researchers collecting or scoring data are kept unaware of condition allocation. Addresses demand characteristics (participants cannot adjust behaviour to perceived expectations) and experimenter bias (researchers cannot unconsciously treat participants differently or interpret ambiguous data in line with their expectations). Commonly used in drug trials where neither the patient nor the administering clinician knows who received the active drug.
Test your knowledge
Frequently asked questions
When should you use a repeated measures design instead of independent groups?+
Repeated measures is preferable when: (a) participant variability is high and you want to eliminate it as a source of noise; (b) you have limited participants available; (c) the conditions are sufficiently different in character that order effects can be managed through counterbalancing. Independent groups is preferable when: (a) participating in one condition would make the second meaningless (e.g. a study testing the effect of seeing a correct answer on a puzzle); (b) the two conditions are so similar that carry-over effects would be unavoidable; or (c) the study is long enough that fatigue effects would be severe.
Does counterbalancing completely solve the order effects problem?+
No. Counterbalancing balances order effects across conditions — ensuring that any practice or fatigue benefit is distributed equally — but it does not eliminate them. If a carry-over effect is asymmetric (AB produces a larger effect than BA), counterbalancing will not fully solve this. More fundamentally, counterbalancing assumes that individual participants' order effects average out across the sample; for small samples, this assumption may not hold. Where carry-over effects are severe and unavoidable, switching to an independent groups design may be the only viable option.
Why can natural experiments not prove causation?+
In a natural experiment, participants are not randomly allocated to conditions — they end up in one "condition" or another through circumstance (geography, policy, timing). This means pre-existing differences between groups may explain any observed difference in outcomes. Without random allocation, you cannot rule out the possibility that the groups differed on some relevant characteristic before the "natural manipulation" occurred. You can strengthen causal inference through pre/post measurement, statistical control, and multiple replications, but the fundamental limitation — no researcher-controlled random assignment — means causation cannot be established with the certainty of a true experiment.
Sources
Last reviewed July 2025- 1.
Shadish W.R., Cook T.D., & Campbell D.T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.
+About this source
Definitive text on experimental design, covering all major designs including natural experiments and threats to validity.
- 2.
Campbell D.T. & Stanley J.C. (1966). Experimental and Quasi-Experimental Designs for Research. Rand McNally.
+About this source
Foundational text introducing the experimental design typology and the distinction between lab, field, and natural experiments.