Understanding validity & reliability

A study can be perfectly consistent and still measure the wrong thing entirely. These two pillars of psychological measurement — reliability and validity — are related but distinct, and conflating them is one of the most common errors in interpreting research.

Lee Cronbach

1916–2001

American educational psychologist who developed Cronbach's alpha (1951), the most widely used measure of internal consistency reliability for multi-item scales. Alpha estimates the average inter-item correlation, capturing whether items on a scale that purports to measure a single construct actually move together. Cronbach also articulated the generalisability theory framework for understanding measurement error — moving beyond single reliability coefficients to a more nuanced understanding of consistency across persons, items, and occasions.

Jacob Cohen

1923–1998

Statistician and psychologist who formalised the concept of statistical power and effect sizes in Psychological Bulletin (1962) and in Statistical Power Analysis for the Behavioral Sciences (1969/1988). Cohen demonstrated that most psychology studies at the time were drastically underpowered, identified conventions for small, medium, and large effects (d = 0.2, 0.5, 0.8), and laid the groundwork for the replication crisis by showing how low-powered studies produce unreliable, irreproducible results.

Samuel Messick

1931–1998

Psychometrician at Educational Testing Service who developed the most comprehensive modern account of construct validity. Messick argued that validity is not a property of a test but of the inferences drawn from test scores, and that validity evidence must address not only whether a test measures the intended construct (content, convergent, discriminant validity) but also the consequences of test use — introducing a social and ethical dimension to measurement evaluation.

Internal validity

The degree to which observed changes in the DV can be attributed to the IV rather than to extraneous factors. High internal validity means the study's design adequately controls confounds, demand characteristics, and other threats so that a causal conclusion is justified. Laboratory experiments with random allocation, standardised procedures, and blind designs maximise internal validity.

External validity

The degree to which findings generalise beyond the immediate study — to other people, settings, times, and measures. Threatened by samples that are unrepresentative of the target population (WEIRD samples: Western, Educated, Industrialised, Rich, Democratic), artificial laboratory settings, and highly specific operationalisations. Internal and external validity are often in tension: the controls that maximise internal validity tend to reduce external validity.

Ecological validity

A sub-type of external validity: whether the study's setting, materials, and tasks resemble the real-world conditions to which findings are meant to generalise. A study is ecologically valid if the behaviour observed reflects what people actually do in natural settings. Many landmark laboratory paradigms — Milgram's obedience setup, the Stroop task, laboratory word-recall tasks — have been criticised for low ecological validity.

Construct validity

Whether a measure accurately reflects the theoretical construct it is intended to measure — as a whole, including all its facets and their relationships to related constructs. A test of "working memory" has construct validity if it correlates with other working memory measures, predicts outcomes theoretically linked to working memory (academic achievement, fluid intelligence), and does not correlate with measures it theoretically should not (e.g. long-term memory capacity).

Face validity

Whether the test appears, on the surface, to measure what it claims. The weakest form of validity evidence — appearance and substance can diverge. Important for practical acceptance (participants who find items irrelevant may not engage honestly) but provides no statistical evidence that the measure captures the target construct.

Criterion validity

Whether scores on a measure correlate with an external criterion known to reflect the same construct. Concurrent validity: criterion measured at the same time as the test (e.g. a new anxiety questionnaire correlated with clinician ratings). Predictive validity: criterion measured later (e.g. aptitude test scores predicting job performance 12 months after hiring).

Test-retest reliability

Consistency over time: the same test administered to the same participants at two time points produces similar scores (high correlation). The critical assumption is that the measured trait is genuinely stable across the chosen interval. Vulnerable to practice effects (better performance on second administration from memory) and genuine change in the attribute — which will lower reliability estimates even if the test itself is perfectly constructed.

Inter-rater reliability

The degree to which two or more independent observers coding the same behaviour reach the same conclusions. Quantified using percentage agreement (simple but inflated by chance coincidences) or Cohen's kappa (adjusts for chance agreement; κ ≥ 0.7 typically considered acceptable). Critical in observational studies and qualitative coding, where subjective judgement is involved.

Standardisation

Keeping the procedure constant for all participants — same instructions, materials, timing, testing environment, and scoring criteria. Reduces variability introduced by procedural inconsistency so that differences in data reflect real variation in the construct rather than variation in how the study was conducted. Also used to refer to establishing norms against a representative population sample in psychometric testing.

Does a reliable test have to be valid?+

No — reliability is necessary but not sufficient for validity. A measure can be perfectly reliable (gives the same result every time) while being invalid (systematically measuring something other than the target construct). The classic analogy: a bathroom scale that consistently reads 5 kg too heavy is reliable but invalid. In psychology, a test that reliably produces high scores for people who read quickly might be a reliable measure of reading speed while being an invalid measure of reading comprehension. The asymmetry runs one way only: an unreliable measure cannot be valid, because random error means it cannot consistently capture any construct.

Is high internal validity enough for a study to be useful?+

Internal validity is necessary for establishing causation, but a study with only internal validity tells you that X causes Y in the specific conditions of that study — nothing more. If the participant sample is unrepresentative (e.g. all male university students), the setting artificial, the stimuli unlike real-world equivalents, and the DV operationalised in an unusual way, the causal finding may not generalise to the populations, settings, and measures that actually matter. High internal validity is the foundation; generalisability (external validity) determines whether the finding is practically or theoretically significant beyond the lab.

What is the difference between inter-rater reliability and test-retest reliability?+

Test-retest reliability assesses consistency across time for the same measure: the same test given to the same people on two occasions should produce correlated scores. Inter-rater reliability assesses consistency across observers for the same data: two people independently coding the same behaviour or responses should reach the same conclusions. Test-retest is relevant for any self-report or performance measure repeated over time. Inter-rater reliability is specifically relevant when human judgement is required to score or classify responses — in observational studies, interview coding, qualitative analysis, or clinical assessment.

Last reviewed July 2025
  1. 1.

    Cronbach L.J. & Meehl P.E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.

    +About this source

    Foundational paper defining construct validity and the nomological network framework for validation.

  2. 2.

    Messick S. (1989). Validity. In R.L. Linn (Ed.), Educational Measurement (3rd ed., pp. 13–103). American Council on Education.

    +About this source

    Comprehensive unified framework for validity integrating content, criterion, and construct aspects.