Validity & Reliability

1 / 9

Term

Internal validity

Tap to reveal definition

Definition

The degree to which observed changes in the DV can be attributed to the IV rather than to extraneous factors. High internal validity means the study's design adequately controls confounds, demand characteristics, and other threats so that a causal conclusion is justified. Laboratory experiments with random allocation, standardised procedures, and blind designs maximise internal validity.

Tap to flip back

Space to flip · ← → to navigate

All 9 Terms & Definitions

Internal validity
The degree to which observed changes in the DV can be attributed to the IV rather than to extraneous factors. High internal validity means the study's design adequately controls confounds, demand characteristics, and other threats so that a causal conclusion is justified. Laboratory experiments with random allocation, standardised procedures, and blind designs maximise internal validity.
External validity
The degree to which findings generalise beyond the immediate study — to other people, settings, times, and measures. Threatened by samples that are unrepresentative of the target population (WEIRD samples: Western, Educated, Industrialised, Rich, Democratic), artificial laboratory settings, and highly specific operationalisations. Internal and external validity are often in tension: the controls that maximise internal validity tend to reduce external validity.
Ecological validity
A sub-type of external validity: whether the study's setting, materials, and tasks resemble the real-world conditions to which findings are meant to generalise. A study is ecologically valid if the behaviour observed reflects what people actually do in natural settings. Many landmark laboratory paradigms — Milgram's obedience setup, the Stroop task, laboratory word-recall tasks — have been criticised for low ecological validity.
Construct validity
Whether a measure accurately reflects the theoretical construct it is intended to measure — as a whole, including all its facets and their relationships to related constructs. A test of "working memory" has construct validity if it correlates with other working memory measures, predicts outcomes theoretically linked to working memory (academic achievement, fluid intelligence), and does not correlate with measures it theoretically should not (e.g. long-term memory capacity).
Face validity
Whether the test appears, on the surface, to measure what it claims. The weakest form of validity evidence — appearance and substance can diverge. Important for practical acceptance (participants who find items irrelevant may not engage honestly) but provides no statistical evidence that the measure captures the target construct.
Criterion validity
Whether scores on a measure correlate with an external criterion known to reflect the same construct. Concurrent validity: criterion measured at the same time as the test (e.g. a new anxiety questionnaire correlated with clinician ratings). Predictive validity: criterion measured later (e.g. aptitude test scores predicting job performance 12 months after hiring).
Test-retest reliability
Consistency over time: the same test administered to the same participants at two time points produces similar scores (high correlation). The critical assumption is that the measured trait is genuinely stable across the chosen interval. Vulnerable to practice effects (better performance on second administration from memory) and genuine change in the attribute — which will lower reliability estimates even if the test itself is perfectly constructed.
Inter-rater reliability
The degree to which two or more independent observers coding the same behaviour reach the same conclusions. Quantified using percentage agreement (simple but inflated by chance coincidences) or Cohen's kappa (adjusts for chance agreement; κ ≥ 0.7 typically considered acceptable). Critical in observational studies and qualitative coding, where subjective judgement is involved.
Standardisation
Keeping the procedure constant for all participants — same instructions, materials, timing, testing environment, and scoring criteria. Reduces variability introduced by procedural inconsistency so that differences in data reflect real variation in the construct rather than variation in how the study was conducted. Also used to refer to establishing norms against a representative population sample in psychometric testing.