What is internal validity, and which type of research design is most effective at maximising it?
A: The degree to which results can be applied to the real world; maximised by field experiments
B: The degree to which observed changes in the DV can be attributed to the IV rather than to extraneous factors; maximised by controlled experiments with random allocation
C: The degree to which a test measures what it is supposed to measure; maximised by factor analysis
D: The consistency of results across repeated measurements; maximised by longitudinal designs
Correct: The degree to which observed changes in the DV can be attributed to the IV rather than to extraneous factors; maximised by controlled experiments with random allocation
Internal validity refers to confidence that the IV — and not some other variable — caused the observed change in the DV. It is threatened by confounding variables, demand characteristics, experimenter bias, and order effects. Laboratory experiments with random allocation to conditions, standardised procedures, and blind designs maximise internal validity by controlling or eliminating these threats. However, the very controls that maximise internal validity (artificial settings, standardised stimuli, restricted participant samples) often reduce external validity — the extent to which findings generalise beyond the immediate study. This internal–external validity trade-off is a recurring tension in research design.
What is ecological validity, and why is it sometimes in tension with internal validity?
A: How well a study measures biological/environmental interactions; tension arises because lab conditions strip away natural stimuli
B: The degree to which a study's setting, tasks, and stimuli reflect real-world conditions, so that findings generalise to everyday behaviour; tension arises because the controls that maximise internal validity often produce artificial conditions that reduce ecological validity
C: Whether a study can be replicated in different ecological contexts; in tension with internal validity because replication reduces control
D: The degree to which participants represent the target population; in tension with internal validity because representative samples are harder to control
Correct: The degree to which a study's setting, tasks, and stimuli reflect real-world conditions, so that findings generalise to everyday behaviour; tension arises because the controls that maximise internal validity often produce artificial conditions that reduce ecological validity
Ecological validity is a sub-type of external validity: whether the conditions of the study (setting, materials, tasks) mirror the real world well enough that the findings generalise to everyday life. Many landmark laboratory studies — Milgram's obedience experiments, Asch's conformity studies, laboratory memory tasks — have been criticised for low ecological validity because their artificially controlled conditions bear little resemblance to naturalistic situations. The tension with internal validity is fundamental: adding more real-world complexity (multiple simultaneous variables, natural settings, participants' normal routines) makes it harder to isolate the IV's effect. Most research involves deliberate trade-offs between the two.
A researcher creates a new test of "emotional intelligence." After administering it, she correlates scores with managers' ratings of employees' emotional skill in the workplace. What type of validity is she assessing?
A: Face validity — whether the test looks like it measures emotional intelligence
B: Construct validity — whether the test measures the theoretical construct of emotional intelligence
C: Criterion validity — whether scores on the new test correlate with an external criterion known to reflect the construct
D: Test-retest validity — whether the test produces the same scores when administered twice
Correct: Criterion validity — whether scores on the new test correlate with an external criterion known to reflect the construct
Criterion validity assesses whether a test's scores correspond to an external criterion that independently measures the same (or a closely related) attribute. It has two forms: concurrent validity (criterion measured at the same time as the test) and predictive validity (criterion measured later — e.g. does the test score predict future job performance?). In this example, workplace ratings are the criterion, and the correlation between test scores and ratings indicates criterion validity. Construct validity is broader: it asks whether the test accurately reflects the theoretical construct as a whole — including correlating with other measures that should theoretically relate to it and not correlating with measures that should not.
What is face validity, and why is it considered the weakest form of validity evidence?
A: Whether the test is aesthetically appealing; weak because aesthetic judgements are subjective
B: Whether the test appears, on the surface, to measure what it claims to measure; weak because superficial resemblance to a construct does not guarantee the test actually measures it
C: Whether the test's items are statistically unrelated to each other; weak because inter-item correlation is insufficient evidence of validity
D: Whether the test produces the same results when used in different countries; weak because translations may alter meaning
Correct: Whether the test appears, on the surface, to measure what it claims to measure; weak because superficial resemblance to a construct does not guarantee the test actually measures it
Face validity is the most basic and subjective form of validity — it simply asks whether the test looks plausible to non-experts. A questionnaire asking "how often do you feel overwhelmed?" would typically have high face validity as a measure of stress. Its weakness is that appearance and substance can diverge: a test can look like it measures the right thing without actually doing so (or, conversely, an abstract test with no obvious face validity might still measure the construct effectively). Face validity is important for practical acceptance — participants who find questions irrelevant may not engage honestly — but it provides no statistical evidence that the measurement instrument actually captures the intended construct.
What does test-retest reliability measure, and what assumption must hold for it to be a valid method?
A: Whether two different tests of the same construct give the same result; requires the tests to be truly parallel forms
B: Whether the same test administered to the same participants at two time points produces similar scores; requires the underlying trait to be stable across the time period
C: Whether two researchers coding the same data reach the same conclusions; requires coders to use the same criteria
D: Whether the test behaves the same way in different populations; requires cultural equivalence of test items
Correct: Whether the same test administered to the same participants at two time points produces similar scores; requires the underlying trait to be stable across the time period
Test-retest reliability assesses consistency over time: administer the same test twice (with a gap of days or weeks), then calculate the correlation between the two sets of scores. A high correlation indicates the test produces stable results. The critical assumption is that the trait being measured is genuinely stable across the chosen time period — if you are measuring a stable personality trait, a few weeks' gap is appropriate; if you are measuring a rapidly fluctuating state (e.g. mood), test-retest reliability will naturally be low even if the test itself is perfectly constructed. The method is also vulnerable to practice effects (participants performing better the second time from memory) and genuine change in the measured attribute.
What is inter-rater reliability and when is it especially important to establish?
A: Agreement between a test score and a real-world criterion; important for tests used in clinical assessment
B: The correlation between two forms of the same test; important when the test will be used to compare large groups
C: The degree to which two or more independent observers coding the same behaviour reach the same conclusions; especially important in observational studies and qualitative coding, where subjective judgement is involved
D: Consistency of results across different testing rooms; especially important in multi-site clinical trials
Correct: The degree to which two or more independent observers coding the same behaviour reach the same conclusions; especially important in observational studies and qualitative coding, where subjective judgement is involved
Inter-rater reliability (also called inter-observer reliability) assesses whether different researchers independently coding the same data reach the same conclusions. It is quantified using Cohen's kappa (for categorical data) or intraclass correlation coefficients (for continuous data). It is especially critical in observational studies — where researchers must decide whether an observed behaviour counts as aggression, anxiety, or attachment behaviour — and in content analysis or qualitative coding. Without high inter-rater reliability, results may reflect the idiosyncratic judgements of a particular coder rather than the phenomenon being studied. Establishing inter-rater reliability requires coders to be trained on the coding scheme before independent coding begins.
What is standardisation in a research context, and how does it contribute to reliability?
A: Converting raw scores to a common scale (e.g. z-scores); ensures meaningful comparison across different measures
B: Ensuring every participant experiences exactly the same procedure, instructions, timing, and materials; reduces the variability introduced by procedural inconsistency so that differences in data reflect real variation in the construct rather than variation in how the study was run
C: Selecting a representative normative sample against which to compare individual scores; contributes to criterion validity
D: Using well-established measures rather than novel ones; contributes to construct validity by using instruments with known psychometric properties
Correct: Ensuring every participant experiences exactly the same procedure, instructions, timing, and materials; reduces the variability introduced by procedural inconsistency so that differences in data reflect real variation in the construct rather than variation in how the study was run
Standardisation means keeping the procedure constant for all participants — using the same instructions (often scripted), the same materials, the same timing, the same testing environment, and the same scoring criteria. This is crucial for reliability because unreliability can come from the test itself (items are ambiguous) or from the testing process (different experimenters explain tasks differently, or participants are tested at different times of day). A standardised procedure ensures that if the same participant were tested again, or if a different researcher ran the study, conditions would be as close to identical as possible. In psychometric testing, standardisation also refers to the process of establishing norms against a representative population sample.
A highly reliable test is necessarily also valid.
Answer: False
Reliability and validity are independent properties — and reliability does not guarantee validity. A bathroom scale that consistently reads 5 kg too heavy is perfectly reliable (it gives the same reading every time) but invalid (the readings are wrong). Likewise, a psychological test can produce highly consistent scores while systematically measuring the wrong construct. For example, a word-reading speed test administered as a measure of "reading comprehension" might produce very reliable scores across repeated administrations, while actually measuring reading fluency rather than comprehension. The important asymmetry is directional: an unreliable measure cannot be valid (random error means it cannot consistently capture anything), but a reliable measure may still be invalid if it captures something other than the target construct.
Validity & Reliability
What is internal validity, and which type of research design is most effective at maximising it?
About this quiz
A study can be perfectly consistent and still measure the wrong thing. Reliability and validity are the two pillars that hold up any psychological measurement — reliability asks whether the measure is consistent, validity asks whether it actually captures what it claims to capture.
This quiz covers internal and external validity, their major threats, the forms of validity that apply to measurement instruments (face, construct, and criterion), and the two main ways to assess the reliability of a test or observation.