What is the key difference between quantitative and qualitative data?
A: Quantitative data comes from experiments; qualitative data comes from observations
B: Quantitative data consists of numbers (frequencies, scores, measurements) that can be analysed statistically; qualitative data consists of non-numerical information (words, descriptions, themes) that is analysed interpretively
C: Quantitative data is collected from large samples; qualitative data is collected from small ones
D: Quantitative data measures attitudes; qualitative data measures behaviour
Correct: Quantitative data consists of numbers (frequencies, scores, measurements) that can be analysed statistically; qualitative data consists of non-numerical information (words, descriptions, themes) that is analysed interpretively
Quantitative data is numerical — scores on a questionnaire, reaction times, frequencies of behaviour, physiological measurements — and lends itself to statistical analysis, hypothesis testing, and comparison across groups. Qualitative data is non-numerical — interview transcripts, open-ended survey responses, field notes, case study narratives — and is typically analysed through thematic analysis, content analysis, or grounded theory approaches. Each has strengths: quantitative data enables objective comparison and generalisation; qualitative data provides depth, nuance, and insight into meaning and experience. Mixed methods research combines both. Neither is inherently superior — the right choice depends on the research question.
What does a positive correlation between two variables mean?
A: As one variable increases, the other decreases
B: The relationship between the two variables is statistically significant
C: As one variable increases, the other also tends to increase
D: One variable is the cause of the other
Correct: As one variable increases, the other also tends to increase
A positive (or direct) correlation means that as values on one variable increase, values on the other tend to increase as well. For example, there is typically a positive correlation between hours of study and exam performance: more study hours tend to be associated with higher scores. A negative (or inverse) correlation means the variables move in opposite directions: as one increases, the other decreases (e.g. higher stress tends to be associated with lower quality of sleep). Zero correlation means no systematic linear relationship exists between the variables. Note that the direction (positive/negative) is logically separate from the strength and the question of statistical significance.
What does a correlation coefficient of –0.82 tell you?
A: A strong negative relationship: as one variable increases, the other tends to decrease, with scores clustering closely around the trend line
B: A weak negative relationship: there is a slight tendency for one variable to decrease as the other increases, but the pattern is unreliable
C: A strong positive relationship: the negative sign indicates the study found an unexpected result
D: No meaningful relationship: correlation coefficients below zero are not interpretable
Correct: A strong negative relationship: as one variable increases, the other tends to decrease, with scores clustering closely around the trend line
A correlation coefficient (r) ranges from –1.00 to +1.00. The sign indicates direction: negative means the variables move in opposite directions. The absolute value indicates strength: closer to 1 means stronger (data points cluster tightly around the trend line), closer to 0 means weaker (data points scatter widely). An r of –0.82 therefore indicates a strong negative relationship — for example, this could describe the relationship between perceived stress and self-reported wellbeing: higher stress scores associate consistently and closely with lower wellbeing scores. Common benchmarks (Cohen, 1988): |r| ≈ 0.1 = small, |r| ≈ 0.3 = medium, |r| ≈ 0.5 = large; –0.82 exceeds even the "large" threshold.
Why does correlation not establish causation?
A: Because correlational studies cannot be repeated, so the result may be a fluke
B: Because a correlation only shows that two variables are statistically associated — there may be reverse causation (B causes A rather than A causing B) or a third variable causing both, and no manipulation of one variable takes place to isolate its effect
C: Because correlations can only describe linear relationships, not non-linear causes
D: Because qualitative methods are required to understand causal mechanisms
Correct: Because a correlation only shows that two variables are statistically associated — there may be reverse causation (B causes A rather than A causing B) or a third variable causing both, and no manipulation of one variable takes place to isolate its effect
Three reasons prevent correlational evidence from establishing causation. First, reverse causation: A correlates with B does not tell us which drives which (depression correlates with social withdrawal — does depression cause withdrawal, or does withdrawal deepen depression, or both?). Second, the third-variable problem: a confounding variable C may cause both A and B, producing a correlation between them with no direct causal link (e.g. ice cream sales and drowning rates both correlate with summer temperature). Third, no manipulation: in a true experiment, the researcher changes the IV and observes the effect on the DV, isolating the causal direction. Correlational designs make no manipulation, so the direction of influence remains undetermined.
A study reports p = .03 for its main finding, with a significance threshold of p < .05. What does this mean?
A: There is a 3% chance that the experimental hypothesis is correct
B: There is a 3% probability of obtaining results at least this extreme if the null hypothesis were true; because p < .05, the null hypothesis is rejected
C: The finding is 97% accurate
D: The effect is three times stronger than would be expected by chance
Correct: There is a 3% probability of obtaining results at least this extreme if the null hypothesis were true; because p < .05, the null hypothesis is rejected
The p-value answers a specific question: given that the null hypothesis (no effect) is true, how probable is it that we would observe a result at least as extreme as the one we obtained? A p-value of .03 means this probability is 3% — a fairly rare outcome under the null hypothesis. Researchers typically set an alpha (significance) threshold of .05, meaning they accept a 5% risk of falsely rejecting the null hypothesis. Since .03 < .05, the null hypothesis is rejected and the result is declared statistically significant. Common misconceptions: the p-value is not the probability that the null hypothesis is true, not the probability the result is a false positive, and says nothing about the effect's size or practical importance.
What is a Type I error and what determines its maximum acceptable rate?
A: Failing to detect a real effect; its rate is determined by statistical power
B: Rejecting the null hypothesis when it is actually true (a false positive); its maximum acceptable rate is set by the significance threshold alpha (typically .05)
C: Using the wrong statistical test for the data type; its rate is controlled by pre-registration
D: A measurement error caused by poorly operationalised variables; controlled by piloting
Correct: Rejecting the null hypothesis when it is actually true (a false positive); its maximum acceptable rate is set by the significance threshold alpha (typically .05)
A Type I error is a false positive: the researcher concludes there is an effect when, in reality, there is none (the null hypothesis is true but is incorrectly rejected). The alpha level sets the maximum tolerable Type I error rate. With α = .05, the researcher accepts that in 5% of studies where there is truly no effect, chance variation will produce results significant enough to reject the null. A stricter alpha (e.g. α = .01) reduces Type I errors but increases Type II errors. A Type II error is the opposite: failing to detect a real effect (false negative). The probability of a Type II error is β; statistical power = 1 – β, and it increases with larger sample sizes, larger effects, and more sensitive measures.
A study finds a statistically significant difference between two groups (p = .004) but the effect size (Cohen's d) is 0.15. What does this combination tell you?
A: The result is both statistically and practically significant — p = .004 confirms a large effect
B: The result is statistically significant (unlikely to be due to chance) but practically small — the groups barely differ in real-world terms, likely because the sample was large enough to detect even tiny effects
C: There must be an error — significant p-values always accompany large effect sizes
D: The effect size needs to be converted before interpretation, since d = 0.15 is undefined
Correct: The result is statistically significant (unlikely to be due to chance) but practically small — the groups barely differ in real-world terms, likely because the sample was large enough to detect even tiny effects
Statistical significance and effect size measure different things. A p-value tells you how likely the result is due to chance; with a large enough sample, even a tiny, trivial difference will be statistically significant. Effect size tells you how large the difference actually is — what proportion of the variability is explained, or how many standard deviations apart the group means are. Cohen's d benchmarks: d = 0.2 (small), d = 0.5 (medium), d = 0.8 (large). A d of 0.15 is below even the "small" threshold — the groups differ by less than one-fifth of a standard deviation, meaning the difference is probably too small to matter in practice. With a very large sample (e.g. n = 10,000), even this negligible difference becomes statistically significant because the study has the power to detect it.
What is Cohen's d and how is it interpreted?
A: A measure of the probability that a result is due to chance; d values below .05 indicate significance
B: A standardised effect size measure for the difference between two group means, expressed in pooled standard deviation units; d ≈ 0.2 = small, d ≈ 0.5 = medium, d ≈ 0.8 = large
C: The correlation coefficient between two continuous variables; d values closer to 1.0 indicate a stronger relationship
D: A measure of the internal consistency of a scale; d > 0.7 indicates acceptable reliability
Correct: A standardised effect size measure for the difference between two group means, expressed in pooled standard deviation units; d ≈ 0.2 = small, d ≈ 0.5 = medium, d ≈ 0.8 = large
Cohen's d is calculated as the difference between two group means divided by the pooled standard deviation: d = (M₁ − M₂) / SD_pooled. Expressing the difference in standard deviation units makes it comparable across studies that use different scales — you can directly compare the effect of a memory training programme measured in words recalled (SD = 4) with one measured in a percentage score (SD = 12), because d puts both on the same standardised metric. Jacob Cohen's (1988) benchmarks — small (0.2), medium (0.5), large (0.8) — were intended as rough guides, not rigid rules; a "small" effect in medicine (where it might save many lives at population scale) may be practically important even if it looks tiny. Effect sizes are the basis of meta-analysis, where studies are combined by averaging their standardised effects.
Data Analysis & Statistics
What is the key difference between quantitative and qualitative data?
About this quiz
Collecting data is only half the work. Interpreting it — knowing how confident to be in a result, what a correlation can and cannot tell you, and how much difference is "enough" to matter — is where the real decisions are made.
This quiz covers the qualitative-vs-quantitative distinction, correlation and its coefficient, why correlation does not equal causation, statistical significance and p-values, Type I and Type II errors, and effect sizes including Cohen's d.