Causality is the hardest claim to make in health research, yet some studies with no random groups still come close to a trial. In 2024, a team used data on 1,097 people with a first bout of psychosis to mimic a trial of antipsychotic drugs. Their risks of stopping treatment and of going to hospital were close to those of an earlier randomised trial.1
Not every study design earns that result. In fact, clinical research splits into two kinds: in experiments, the researcher assigns the exposure; in observational studies, the researcher only watches.2 Here we test three designs against the Bradford Hill criteria: a randomised controlled trial (RCT), a case-control study, and propensity score matching (PSM).
Key point: a study’s design sets the ceiling on what it can claim about causality.
| Objective | Outcome |
|---|---|
| Know the nine Bradford Hill criteria | Judge a causal claim on more than one test |
| Compare case-control, PSM, and RCT designs | See how each handles bias and time order |
| Map each design to the criteria | Know which evidence a design cannot supply |
What the Bradford Hill Criteria Ask of Causality
In 1965, Austin Bradford Hill set out a way to judge whether a link between something in our surroundings and a disease is a cause.3 Reviews still apply all nine: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy.4 For example, a 2021 review of river blindness and epilepsy shows the method at work. The data met strength, consistency, temporality, and biological gradient, but gave little from experiment.4 So think of the criteria as a lens on causality, not a verdict. Indeed, Rothman and Greenland argue that no one can prove a causal claim, so the real task is to measure an effect rather than to tick boxes.5
No single criterion proves cause; the pattern across criteria builds the case.
RCT, Case-Control, and PSM: Three Routes of Study Design
A case-control study starts with the outcome. Researchers find people with a disease (cases) and people without it (controls), then compare each group’s past exposures.6 This reverse direction saves time and money, but it invites selection bias in how researchers choose controls and recall bias in how people report the past.6 A prospective version enrols new cases as they arise, or nests the study inside a cohort. As a result, cases and controls come from the same group of people.7
Propensity score matching reuses old records. First, it gives each person a score, the chance they got the treatment given the traits in the records; then it pairs treated and untreated people with close scores.8 Since the score uses only recorded traits, it cannot balance a confounder that no one wrote down. An RCT closes that gap for causality, because chance decides who gets the exposure.2
| Design | Starts from | Controls confounding by | Main weakness |
|---|---|---|---|
| Case-control | Outcome, then looks back | Matching or adjustment for measured factors | Selection and recall bias |
| PSM | Existing records of exposure | Balancing observed covariates | Unmeasured confounding |
| RCT | Random assignment of exposure | Chance allocation | Cost; some exposures cannot be assigned |
Only chance allocation deals with confounders that no one thought to measure.
Mapping Each Design to the Hill Criteria
Design decides which criteria one study can meet, and so how far it can go toward causality. In contrast, an RCT meets experiment and temporality head on. A PSM study can show temporality when the records log exposure before the outcome. Trial emulation shows how close careful work of this kind can come.1 A case-control study mainly shows strength and, with graded exposure data, a dose-response. However, consistency comes only from many studies, which is why reviews apply the criteria to a whole body of work.4
Observational designs need support from other studies to match what one good trial supplies.
Takeaway
No study design proves causality alone. An RCT best meets experiment and temporality, while case-control and PSM studies need the other Bradford Hill criteria, drawn from several studies, to back a causal claim. Before you trust a claim, ask which criteria its design can satisfy, because a p value speaks only to chance.2
“Causal inference in epidemiology is better viewed as an exercise in measurement of an effect rather than as a criterion-guided process …”5
Kenneth Rothman and Sander Greenland, American Journal of Public Health, 2005
References
- Szmulewicz AG, Martínez-Alés G, Logan R, et al. Antipsychotic drugs in first-episode psychosis: a target trial emulation in the FEP-CAUSAL Collaboration. Am J Epidemiol. 2024;193(8):1081-7. doi:10.1093/aje/kwae029
- Grimes DA, Schulz KF. An overview of clinical research: the lay of the land. Lancet. 2002;359(9300):57-61. doi:10.1016/S0140-6736(02)07283-5
- Hill AB. The environment and disease: association or causation? Proc R Soc Med. 1965;58(5):295-300. doi:10.1177/003591576505800503
- Colebunders R, Njamnshi AK, Menon S, et al. Onchocerca volvulus and epilepsy: a comprehensive review using the Bradford Hill criteria for causation. PLoS Negl Trop Dis. 2021;15(1):e0008965. doi:10.1371/journal.pntd.0008965
- Rothman KJ, Greenland S. Causation and causal inference in epidemiology. Am J Public Health. 2005;95 Suppl 1:S144-50. doi:10.2105/AJPH.2004.059204
- Schulz KF, Grimes DA. Case-control studies: research in reverse. Lancet. 2002;359(9304):431-4. doi:10.1016/S0140-6736(02)07605-5
- Wacholder S, Silverman DT, McLaughlin JK, Mandel JS. Selection of controls in case-control studies. III. Design options. Am J Epidemiol. 1992;135(9):1042-50. doi:10.1093/oxfordjournals.aje.a116398
- Austin PC. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate Behav Res. 2011;46(3):399-424. doi:10.1080/00273171.2011.568786
Useful Resources
- CDC Principles of Epidemiology: Analytic Epidemiology: how experimental, cohort, case-control, and cross-sectional studies test a hypothesis.
- StatPearls: Case Control Studies (NCBI Bookshelf): methods, strengths, and weaknesses of the case-control design.
- National Library of Medicine: Randomized Clinical Trials: why trials support causal conclusions, and where they fall short.
- MeSH: Randomized Controlled Trial: the formal definition used to index trials in PubMed.
Frequently Asked Questions
Can a case-control study prove causation?
No single study proves causality. Case-control studies answer questions quickly and cheaply, but selection and recall bias can distort them, so they work best as one strand of evidence. The StatPearls overview sets out the trade-offs.
Is propensity score matching as good as an RCT?
Not by default. Matching balances only the traits you measured, so hidden confounders remain. With rich data and a clear target trial, results can come close, as our guide to propensity score matching explains.
Which Bradford Hill criterion matters most?
Temporality comes closest to essential, since a cause must come before its effect. The others add weight rather than proof. The CDC epidemiology lesson shows how designs establish time order.
Why do randomised trials rank so high?
With random assignment, no trait of the people in the trial decides who gets the exposure. That guards causality claims against confounders nobody measured. The National Library of Medicine module covers their strengths and limits.
How does meta-analysis help with the consistency criterion?
Consistency asks whether different studies find the same link. A meta-analysis pools those studies and tests whether they agree. Our Meta-Analysis wiki entry explains how.






Leave a Reply