Written by Lee Stoner and Erik D Hanson.
One of the most accurate pneumonia detectors ever trained had a secret: it was partly reading the hospital, not the lungs. In other words, machine learning, software that learns patterns from data to make predictions, can be right for the wrong reasons.
Researchers trained the model on 158,323 chest X-rays from three hospital systems. For two of those hospitals, it could tell which one an image came from in more than 99.9% of cases. One of them also had far more pneumonia than the other.1 As a result, the model took a shortcut: the hospital that took an X-ray predicted the diagnosis, even though no hospital gives anyone pneumonia.
Here is the core problem: a variable that predicts an outcome is not always a cause of it, and machine learning cannot tell the difference on its own.2 Causal inference, the science of working out what happens when you change something, can.
| Objective | Outcome |
|---|---|
| Why accurate models can mislead | Spot the confounder behind a prediction |
| Prediction versus causation | Ask the right question of any data claim |
| How researchers test cause without a trial | Judge whether a study earns causal language |
Why Machine Learning Predicts Without Understanding
In predictive modeling, a machine learning model has one job: shrink its prediction error. Because of that, it has no reason to ask why a pattern holds, so if a variable tracks the outcome, the model uses it. Statisticians call that association, and association alone says nothing about cause.3
The pneumonia model shows the trap, because each hospital’s patient mix shaped both its X-rays and its pneumonia rate. That shared influence is a confounder: a third factor that links two things that do not affect each other. How do you catch one? Some researchers, especially those interested in causality, draw a directed acyclic graph (DAG), a map of arrows showing what they believe causes what. The map shows which confounding variables to account for before reading any link as causal.4
High accuracy proves a model found a pattern. It does not prove the pattern will hold once you change something.
Correlation vs Causation: When Accurate Models Give Bad Advice
Correlation can be absurd. A 2012 note in a leading medical journal showed that countries eating more chocolate also produce more Nobel laureates.5 Of course, nobody should buy truffles to win a prize, yet a model fed country data would happily use chocolate to predict one.
The real difference, then, is the question. Prediction asks what will happen, whereas causation asks a counterfactual question: what will happen if we change something? For that second question, randomized controlled trials (RCTs) assign treatment by chance, which is why medicine treats them as its gold standard.6 Even so, many observational papers still hide a causal goal behind the phrase “associated with”.7
| Question | Example | Best tool | Blind spot |
|---|---|---|---|
| What will happen? | Which patients are at highest risk? | Machine learning model | Cannot say what to change |
| What if we change X? | Does this drug lower risk? | Randomized controlled trial | Slow, costly, sometimes impossible |
| What if, without a trial? | What did the drug do in real users? | Causal inference methods | Only as good as the confounders measured |
Before trusting a data-driven recommendation, ask which of these questions the analysis actually answered.
Causal Inference Without a Trial
However, trials are not always possible. When they are not, causal inference offers a workaround: design the analysis like the trial you wish you could run. Researchers call this emulating a target trial.8
A SciencEye-led study applied this logic to wearable data. It tracked 66 people for 12 weeks after they started a GLP-1 receptor agonist, a class of diabetes and weight-loss drugs. Next, propensity score matching paired them with wearable users of similar measured traits not on the drug. Compared with those matches, users lost 10.0% of body weight (95% confidence interval, 8.5% to 11.2%), and resting heart rate rose 3.2 beats per minute.9
Good causal inference starts with study design, not with a bigger dataset.
Takeaway
Machine learning is great at prediction and weak at explanation. Causal inference fills that gap, and newer causal machine learning methods now aim to estimate what happens after a change.10 Until those tools mature, carry one question into every machine learning headline: was this shown to cause the outcome, or only to predict it?
“Using the term ‘causal’ is necessary to improve the quality of observational research.”7
Miguel A. Hernán, American Journal of Public Health, 2018
References
SciencEye checked each source below in PubMed, and every link opens the original article. Reference 9 comes from a SciencEye study, cited because its design shows causal inference at work.
- Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018;15(11):e1002683. doi:10.1371/journal.pmed.1002683
- Obermeyer Z, Emanuel EJ. Predicting the future: big data, machine learning, and clinical medicine. N Engl J Med. 2016;375(13):1216-9. doi:10.1056/NEJMp1606181
- Altman N, Krzywinski M. Association, correlation and causation. Nat Methods. 2015;12(10):899-900. doi:10.1038/nmeth.3587
- Tennant PWG, Murray EJ, Arnold KF, Berrie L, Fox MP, Gadd SC, et al. Use of directed acyclic graphs (DAGs) to identify confounders in applied health research: review and recommendations. Int J Epidemiol. 2021;50(2):620-32. doi:10.1093/ije/dyaa213
- Messerli FH. Chocolate consumption, cognitive function, and Nobel laureates. N Engl J Med. 2012;367(16):1562-4. doi:10.1056/NEJMon1211064
- Bothwell LE, Greene JA, Podolsky SH, Jones DS. Assessing the gold standard: lessons from the history of RCTs. N Engl J Med. 2016;374(22):2175-81. doi:10.1056/NEJMms1604593
- Hernán MA. The C-word: scientific euphemisms do not improve causal inference from observational data. Am J Public Health. 2018;108(5):616-9. doi:10.2105/AJPH.2018.304337
- Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. Am J Epidemiol. 2016;183(8):758-64. doi:10.1093/aje/kwv254
- Grosicki GJ, Kim J, Fielding F, Jasinski SR, Chapman C, Hippel WV, et al. Heart and health behavior responses to GLP-1 receptor agonists: a 12-wk study using wearable technology and causal inference. Am J Physiol Heart Circ Physiol. 2025;328(2):H235-H241. doi:10.1152/ajpheart.00809.2024
- Sanchez P, Voisey JP, Xia T, Watson HI, O’Neil AQ, Tsaftaris SA. Causal machine learning for healthcare and precision medicine. R Soc Open Sci. 2022;9(8):220638. doi:10.1098/rsos.220638
Useful Resources
- Causal Inference Book (Harvard T.H. Chan School of Public Health): a free textbook on causal methods for observational data.
- NIH Clinical Research Trials and You: The Basics: how randomization and clinical trials work.
- FDA: Artificial Intelligence-Enabled Medical Devices: the AI tools cleared for use in medicine.
- NIST AI Risk Management Framework: federal guidance on judging whether an AI system can be trusted.
FAQs
Is machine learning useless for health decisions?
No. It shines at prediction, such as flagging high-risk patients or reading scans, as Obermeyer and Emanuel explain. However, trouble starts when someone treats a prediction as a cause. Our post on pulse wave velocity shows prediction used well.
What is a confounding variable?
A confounder is a factor that influences both a supposed cause and its outcome, creating a link between them. That is why drawing a directed acyclic graph helps you find confounders before analysis.
Can observational data ever show causation?
Yes, with careful design and honest assumptions. Methods such as propensity score matching and target trial emulation make the comparison fair on measured traits. Still, unmeasured confounders remain the weak point.
Why are randomized controlled trials the gold standard?
Because chance, not choice, decides who gets treatment, the groups tend to differ only in the treatment itself. The National Institutes of Health guide to clinical trials explains randomization in plain terms.
What is causal machine learning?
It pairs the pattern-finding power of machine learning with causal reasoning, so models can estimate what a change would do. For a deeper look, this review of causal machine learning shows early uses in medicine.






Leave a Reply