SCIENCEYE BLOG

Contact

Propensity Score Matching: Causality Without a Trial

6 min read

Propensity Score Matching PSM
Propensity Score Matching: Causality Without a Trial

Some of the most important questions in health will never get a randomized trial. Nobody will randomize people to decades of shift work, a caregiving role or a new tax on sugary drinks. Yet policymakers still need causal answers, and observational research is often the only evidence on the table. Regulators now weigh real-world evidence from health records and insurance claims when they judge medical products.8

Propensity score matching (PSM) is one of the most widely used tools for turning those data into causal claims. Used rigorously, it approximates the trial you could never run.5 Used carelessly, it gives confident answers to the wrong question. Here’s what separates the two.

What you’ll learnWhat you’ll be able to do
How PSM builds a trial-like comparisonExplain what matching can and cannot balance
How to check whether matching workedRead and judge a covariate balance table
How to handle unmeasured confoundersReport a sensitivity analysis such as the E-value
Learning objectives for this article.

How Propensity Score Matching Mimics a Randomized Trial

How propensity score matching builds balanced groups Four steps left to right: baseline covariates, propensity score, matched pairs, balanced groups on measured covariates only. From many covariates to one score to balanced groups Baselinecovariates Propensity score(probability oftreatment) Matchedtreated/untreated pairs Balanced groups(measured only) Unmeasured confounders pass through every step unbalanced.
Figure 1. The propensity score collapses many measured covariates into one number that matching can balance.1,2

Rosenbaum and Rubin defined the propensity score as each person’s probability of receiving treatment, given their observed baseline characteristics.1 Their key result is elegant. People with the same score share, on average, the same distribution of those measured covariates, whether or not they were treated. So instead of matching on 30 variables at once, you match on one number.

Think of it as a seating chart. Randomization shuffles everyone into two rooms, so the rooms look alike on everything, measured or not. Matching seats each treated person beside an untreated person with a near-identical profile. The rooms end up alike, but only on the traits written on the chart.2 Matching also changes the question: most matched designs estimate the effect in people who were treated, so state your target estimand before you match.6

Matching balances what you measured. It can’t touch what you didn’t.

Design First, Then Check Covariate Balance

For causality claims, Rubin’s rule is simple and often ignored: build the matched groups before you look at any outcome data.7 That keeps the design honest, because you can’t tune the match until the result looks the way you hoped. Treat it like a trial protocol and write it down first.

Next, prove that propensity score matching worked. A t-test p-value tells you little here, since it rises and falls with sample size. Compare groups with standardized differences, variance ratios and plots of each covariate’s distribution instead.3 Comparing the propensity score itself across groups isn’t enough, because that comparison says almost nothing about covariate balance.3

DiagnosticWhat it checksCommon benchmark
Standardized mean differenceDifference in each covariate’s mean or proportion, in standard deviation unitsBelow 0.1
Variance ratioWhether continuous covariates have similar spread in both groupsClose to 1
Distribution plotsThe whole shape of each covariate, not just its averageClear overlap across the range
Table 1. Balance diagnostics for propensity score matched samples.2,3

A PSM study is only as credible as the balance table it reports.

Unmeasured Confounders and Real-World Evidence

No match can balance a confounder nobody recorded. The honest response is to measure how strong that hidden confounder would need to be. The E-value does exactly this. It gives the minimum risk ratio an unmeasured confounder would need with both treatment and outcome to explain away the result.4

Does rigor pay off? When researchers emulated 32 clinical trials with claims data and propensity score matching, the results tracked the trials closely (r = 0.82).10 Agreement rose to 0.93 when researchers could closely emulate the trial design and fell to 0.53 when they couldn’t.10 An earlier emulation of 10 cardiovascular trials told a similar story.9

Agreement between trial results and database emulations Bar chart of Pearson correlations: 0.82 for all 32 emulated trials, 0.93 for 16 closely emulated trials, 0.53 for 16 less closely emulated trials. r = 0.82 All trials n = 32 r = 0.93 Close emulation n = 16 r = 0.53 Less close emulation n = 16
Figure 2. Correlation between randomized trial results and propensity score matched database emulations of the same trials.10

Close emulation of the trial you wish you had is what earns the causal claim.

Conclusion

Propensity score matching doesn’t create causality from data. It creates a fair comparison. You earn the causal claim by showing your work: fix the design before seeing outcomes, prove covariate balance and report how much hidden confounding would overturn the result. Do that, and observational research becomes real-world evidence that can guide policy.8,10

“Observational studies can and should be designed to approximate randomized experiments as closely as possible.”

Donald B. Rubin, Statistics in Medicine, 20077

Reviewing a PSM study this week? Before you trust it, ask for three things: the prespecified design, the balance table and the E-value.

Useful Resources

FAQs

Is propensity score matching better than regression adjustment?

Not automatically. Both rely on the same assumption of no unmeasured confounding, but matching separates design from analysis and makes balance visible in a table. Austin (2011) compares the approaches directly.

What standardized difference counts as balanced?

A common benchmark is below 0.1 for every covariate, checked alongside variance ratios and plots. Austin (2009) explains why these diagnostics beat significance tests.

Can PSM prove causality?

No single method can. PSM makes a causal claim credible only if its assumptions hold, which is why the E-value should accompany the estimate.

How does PSM relate to target trial emulation?

Target trial emulation is the design framework, and PSM is one way to adjust for confounding within it. Hernán and Robins (2016) set out the framework.

Where does real-world evidence fit in policy?

Regulators increasingly use it to support decisions about medical products. The FDA real-world evidence program describes how.

Future Blog Topics

TopicWhy it matters
Matching or weighting? Choosing a propensity score methodInverse probability weighting keeps your full sample but brings its own trade-offs
Target trial emulation, step by stepA design checklist that prevents the most common observational biases
Reading an E-value like a reviewerTurns a sensitivity analysis into a judgment you can defend
Coming soon on the SciencEye blog.

References

  1. Rosenbaum PR, Rubin DB. The central role of the propensity score in observational studies for causal effects. Biometrika. 1983;70(1):41-55. doi:10.1093/biomet/70.1.41
  2. Austin PC. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate Behav Res. 2011;46(3):399-424. doi:10.1080/00273171.2011.568786
  3. Austin PC. Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples. Stat Med. 2009;28(25):3083-107. doi:10.1002/sim.3697
  4. VanderWeele TJ, Ding P. Sensitivity analysis in observational research: introducing the E-value. Ann Intern Med. 2017;167(4):268-74. doi:10.7326/M16-2607
  5. Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. Am J Epidemiol. 2016;183(8):758-64. doi:10.1093/aje/kwv254
  6. Stuart EA. Matching methods for causal inference: a review and a look forward. Stat Sci. 2010;25(1):1-21. doi:10.1214/09-STS313
  7. Rubin DB. The design versus the analysis of observational studies for causal effects: parallels with the design of randomized trials. Stat Med. 2007;26(1):20-36. doi:10.1002/sim.2739
  8. Sherman RE, Anderson SA, Dal Pan GJ, Gray GW, Gross T, Hunter NL, et al. Real-world evidence: what is it and what can it tell us? N Engl J Med. 2016;375(23):2293-7. doi:10.1056/NEJMsb1609216
  9. Franklin JM, Patorno E, Desai RJ, Glynn RJ, Martin D, Quinto K, et al. Emulating randomized clinical trials with nonrandomized real-world evidence studies: first results from the RCT DUPLICATE initiative. Circulation. 2021;143(10):1002-13. doi:10.1161/CIRCULATIONAHA.120.051718
  10. Wang SV, Schneeweiss S, Franklin JM, Desai RJ, Feldman W, Garry EM, et al. Emulation of randomized clinical trials with nonrandomized database analyses: results of 32 clinical trials. JAMA. 2023;329(16):1376-85. doi:10.1001/jama.2023.4221

© 2026 SciencEye. All rights reserved. | SciencEye.com

Leave a Reply

Your email address will not be published. Required fields are marked *