Computational Provenance & Reproducibility Record

RAISINS - R and AI Solutions for INferential Statistics · Online Statistical Analysis Platform for Agricultural Research

Computational Provenance & Reproducibility Record Propensity Score Matching · 2.0.0 · DOI 10.5281/zenodo.23013711

Computational Provenance & Reproducibility Record

RAISINS · Propensity Score Matching Module

This Computational Provenance Record documents the statistical computing environment, software dependencies, computational provenance, and bibliographic references associated with the RAISINS Propensity Score Matching module. It is intended to support computational reproducibility and software transparency. Detailed statistical methodology, mathematical derivations, and user guidance are provided separately in the official module documentation.

Any issues or updates required please comment here

Go to the App from here

1 Module Metadata

Parameter Specification
Module Propensity Score Matching
Module Version 2.0.0
DOI 10.5281/zenodo.23013711
Document Type Computational workflow
Statistical Engine R
R Version 4.5.2
Reproducibility Execution Environment: Posit Connect · GCR · renv-locked

2 Statistical Dependency Manifest

Package Version Repository Core Statistical Functions
MatchIt 4.7.2 CRAN matchit(), match.data(), summary.matchit()
optmatch 0.10.8 CRAN Solver used by matchit(method = "optimal")
cobalt 4.6.2 CRAN love.plot(), bal.plot()
marginaleffects 0.32.0 CRAN avg_comparisons()
sandwich 3.1-1 CRAN Cluster-robust variance used by avg_comparisons(vcov = ~subclass)
margins 0.3.28 CRAN margins()
broom 1.0.13 CRAN tidy()
stats 4.5.2 Base R glm(), lm(), predict(), logLik(), pchisq(), t.test()

3 Statistical Function Registry

Analytical Role Primary Function(s)
Propensity-score model stats::glm(treatment ~ covariates, family = binomial)
Coefficients, odds ratios, marginal effects broom::tidy(), margins::margins()
Model fit (log-likelihood, LR chi-square, McFadden R²) stats::logLik(), stats::pchisq()
Matching MatchIt::matchit(distance = "glm")
Matched sample MatchIt::match.data()
Covariate balance before and after matching MatchIt::summary.matchit()
Love plot and propensity-score balance plot cobalt::love.plot(), cobalt::bal.plot()
Treatment effect on the treated (matched sample) stats::lm() + marginaleffects::avg_comparisons()
Welch and paired t-tests (matched sample) stats::t.test()
IPWRA and AIPW (full sample) stats::glm(), stats::lm(); bootstrap standard errors

4 Default Methods & Parameters

Analysis Step / Parameter Default Method / Value
Propensity
Score Model
Model Logistic regression of the binary treatment on the selected covariates, entered additively; the outcome is never used as a covariate
Treatment coding Exactly two levels required; the second level (in sorted order) is the treated group
Covariates Qualitative covariates are treated as factors; quantitative covariates as numeric
Matching Method Nearest-neighbour (default) or optimal, 1:1, without replacement, distance = "glm"
Caliper Optional, 0.1–0.5 standard deviations of the propensity score; nearest-neighbour matching only
Balance Standardised mean difference, variance ratio and eCDF statistics before and after matching. |SMD| < 0.1 well balanced; 0.1–0.25 borderline; > 0.25 imbalanced
Diagnostics Before vs after McFadden pseudo R², LR chi-square and mean standardised bias, from the original model and from the same model refitted on the matched sample
Treatment
Effect
Matched-sample ATT Linear outcome model with treatment × covariate interactions on the matched sample; ATT from avg_comparisons() over the treated units, with standard errors clustered by matched pair
t-tests Welch two-sample t-test and paired t-test of matched-pair differences, on the matched sample
IPWRA and AIPW ATE and ATT on the full sample, as robustness checks; propensity scores bounded away from 0 and 1; standard errors from 500 bootstrap resamples, stratified by treatment, with a fixed seed

5 R Code for Key Analytical Steps

The code blocks below demonstrate the computation behind each reported result using the lalonde dataset supplied with the MatchIt package.

5.1 Propensity-Score Model

library(MatchIt)
data("lalonde", package = "MatchIt")
form <- treat ~ age + educ + race + married + nodegree + re74 + re75

ps_mod <- glm(form, data = lalonde, family = binomial)
broom::tidy(ps_mod)                       # coefficients, SE, p-values
exp(coef(ps_mod))                         # odds ratios
summary(margins::margins(ps_mod))         # average marginal effects

ll_null <- -ps_mod$null.deviance / 2
1 - as.numeric(logLik(ps_mod)) / ll_null  # McFadden pseudo R-squared

5.2 Matching and Covariate Balance

m.out <- matchit(form, data = lalonde,
                 method = "nearest",      # or "optimal"
                 distance = "glm", ratio = 1, replace = FALSE,
                 caliper = 0.2)           # optional; nearest-neighbour only

summary(m.out)$nn                         # matched / unmatched counts
summary(m.out)$sum.all                    # balance before matching
summary(m.out)$sum.matched                # balance after matching

cobalt::love.plot(m.out, stars = "std", abs = TRUE)
cobalt::bal.plot(m.out, var.name = "distance", which = "both")

5.3 Treatment Effect on the Matched Sample

library(marginaleffects)
m.dat <- match.data(m.out)

fit <- lm(re78 ~ treat * (age + educ + race + married + nodegree + re74 + re75),
          data = m.dat, weights = weights)

avg_comparisons(fit, variables = "treat",
                newdata = subset(m.dat, treat == 1),
                wts = "weights", vcov = ~subclass)   # ATT

t.test(re78 ~ treat, data = m.dat)                    # Welch t-test

5.4 IPWRA and AIPW (Full Sample)

X <- lalonde[, c("age", "educ", "race", "married", "nodegree", "re74", "re75")]
A <- lalonde$treat; Y <- lalonde$re78
e <- fitted(glm(A ~ ., data = cbind(A = A, X), family = binomial))

# IPWRA, ATT: treated weighted 1, controls weighted e/(1-e)
f1 <- lm(Y ~ ., data = cbind(Y = Y, X)[A == 1, ])
f0 <- lm(Y ~ ., data = cbind(Y = Y, X)[A == 0, ], weights = (e / (1 - e))[A == 0])
mean((predict(f1, X) - predict(f0, X))[A == 1])

# AIPW, ATE
m1 <- predict(lm(Y ~ ., data = cbind(Y = Y, X)[A == 1, ]), X)
m0 <- predict(lm(Y ~ ., data = cbind(Y = Y, X)[A == 0, ]), X)
mean(A * (Y - m1) / e - (1 - A) * (Y - m0) / (1 - e) + (m1 - m0))

Explore the entire Propensity Score Matching module in preview mode using our demo datasets. To submit suggestions or report a workflow issue, please use the discussion section below, or visit the official RAISINS website.

6 RAISINS Native Statistical Framework

RAISINS uses R for all its statistical computations. Every package used to generate major results is listed and demonstrated with examples, so results can be reproduced independently. These results are then organized and formatted on the RAISINS website along with visualisation to make them easier to use and interpret. RAISINS also has its own custom-built statistical tools for managing workflows, validating results, and generating reports. Details of these are not fully covered here, they’re shared with outside researchers only on request, and are subject to licensing terms.

7 Package References

R Core Team. (2025). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/

Ho, D., Imai, K., King, G., & Stuart, E. A. (2011). MatchIt: Nonparametric Preprocessing for Parametric Causal Inference. Journal of Statistical Software, 42(8), 1-28. https://doi.org/10.18637/jss.v042.i08

Hansen, B. B., & Klopfer, S. O. (2006). Optimal full matching and related designs via network flows. Journal of Computational and Graphical Statistics, 15(3), 609-627. https://doi.org/10.1198/106186006X137047

Greifer, N. (2024). cobalt: Covariate Balance Tables and Plots (R package version 4.6.2). https://doi.org/10.32614/CRAN.package.cobalt

Arel-Bundock, V., Greifer, N., & Heiss, A. (2024). How to Interpret Statistical Models Using marginaleffects for R and Python. Journal of Statistical Software, 111(9), 1-32. https://doi.org/10.18637/jss.v111.i09

Zeileis, A., Köll, S., & Graham, N. (2020). Various Versatile Variances: An Object-Oriented Implementation of Clustered Covariances in R. Journal of Statistical Software, 95(1), 1-36. https://doi.org/10.18637/jss.v095.i01

Leeper, T. J. (2024). margins: Marginal Effects for Model Objects (R package version 0.3.28). https://doi.org/10.32614/CRAN.package.margins

Robinson, D., Hayes, A., & Couch, S. (2025). broom: Convert Statistical Objects into Tidy Tibbles (R package version 1.0.13). https://CRAN.R-project.org/package=broom

Rosenbaum, P. R., & Rubin, D. B. (1983). The Central Role of the Propensity Score in Observational Studies for Causal Effects. Biometrika, 70(1), 41-55. https://doi.org/10.1093/biomet/70.1.41

Austin, P. C. (2011). An Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies. Multivariate Behavioral Research, 46(3), 399-424. https://doi.org/10.1080/00273171.2011.568786

Robins, J. M., Rotnitzky, A., & Zhao, L. P. (1994). Estimation of Regression Coefficients When Some Regressors Are Not Always Observed. Journal of the American Statistical Association, 89(427), 846-866. https://doi.org/10.1080/01621459.1994.10476818

Allaire, J. J., Xie, Y., Dervieux, C., McPherson, J., Luraschi, J., Ushey, K., Atkins, A., Wickham, H., Cheng, J., Chang, W., & Iannone, R. (2026). rmarkdown: Dynamic Documents for R (R package version 2.31). https://github.com/rstudio/rmarkdown

Feedback & Discussion