Computational Provenance & Reproducibility Record

RAISINS - R and AI Solutions for INferential Statistics · Online Statistical Analysis Platform for Agricultural Research

Computational Provenance & Reproducibility Record Endogenous Switching Regression · 1.0.0 · DOI 10.5281/zenodo.23114309

Computational Provenance & Reproducibility Record

RAISINS · Endogenous Switching Regression Module

This Computational Provenance Record documents the statistical computing environment, software dependencies, computational provenance, and bibliographic references associated with the RAISINS Endogenous Switching Regression module. It is intended to support computational reproducibility and software transparency. Detailed statistical methodology, mathematical derivations, and user guidance are provided separately in the official module documentation.

Any issues or updates required please comment here

Go to the App from here

1 Module Metadata

Parameter Specification
Module Endogenous Switching Regression
Module Version 1.0.0
DOI 10.5281/zenodo.23114309
Document Type Computational workflow
Statistical Engine R
R Version 4.5.2
Reproducibility Execution Environment: Posit Connect · GCR · renv-locked

2 Statistical Dependency Manifest

Package Version Repository Core Statistical Functions
endoSwitch 1.0.0 GitHub (cbw1243/endoSwitch@d572d57) endoSwitch2Stage(), endoSwitch(), summary.endoSwitch(), treatmentEffect()
maxLik 1.5-2.2 CRAN maxLik(method = "BFGS"), the optimiser called by endoSwitch()
msm 1.8.2 CRAN deltamethod(), standard errors of the back-transformed σ and ρ
data.table 1.18.6.1 CRAN as.data.table(), the data structure required by endoSwitch
ivreg 0.6-8 CRAN ivreg(), summary(diagnostics = TRUE)
car 3.1-5 CRAN vif(), leveneTest()
stats 4.5.2 Base R glm(), lm(), logLik(), pchisq(), pnorm(), dnorm(), qnorm(), t.test(), prop.test(), shapiro.test()

3 Statistical Function Registry

Analytical Role Primary Function(s)
Instrument falsification test stats::glm(family = binomial("probit")), stats::lm() on non-adopters
Endogeneity test (weak instruments, Wu-Hausman, Sargan) ivreg::ivreg() + summary(diagnostics = TRUE)
Adopter vs non-adopter characteristics stats::prop.test(), car::leveneTest(center = median), stats::t.test()
Multicollinearity car::vif() on stats::lm() of adoption on all covariates
Two-stage estimation (probit + OLS with inverse Mills ratio) endoSwitch::endoSwitch2Stage()
Probit average marginal effects stats::pnorm(), stats::dnorm(), stats::vcov(); delta-method standard errors
Full information maximum likelihood endoSwitch::endoSwitch() → maxLik::maxLik()
σ and ρ with standard errors endoSwitch:::calcPar() → msm::deltamethod()
LR test of independent equations stats::logLik(), stats::pchisq(df = 2)
Treatment effects (ATT, ATU, heterogeneity) summary(fit)$treatmentEffect; household-level values from endoSwitch:::treatmentEffect(treatEffect = FALSE)
Residual normality stats::shapiro.test() on the second-stage OLS residuals

4 Default Methods & Parameters

Analysis Step / Parameter Default Method / Value
Model
Specification
Adoption variable Binary, coded 0 (did not adopt) / 1 (adopted); both values must be present
Instrument(s) Chosen separately; added automatically to the selection equation and never to the outcome equation (exclusion restriction)
Covariates Selection equation = selection covariates + instruments; outcome equation = outcome covariates, identical in both regimes
Text covariates Trimmed, converted to upper case and expanded into 0/1 dummy columns with model.matrix(); the first category is the reference
Sample Files with a completely blank row or column are rejected before analysis; remaining rows with a missing value in any selected column are dropped (stats::na.omit() inside endoSwitch()). A warning is shown below 200 complete observations, because two-step estimates become unstable in small samples (Bascle, 2008)
Transformation Log transform (optional, off by default) Base-10 log transformation of the selected numeric columns, used to reduce skewness so that effects on the outcome can be read as percentage changes. Columns with any zero or negative value are first shifted so that their smallest value becomes 1. All later steps use the transformed values
Pre-estimation
Tests
Falsification test An instrument is valid when it is significant (p < 0.05) in the adoption probit on all selection covariates, and not significant (p ≥ 0.05) in the OLS of the outcome for non-adopters on all instruments jointly plus the outcome covariates (Di Falco, Veronesi & Yesuf, 2011)
Endogeneity test IV regression of the outcome on adoption + outcome covariates, adoption instrumented by the instruments. Reports the weak-instrument F-test, the Wu-Hausman test and the Sargan test (Sargan requires two or more instruments)
Group differences 0/1 variables: two-sample test of proportions (prop.test(), continuity-corrected; reported as "Chi-square test"). Continuous variables: Student's t-test if the median-centred Levene's test gives p ≥ 0.05, otherwise Welch's t-test
Multicollinearity VIF from an OLS of adoption on all selection and outcome covariates; < 5 acceptable, 5–10 moderate, ≥ 10 high. Perfectly collinear (aliased) covariates stop the analysis
Two-Stage
Estimation
First stage Probit (binomial(link = "probit")) of adoption on the selection covariates
Second stage Separate OLS regressions of the outcome for adopters and for non-adopters on the outcome covariates. Each regression also includes a selection-correction term (the inverse Mills ratio) from the first-stage probit, which captures unobserved factors that affect both the decision to adopt and the outcome. The OLS standard errors are not adjusted for the fact that this term is itself an estimate
Marginal effects Average marginal effects of the probit, which express each coefficient as a change in the probability of adoption. For a continuous variable, this is the average effect of a one-unit increase. For a 0/1 variable, it is the average change in probability when the variable switches from 0 to 1. Standard errors use the delta method
Full MLE Optimiser maxLik() with method = "BFGS" (package default), unweighted (Weight = NA); return code 0 = converged
Starting values The coefficients, error standard deviations (σ) and correlations (ρ) from the two-stage estimation (endoSwitch2Stage()) are used as starting values. A negative two-stage error variance is replaced by 0.1
σ and ρ σ is the standard deviation of the outcome-equation error in each regime. ρ is the correlation between the adoption-equation error and each outcome-equation error, and its sign shows the direction of self-selection. Both are estimated on an unrestricted scale, which keeps σ positive and ρ between −1 and 1, and are then converted back. Standard errors use the delta method
LR test Likelihood-ratio test of whether the adoption and outcome equations are independent, i.e. whether both correlations are zero. It compares the full ESR model with separate probit and OLS models fitted to the same observations, using a chi-square test with 2 degrees of freedom. p < 0.05 is reported as selection bias present
Treatment
Effects
Conditional expectations Four expected outcomes are predicted for every household from the full MLE estimates: adopters' actual outcome, adopters' counterfactual outcome had they not adopted, non-adopters' actual outcome, and non-adopters' counterfactual outcome had they adopted. Each prediction combines the household's own characteristics, the coefficients of the relevant regime and a selection-correction term (Lokshin & Sajaia, 2004)
ATT, ATU ATT is the average difference between adopters' actual and counterfactual outcomes, i.e. the gain adopters obtained from adopting. ATU is the same average difference for non-adopters, i.e. the gain they would have obtained had they adopted. Standard errors come from a paired t-test of the two sets of household-level predictions
Heterogeneity Base heterogeneity (Y1 and Y0 columns) shows whether adopters and non-adopters differ systematically in outcome under the same adoption status, i.e. whether adopters are inherently better or worse off. Transitional heterogeneity shows whether adoption benefits adopters more or less than it would benefit non-adopters. Standard errors come from Welch two-sample t-tests. Values in parentheses under Y1 / Y0 are standard deviations of the household-level predictions

5 R Code for Key Analytical Steps

The code blocks below demonstrate the computation behind each reported result using the ImpactData dataset supplied with the endoSwitch package (adoption of conservation agriculture by 408 farm households in Zambia).

5.1 Data and Model Specification

library(endoSwitch)
data("ImpactData", package = "endoSwitch")
d <- as.data.frame(ImpactData)
d$Output <- log10(d$Output)                            # optional log transformation

adopt   <- "CA"                                        # adoption (0/1)
yvar    <- "Output"                                    # outcome
instr   <- c("Distance_to_market", "Association")      # instruments
sel_cov <- unique(c("Age", "Education", "Farm_size", "Access_to_credit",
                    "Perception", instr))              # instruments added to selection
out_cov <- c("Age", "Education", "Farm_size", "Fertilizer_ha", "Access_to_credit")
d <- na.omit(d[, unique(c(adopt, yvar, sel_cov, out_cov))])

5.2 Instrument Falsification and Endogeneity Tests

# relevance: instruments in the adoption probit (valid if p < 0.05)
f1 <- glm(reformulate(sel_cov, adopt), family = binomial("probit"), data = d)
coef(summary(f1))[instr, ]

# exclusion: instruments in the outcome OLS of non-adopters (valid if p >= 0.05)
f2 <- lm(reformulate(c(instr, out_cov), yvar), data = d[d[[adopt]] == 0, ])
coef(summary(f2))[instr, ]

# IV diagnostics: weak instruments, Wu-Hausman, Sargan
iv <- ivreg::ivreg(as.formula(paste(yvar, "~", paste(c(adopt, out_cov), collapse = " + "), "|",
                                    paste(c(instr, out_cov), collapse = " + "))), data = d)
summary(iv, diagnostics = TRUE)$diagnostics

5.3 Group Differences and Multicollinearity

grp <- factor(d[[adopt]] == 1)

x  <- d$Age                                            # continuous variable
eq <- car::leveneTest(x ~ grp, center = median)$`Pr(>F)`[1] >= 0.05
t.test(x[grp == "TRUE"], x[grp == "FALSE"], var.equal = eq)   # Student's or Welch's

b <- d$Access_to_credit                                # 0/1 variable
prop.test(c(sum(b[grp == "TRUE"]), sum(b[grp == "FALSE"])),
          c(sum(grp == "TRUE"), sum(grp == "FALSE")))

car::vif(lm(reformulate(unique(c(sel_cov, out_cov)), adopt), data = d))

5.4 Two-Stage Estimation and Marginal Effects

ts <- endoSwitch2Stage(d, yvar, adopt, out_cov, sel_cov)
summary(ts$FirstStageReg)                              # probit selection equation
summary(ts$SecondStageReg.0)                           # non-adopters + Mills ratio
summary(ts$SecondStageReg.1)                           # adopters + Mills ratio
ts$distParEst                                          # sigma0, sigma1, rho0, rho1

# average marginal effect of a continuous regressor, delta-method SE
X  <- model.matrix(ts$FirstStageReg); b <- coef(ts$FirstStageReg); V <- vcov(ts$FirstStageReg)
xb <- as.vector(X %*% b); k <- "Farm_size"
me   <- mean(dnorm(xb)) * b[[k]]
grad <- -b[[k]] * colMeans(xb * dnorm(xb) * X); grad[k] <- grad[k] + mean(dnorm(xb))
c(me = me, se = sqrt(as.numeric(t(grad) %*% V %*% grad)))

5.5 Full Information Maximum Likelihood and LR Test

fit <- endoSwitch(d, yvar, adopt, out_cov, sel_cov, Weight = NA, treatEffect = TRUE)
s <- summary(fit)
s$estimate                                             # Select.*, Outcome.0.*, Outcome.1.*, Sigma, Rho
c(s$maxinType, s$loglik, s$returnCode)                 # BFGS; return code 0 = converged

ll_ind <- as.numeric(logLik(f1) +
  logLik(lm(reformulate(out_cov, yvar), data = d[d[[adopt]] == 0, ])) +
  logLik(lm(reformulate(out_cov, yvar), data = d[d[[adopt]] == 1, ])))
lr <- 2 * (s$loglik - ll_ind)
pchisq(lr, df = 2, lower.tail = FALSE)                 # H0: rho0 = rho1 = 0

5.6 Treatment Effects and Residual Normality

s$treatmentEffect                                      # ATT, ATU, heterogeneity

hp <- endoSwitch:::treatmentEffect(fit$MLE.Results, data.table::as.data.table(d),
                                   yvar, adopt, out_cov, sel_cov, FALSE)
mean(hp$EYA1$EY1.A1 - hp$EYA1$EY0.A1)                  # ATT
mean(hp$EYA0$EY1.A0 - hp$EYA0$EY0.A0)                  # ATU
t.test(hp$EYA1$EY1.A1, hp$EYA1$EY0.A1, paired = TRUE)$stderr   # SE of ATT

shapiro.test(residuals(ts$SecondStageReg.0))           # non-adopters
shapiro.test(residuals(ts$SecondStageReg.1))           # adopters

Explore the entire Endogenous Switching Regression module in preview mode using our demo datasets. To submit suggestions or report a workflow issue, please use the discussion section below, or visit the official RAISINS website.

6 RAISINS Native Statistical Framework

RAISINS uses R for all its statistical computations. Every package used to generate major results is listed and demonstrated with examples, so results can be reproduced independently. These results are then organized and formatted on the RAISINS website along with visualisation to make them easier to use and interpret. RAISINS also has its own custom-built statistical tools for managing workflows, validating results, and generating reports. Details of these are not fully covered here, they’re shared with outside researchers only on request, and are subject to licensing terms.

7 Package References

R Core Team. (2025). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/

Chen, B., Yun, S., & Gramig, B. (n.d.). endoSwitch: Endogenous Switching Regression Models (R package version 1.0.0, GitHub commit d572d57). https://github.com/cbw1243/endoSwitch

Henningsen, A., & Toomet, O. (2011). maxLik: A package for maximum likelihood estimation in R. Computational Statistics, 26(3), 443-458. https://doi.org/10.1007/s00180-010-0217-1

Jackson, C. H. (2011). Multi-State Models for Panel Data: The msm Package for R. Journal of Statistical Software, 38(8), 1-29. https://doi.org/10.18637/jss.v038.i08

Barrett, T., Dowle, M., Srinivasan, A., Gorecki, J., Chirico, M., Hocking, T., Schwendinger, B., & Krylov, I. (2025). data.table: Extension of ‘data.frame’ (R package version 1.18.6.1). https://doi.org/10.32614/CRAN.package.data.table

Fox, J., Kleiber, C., & Zeileis, A. (2025). ivreg: Instrumental-Variables Regression by ‘2SLS’, ‘2SM’, or ‘2SMM’, with Diagnostics (R package version 0.6-8). https://doi.org/10.32614/CRAN.package.ivreg

Fox, J., & Weisberg, S. (2019). An R Companion to Applied Regression (3rd ed.). Sage, Thousand Oaks, CA. https://www.john-fox.ca/Companion/

Heckman, J. J. (1979). Sample Selection Bias as a Specification Error. Econometrica, 47(1), 153-161. https://doi.org/10.2307/1912352

Lokshin, M., & Sajaia, Z. (2004). Maximum likelihood estimation of endogenous switching regression models. The Stata Journal, 4(3), 282-289. https://doi.org/10.1177/1536867X0400400306

Di Falco, S., Veronesi, M., & Yesuf, M. (2011). Does Adaptation to Climate Change Provide Food Security? A Micro-Perspective from Ethiopia. American Journal of Agricultural Economics, 93(3), 829-846. https://doi.org/10.1093/ajae/aar006

Bascle, G. (2008). Controlling for endogeneity with instrumental variables in strategic management research. Strategic Organization, 6(3), 285-327. https://doi.org/10.1177/1476127008094339

Allaire, J. J., Xie, Y., Dervieux, C., McPherson, J., Luraschi, J., Ushey, K., Atkins, A., Wickham, H., Cheng, J., Chang, W., & Iannone, R. (2025). rmarkdown: Dynamic Documents for R (R package version 2.30). https://github.com/rstudio/rmarkdown

Feedback & Discussion