Computational Provenance & Reproducibility Record

RAISINS - R and AI Solutions for INferential Statistics · Online Statistical Analysis Platform for Agricultural Research

Computational Provenance & Reproducibility Record Pooled Strip Plot · 2.0.0 · DOI 10.5281/zenodo.23081818

Computational Provenance & Reproducibility Record

RAISINS · Pooled Strip Plot Module

This Computational Provenance Record documents the statistical computing environment, software dependencies, computational provenance, and bibliographic references associated with the RAISINS Pooled Strip Plot module. It is intended to support computational reproducibility and software transparency. Detailed statistical methodology, mathematical derivations, and user guidance are provided separately in the official module documentation.

Any issues or updates required please comment here

Go to the App from here

1 Module Metadata

Parameter Specification
Module Pooled Strip Plot
Module Version 2.0.0
DOI 10.5281/zenodo.23081818
Document Type Computational workflow
Statistical Engine R
R Version 4.5.2
Design Row-strip factor (A) and column-strip factor (B) crossed in RCBD blocks nested within Location; pooled over Locations
Error strata Three — Pooled Error (a) = L:Block:A, Pooled Error (b) = L:Block:B, Pooled Error (c) = Residuals
Reproducibility Execution Environment: Posit Connect · GCR · renv-locked

2 Statistical Dependency Manifest

Package Version Repository Core Statistical Functions
stats 4.5.2 Base R lm(), aov(), bartlett.test(), pf(), residuals(), df.residual(), prcomp()
car 3.1-3 CRAN Anova()
agricolae 1.3-7 CRAN LSD.test(), HSD.test(), duncan.test()
gtools 3.9.5 CRAN mixedsort(), mixedorder()
dplyr 1.1.4 CRAN group_by(), summarise()
moments 0.14.1 CRAN skewness(), kurtosis()
factoextra 1.0.7 CRAN get_eigenvalue(), fviz_eig(), fviz_pca_biplot()
scales 1.4.0 CRAN rescale()
phia 0.3-2 CRAN interactionMeans()
GGally 2.4.0 CRAN ggpairs()
ggdist 3.3.3 CRAN stat_halfeye()

3 Statistical Function Registry

Analytical Role Primary Function(s)
Pooled strip-plot model fit stats::lm(Y ~ L + L:Block + A + L:A + L:Block:A + B + L:B + L:Block:B + A:B + L:A:B)
Sums of squares car::Anova(type = "II")
Pooled Error (a) — row-strip stratum row 8 of the Anova() table, L:Block:A
Pooled Error (b) — column-strip stratum row 9 of the Anova() table, L:Block:B
Pooled Error (c) — interaction stratum row 11 of the Anova() table, Residuals
F-test p-values against the assigned stratum stats::pf()
Homogeneity of error variances over Locations stats::bartlett.test() on the per-Location aov() residuals, labelled by Location
Aitken transformation for heterogeneous errors per-Location division by sqrt(MSE)
Post-hoc mean comparison and letter grouping agricolae::LSD.test(), agricolae::HSD.test(), agricolae::duncan.test()
Natural ordering of treatment labels gtools::mixedsort(), gtools::mixedorder()
Summary statistics by Location × Row × Column dplyr::group_by(), dplyr::summarise(), moments::skewness(), moments::kurtosis()
Principal component analysis on Location × Row × Column means stats::prcomp(center = TRUE, scale = TRUE)
PCA eigenvalues, scree plot, biplot, index scores factoextra::get_eigenvalue(), factoextra::fviz_eig(), factoextra::fviz_pca_biplot(), scales::rescale()
Interaction plot means phia::interactionMeans()

4 Default Methods & Parameters

Analysis Step / Parameter Default Method / Value
ANOVA
Structure
Model and sums of squares A single fit, lm(Y ~ L + L:Block + A + L:A + L:Block:A + B + L:B + L:Block:B + A:B + L:A:B), with sums of squares from car::Anova(type = "II"). Blocks are nested within Location. The ANOVA is run on the Aitken-transformed values for any character that failed Bartlett's test
Pooled Error (a) L:Block:A, df = Locations × (Blocks − 1) × (Row levels − 1). Tests Row treatments (A) and Location × Row
Pooled Error (b) L:Block:B, df = Locations × (Blocks − 1) × (Column levels − 1). Tests Column treatments (B) and Location × Column
Pooled Error (c) Residuals, df = Locations × (Blocks − 1) × (Row levels − 1) × (Column levels − 1). Tests the Location main effect, Location × Block, Row × Column and Location × Row × Column
Homogeneity
& Pooling
Bartlett's test Per character. Within each Location, aov(Y ~ Block + FactorA + Block:FactorA + FactorB + Block:FactorB + FactorA:FactorB) is fitted, so the residuals are the strip-plot interaction error; the residuals from all Locations are combined, labelled by Location, and tested with bartlett.test(). At least 2 Locations with ≥ 2 levels each of Block, FactorA and FactorB are required
Aitken transformation Applied when Bartlett's p ≤ 0.05. Within each Location, the character is divided by sqrt(MSE), where MSE = sum(resid(m)^2) / df.residual(m) from the model above
User
Transformation

(optional, off by default)
Log log10(x); if any value ≤ 0, log10(x - min(x) + 1)
Square root sqrt(x); sqrt(x + 0.5) if any value is 0; not applied if any value is negative
Arcsine asin(sqrt(x)) for proportions in [0, 1]; 0 and 1 are replaced by 1/(4n) and 1 - 1/(4n); not applied if any value is outside [0, 1]
Multiple
Comparison
Procedure LSD (default), Tukey's HSD or DMRT, at α = 0.05 (default) or 0.01
Error term used Pooled Error (a) for Row and Location × Row; Pooled Error (b) for Column and Location × Column; Pooled Error (c) for Location, Row × Column and Location × Row × Column
Critical difference Column 6 of $statistics for LSD, column 5 for Tukey's HSD; for DMRT, the critical ranges from $duncan. CD and letters are shown only when the effect is significant at α
SE(m), SE(d), CV (%) sqrt(MSerror / r), sqrt(2 * MSerror / r) and $statistics[1, 4] from the corresponding agricolae call
Summary
Statistics
Per cell For each Location × Row × Column combination: N, Mean, SD, SE = SD / √N, Min, Max, CV, moments::skewness(), moments::kurtosis()
Principal
Component
Analysis
Input Character means per Location × Row × Column combination, ordered by gtools::mixedorder(); prcomp(center = TRUE, scale = TRUE); requires at least two response variables
Index scores PC signs are reversed after fitting ($x and $rotation multiplied by −1). Index 1 and 2 are the PC1 and PC2 scores, also rescaled to [0, 1]; the selection cutoff defaults to 0.75
Presentation Decimal places, font 2 decimal places (1–4); table font Cambria. Display only

5 R Code for Key Analytical Steps

The chunks below reproduce the module’s computations on datasets::CO2. Keeping three of its seven CO2 concentrations gives a balanced 2 Locations × 3 Blocks × 2 Row × 3 Column = 36-observation layout with the structure the module expects. The numbers in the comments are the actual outputs.

5.1 Model Dataset

data(CO2, package = "datasets")
d <- subset(CO2, conc %in% c(175, 350, 675))
d <- data.frame(
  Location   = factor(d$Type, labels = c("Quebec", "Mississippi")),
  FactorA    = factor(d$Treatment, labels = c("A_nonchilled", "A_chilled")),   # row strips
  FactorB    = factor(paste0("B_", d$conc)),                                     # column strips
  Block      = factor(substr(as.character(d$Plant), 3, 3)),
  Uptake     = d$uptake,
  Efficiency = d$uptake / d$conc * 100
)
nrow(d)                                                         # 36

5.2 Homogeneity of Error Variances and the Aitken Transformation

d_tr <- d
for (ch in c("Uptake", "Efficiency")) {
  fits <- lapply(split(d, d$Location), function(sub) {
    sub <- droplevels(sub)
    aov(as.formula(paste(ch, "~ Block + FactorA + Block:FactorA + FactorB + Block:FactorB + FactorA:FactorB")),
        data = sub)
  })
  mse  <- sapply(fits, function(m) sum(resid(m)^2) / df.residual(m))
  res  <- lapply(fits, residuals)
  test <- bartlett.test(unlist(res), rep(names(fits), sapply(res, length)))
  if (test$p.value <= 0.05) {
    for (loc in names(mse)) {
      rows <- d$Location == loc
      d_tr[rows, ch] <- d[rows, ch] / sqrt(mse[loc])
    }
  }
}
#> Uptake:     chisq = 33.09  p < 0.0001  MSE Quebec = 4.38, Mississippi = 0.16  -> Aitken applied
#> Efficiency: chisq = 4.06   p = 0.0439  MSE Quebec = 1.50, Mississippi = 0.55  -> Aitken applied

5.3 Pooled ANOVA with Three Error Strata

library(car)
L <- d_tr$Location; A <- d_tr$FactorA; B <- d_tr$FactorB; Block <- d_tr$Block
y <- d_tr$Uptake

fit <- lm(y ~ L + L:Block + A + L:A + L:Block:A + B + L:B + L:Block:B + A:B + L:A:B)
at  <- as.data.frame(car::Anova(fit, type = "II"))
rownames(at)
#> "L" "A" "B" "L:Block" "L:A" "L:B" "A:B" "L:Block:A" "L:Block:B" "L:A:B" "Residuals"

MS <- at[, "Sum Sq"] / at[, "Df"]
Ea <- 8; Eb <- 9; Ec <- 11                       # Pooled Error (a), (b), (c)
Fp <- function(k, e) {
  Fv <- MS[k] / MS[e]
  c(F = Fv, p = pf(Fv, at[k, "Df"], at[e, "Df"], lower.tail = FALSE))
}
MS[Ec]       # 1 - after the Aitken transformation the pooled residual is exactly 1
Fp(2,  Ea)   # Row treatments (A)    F = 17.42     p = 0.0140
Fp(5,  Ea)   # Location x Row        F = 12.20     p = 0.0250
Fp(3,  Eb)   # Column treatments (B) F = 32.24     p = 0.0001
Fp(6,  Eb)   # Location x Column     F = 7.53      p = 0.0145
Fp(1,  Ec)   # Location              F = 12916.56  p < 0.0001
Fp(7,  Ec)   # Row x Column          F = 79.85     p < 0.0001
Fp(10, Ec)   # Location x Row x Col  F = 93.60     p < 0.0001

5.4 Post-hoc Comparison, CD and Letter Grouping

library(agricolae)
alpha <- 0.05

# Row treatments -> Pooled Error (a)
outA <- LSD.test(y, A, DFerror = at[Ea, "Df"], MSerror = MS[Ea], alpha = alpha)
# (values are on the Aitken-transformed scale)
outA$groups                          # A_nonchilled 42.69 a, A_chilled 28.55 b
outA$statistics[1, 6]                # CD = 9.40
r <- outA$means[1, 3]
sqrt(outA$statistics[1, 1] / r)      # SE(m) = 2.39
sqrt(2 * outA$statistics[1, 1] / r)  # SE(d) = 3.39
outA$statistics[1, 4]                # CV (%) = 28.52

# Column treatments -> Pooled Error (b)
outB <- LSD.test(y, B, DFerror = at[Eb, "Df"], MSerror = MS[Eb], alpha = alpha)
outB$groups                          # B_675 40.10 a, B_350 38.32 a, B_175 28.44 b
outB$statistics[1, 6]                # CD = 3.61
HSD.test(y, B, DFerror = at[Eb, "Df"], MSerror = MS[Eb], alpha = alpha)$statistics[1, 5]   # 4.47
duncan.test(y, B, DFerror = at[Eb, "Df"], MSerror = MS[Eb], alpha = alpha)$duncan          # critical ranges 3.61, 3.76

# Location -> Pooled Error (c)
LSD.test(y, L, DFerror = at[Ec, "Df"], MSerror = MS[Ec], alpha = alpha)$statistics[1, 6]   # CD = 0.77

5.5 Principal Component Analysis and Index Scores

library(gtools); library(factoextra); library(scales)
d$AxBxL <- interaction(d$FactorA, d$FactorB, d$Location, sep = "x")
means <- aggregate(cbind(Uptake, Efficiency) ~ AxBxL, data = d, FUN = mean)
rownames(means) <- means$AxBxL
means <- means[mixedorder(rownames(means)), -1]          # 12 combinations

res.pca <- prcomp(means, center = TRUE, scale = TRUE)
res.pca$x        <- -res.pca$x                           # sign convention used by the app
res.pca$rotation <- -res.pca$rotation
get_eigenvalue(res.pca)                                  # PC1 = 1.122 (56.08%), PC2 = 0.878 (43.92%)
rescale(res.pca$x[, 1], to = c(0, 1))                    # scaled Index 1

Explore the entire Pooled Strip Plot module in preview mode using our demo datasets. To submit suggestions or report a workflow issue, please use the discussion section below, or visit the official RAISINS website.

6 RAISINS Native Statistical Framework

RAISINS uses R for all its statistical computations. Every package used to generate major results is listed and demonstrated with examples, so results can be reproduced independently. These results are then organized and formatted on the RAISINS website along with visualisation to make them easier to use and interpret. RAISINS also has its own custom-built statistical tools for managing workflows, validating results, and generating reports. Details of these are not fully covered here, they’re shared with outside researchers only on request, and are subject to licensing terms.

7 Package References

R Core Team. (2025). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/

Fox, J., & Weisberg, S. (2024). car: Companion to Applied Regression (R package version 3.1-3). https://doi.org/10.32614/CRAN.package.car

de Mendiburu, F. (2023). agricolae: Statistical Procedures for Agricultural Research (R package version 1.3-7). https://doi.org/10.32614/CRAN.package.agricolae

Warnes, G. R., Bolker, B., & Lumley, T. (2023). gtools: Various R Programming Tools (R package version 3.9.5). https://doi.org/10.32614/CRAN.package.gtools

Wickham, H., François, R., Henry, L., Müller, K., & Vaughan, D. (2023). dplyr: A Grammar of Data Manipulation (R package version 1.1.4). https://doi.org/10.32614/CRAN.package.dplyr

Komsta, L., & Novomestky, F. (2022). moments: Moments, Cumulants, Skewness, Kurtosis and Related Tests (R package version 0.14.1). https://doi.org/10.32614/CRAN.package.moments

Kassambara, A., & Mundt, F. (2020). factoextra: Extract and Visualize the Results of Multivariate Data Analyses (R package version 1.0.7). https://doi.org/10.32614/CRAN.package.factoextra

Wickham, H., Pedersen, T. L., & Seidel, D. (2025). scales: Scale Functions for Visualization (R package version 1.4.0). https://doi.org/10.32614/CRAN.package.scales

De Rosario-Martinez, H. (2015). phia: Post-Hoc Interaction Analysis (R package version 0.3-2). https://doi.org/10.32614/CRAN.package.phia

Schloerke, B., Cook, D., Larmarange, J., Briatte, F., Marbach, M., Thoen, E., Elberg, A., & Crowley, J. (2025). GGally: Extension to ‘ggplot2’ (R package version 2.4.0). https://doi.org/10.32614/CRAN.package.GGally

Kay, M. (2025). ggdist: Visualizations of Distributions and Uncertainty (R package version 3.3.3). https://doi.org/10.32614/CRAN.package.ggdist

Gomez, K. A., & Gomez, A. A. (1984). Statistical Procedures for Agricultural Research (2nd ed.). John Wiley & Sons.

Feedback & Discussion