Computational Provenance & Reproducibility Record

RAISINS - R and AI Solutions for INferential Statistics · Online Statistical Analysis Platform for Agricultural Research

Computational Provenance & Reproducibility Record Pooled Analysis in RBD · 2.0.0 · DOI 10.5281/zenodo.23076715

Computational Provenance & Reproducibility Record

RAISINS · Pooled Analysis in RBD Module

This Computational Provenance Record documents the statistical computing environment, software dependencies, computational provenance, and bibliographic references associated with the RAISINS Pooled Analysis in RBD module. It is intended to support computational reproducibility and software transparency. Detailed statistical methodology, mathematical derivations, and user guidance are provided separately in the official module documentation.

Any issues or updates required please comment here

Go to the App from here

1 Module Metadata

Parameter Specification
Module Pooled Analysis in RBD
Module Version 2.0.0
DOI 10.5281/zenodo.23076715
Document Type Computational workflow
Statistical Engine R
R Version 4.5.2
Reproducibility Execution Environment: Posit Connect · GCR · renv-locked

2 Statistical Dependency Manifest

Package Version Repository Core Statistical Functions
agricolae 1.3-7 CRAN LSD.test(), HSD.test(), duncan.test()
car 3.1-5 CRAN Anova()
nortest 1.0-4 CRAN ad.test()
scales 1.4.0 CRAN rescale()
stats 4.5.2 Base R lm(), aov(), bartlett.test(), pf(), prcomp(), cor(), shapiro.test(), sd(), mean()

3 Statistical Function Registry

Analytical Role Primary Function(s)
Per-environment error variance (input to pooling) stats::aov() (RBD: y ~ Block + Treatment per environment)
Homogeneity of variances across environments stats::bartlett.test()
Pooled ANOVA (Location, block-within-location, Treatment, interaction) & F-tests stats::lm(), car::Anova() (Type II), stats::pf()
Multiple comparison / critical difference agricolae::LSD.test(), agricolae::HSD.test(), agricolae::duncan.test()
Multivariate selection index & trait correlation stats::prcomp(), scales::rescale(), stats::cor()
Normality diagnostics stats::shapiro.test(), nortest::ad.test()
Descriptive statistics base::mean(), stats::sd()

4 Default Methods & Parameters

Analysis Step / Parameter Default Method / Value
Pooled Model Model selection Default Model 2 (Treatment fixed, Location/Season random). The choice sets the error term used to test each effect: Model 1 (both random), Model 2 (treatment fixed, environment random), Model 3 (environment fixed, treatment random), Model 4 (both fixed)
Homogeneity
& Aitken

(bartlett.test())
Test Bartlett's test on the list of per-environment RBD aov(y ~ Block + Treatment) error variances, run once per response variable at α = 0.05
Consequence p ≤ 0.05 ⇒ error variances heterogeneous, the Aitken transformation is applied to that trait (each environment's values divided by √MSE of that environment) before pooling; p > 0.05 ⇒ the raw data are pooled unchanged
Pooled ANOVA
(lm() + car::Anova())
Model & sums of squares lm(y ~ Location + Location:Block + Treatment + Location:Treatment); Type II sums of squares via car::Anova(type = "II"); mean square = SS / df. The Location:Block term is the block (replication) effect within environments — unique to the RBD
F-tests F-ratios formed against the error term implied by the selected model (e.g. under Model 2 the treatment effect is tested against the interaction mean square and the environment effect against the pooled error); p-values from stats::pf()
Multiple
Comparison
Test Default LSD (LSD.test()); options TUKEY (HSD.test()) and DMRT (duncan.test()). Applied to Treatment, Location, and the Location×Treatment interaction
Significance & adjustment α default 0.05 (0.01 optional); p-adjustment default none (standard p.adjust() methods selectable, LSD only)
Critical difference CD / critical range reported only when the corresponding effect is significant; suppressed (blank) otherwise
Normalising
Transformations

(manual, optional)
Formulae Logarithmic log10(x) (or log10(x − min + 1) when values are non-positive); square-root sqrt(x) (or sqrt(x + 0.5) when zeros are present); arcsine asin(sqrt(x))
Scope User-selected per variable; independent of the automatic Aitken transformation
Multivariate
Index

(prcomp())
PCA prcomp(center = TRUE, scale = TRUE) on the treatment means (transformed means when Aitken was applied); requires at least two traits
Index & correlation Component scores rescaled to [0, 1] with scales::rescale() to form the selection index; trait correlation via stats::cor()
Rounding Displayed decimals Results rounded to 2 decimal places by default (adjustable 1–4)

5 R Code for Key Analytical Steps

The code blocks below demonstrate the exact computation behind each reported result. The ANOVA steps use the built-in warpbreaks dataset, treating wool (2 levels) as the environment and tension (3 levels) as the treatment, with a block (replication) index built so that every block within an environment contains each treatment once — the same structure as a pooled RBD. The multivariate step uses iris to illustrate the several-trait selection index.

5.1 Building the Pooled-RBD Layout

data(warpbreaks)
wb <- warpbreaks                                    # wool = environment, tension = treatment
# block (replication) index: within each environment, blocks 1..9 each hold all 3 treatments
wb$block <- factor(ave(seq_len(nrow(wb)), wb$wool,
                       FUN = function(i) rep(1:9, times = 3)))
head(wb)

5.2 Homogeneity of Variance (Bartlett) and the Aitken Decision

# one RBD per environment; Bartlett across their error variances
per_env <- lapply(split(wb, wb$wool),
                  function(d) aov(breaks ~ block + tension, data = d))
bt <- bartlett.test(per_env)                        # homogeneity across environments
bt

# per-environment MSE (used by Aitken when Bartlett is significant)
mse <- sapply(per_env, function(m) sum(resid(m)^2) / df.residual(m))
mse
# if bt$p.value <= 0.05: divide each environment's values by sqrt(its MSE) before pooling

5.3 Pooled ANOVA

library(car)
fit  <- lm(breaks ~ wool + wool:block + tension + wool:tension,  # env + block(env) + trt + interaction
           data = wb)
atab <- Anova(fit, type = "II")                     # Type II sums of squares
atab

MS <- atab[, "Sum Sq"] / atab[, "Df"]               # mean square = SS / df
MS
# F-ratios use the error term implied by the chosen pooled model; p-values via:
# pf(Fvalue, df1, df2, lower.tail = FALSE)

5.4 Multiple Comparison (Critical Difference)

library(agricolae)
dfe      <- df.residual(fit)
mse_pool <- sum(resid(fit)^2) / dfe
alpha    <- 0.05                                    # default (0.01 optional)

# LSD is the default; TUKEY and DMRT are the alternatives
LSD.test(wb$breaks, wb$tension,
         DFerror = dfe, MSerror = mse_pool, alpha = alpha, p.adj = "none")$groups
HSD.test(wb$breaks, wb$tension,
         DFerror = dfe, MSerror = mse_pool, alpha = alpha)$groups
duncan.test(wb$breaks, wb$tension,
            DFerror = dfe, MSerror = mse_pool, alpha = alpha)$groups

5.5 Normalising Transformations

x <- wb$breaks

log_x  <- if (min(x) <= 0) log10(x - min(x) + 1) else log10(x)   # logarithmic
sqrt_x <- if (any(x == 0)) sqrt(x + 0.5)         else sqrt(x)    # square root
p      <- (x - min(x)) / (max(x) - min(x))
asin_x <- asin(sqrt(p))                                          # arcsine (proportions)

# Aitken (automatic, applied per environment only when Bartlett is significant):
#   x_within_environment / sqrt(MSE_of_that_environment)

5.6 Multivariate Selection Index and Correlation

library(scales)

# illustrative multi-trait treatment means (rows = treatments, cols = traits)
means <- aggregate(cbind(Sepal.Length, Sepal.Width, Petal.Length) ~ Species,
                   data = iris, FUN = mean)
X   <- scale(means[, -1])
pca <- prcomp(X, center = TRUE, scale. = TRUE)      # requires >= 2 traits

index <- rescale(pca$x[, 1], to = c(0, 1))          # PCA-based selection index in [0, 1]
data.frame(Treatment = means$Species, Index = round(index, 3))

cor(iris[, 1:4])                                     # trait correlation matrix

Explore the entire Pooled Analysis in RBD module in preview mode using our demo datasets. To submit suggestions or report a workflow issue, please use the discussion section below, or visit the official RAISINS website.

6 RAISINS Native Statistical Framework

RAISINS uses R for all its statistical computations. Every package used to generate major results is listed and demonstrated with examples, so results can be reproduced independently. These results are then organized and formatted on the RAISINS website along with visualisation to make them easier to use and interpret. RAISINS also has its own custom-built statistical tools for managing workflows, validating results, and generating reports. Details of these are not fully covered here, they’re shared with outside researchers only on request, and are subject to licensing terms.

7 Package References

R Core Team. (2025). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/

de Mendiburu, F. (2023). agricolae: Statistical Procedures for Agricultural Research (R package version 1.3-7). https://doi.org/10.32614/CRAN.package.agricolae

Fox, J., & Weisberg, S. (2019). car: Companion to Applied Regression (R package version 3.1-5). https://doi.org/10.32614/CRAN.package.car

Gross, J., & Ligges, U. (2015). nortest: Tests for Normality (R package version 1.0-4). https://doi.org/10.32614/CRAN.package.nortest

Wickham, H., Pedersen, T. L., & Seidel, D. (2025). scales: Scale Functions for Visualization (R package version 1.4.0). https://doi.org/10.32614/CRAN.package.scales

Feedback & Discussion