This Computational Provenance Record documents the statistical computing environment, software dependencies, computational provenance, and bibliographic references associated with the RAISINS Pooled Analysis in CRD module. It is intended to support computational reproducibility and software transparency. Detailed statistical methodology, mathematical derivations, and user guidance are provided separately in the official module documentation.
Any issues or updates required please comment here
Default Model 2 (Treatment fixed, Location/Season random). The choice sets the error term used to test each effect: Model 1 (both random), Model 2 (treatment fixed, environment random), Model 3 (environment fixed, treatment random), Model 4 (both fixed)
Homogeneity
& Aitken
(bartlett.test())
Test
Bartlett's test on the list of per-environment one-way aov() error variances, run once per response variable at α = 0.05
Consequence
p ≤ 0.05 ⇒ error variances heterogeneous, the Aitken transformation is applied to that trait (each environment's values divided by √MSE of that environment) before pooling; p > 0.05 ⇒ the raw data are pooled unchanged
Pooled ANOVA
(lm() + car::Anova())
Model & sums of squares
lm(y ~ Location + Treatment + Location:Treatment); Type II sums of squares via car::Anova(type = "II"); mean square = SS / df
F-tests
F-ratios formed against the error term implied by the selected model (e.g. under Model 2 the treatment effect is tested against the interaction mean square and the environment effect against the pooled error); p-values from stats::pf()
Multiple
Comparison
Test
Default LSD (LSD.test()); options TUKEY (HSD.test()) and DMRT (duncan.test()). Applied to Treatment, Location, and the Location×Treatment interaction
CD / critical range reported only when the corresponding effect is significant; suppressed (blank) otherwise
Normalising
Transformations
(manual, optional)
Formulae
Logarithmic log10(x) (or log10(x − min + 1) when values are non-positive); square-root sqrt(x) (or sqrt(x + 0.5) when zeros are present); arcsine asin(sqrt(x))
Scope
User-selected per variable; independent of the automatic Aitken transformation
Multivariate
Index
(prcomp())
PCA
prcomp(center = TRUE, scale = TRUE) on the treatment means (transformed means when Aitken was applied); requires at least three traits
Index & correlation
Component scores rescaled to [0, 1] with scales::rescale() to form the selection index; trait correlation via stats::cor()
Rounding
Displayed decimals
Results rounded to 2 decimal places by default (adjustable 1–4)
5 R Code for Key Analytical Steps
The code blocks below demonstrate the exact computation behind each reported result. The ANOVA steps use the built-in warpbreaks dataset, treating wool (2 levels) as the environment and tension (3 levels) as the treatment, with 9 replications per combination — the same structure as a pooled CRD. The multivariate step uses iris to illustrate the several-trait selection index.
5.1 Homogeneity of Variance (Bartlett) and the Aitken Decision
data(warpbreaks)env <- warpbreaks$wool # 2 environments (A, B)# one CRD per environment; Bartlett across their error variancesper_env <-lapply(split(warpbreaks, env),function(d) aov(breaks ~ tension, data = d))bt <-bartlett.test(per_env) # homogeneity across environmentsbt# per-environment MSE (used by Aitken when Bartlett is significant)mse <-sapply(per_env, function(m) sum(resid(m)^2) /df.residual(m))mse# if bt$p.value <= 0.05: divide each environment's values by sqrt(its MSE) before pooling
5.2 Pooled ANOVA
library(car)fit <-lm(breaks ~ wool + tension + wool:tension, # environment + treatment + interactiondata = warpbreaks)atab <-Anova(fit, type ="II") # Type II sums of squaresatabMS <- atab[, "Sum Sq"] / atab[, "Df"] # mean square = SS / dfMS# F-ratios use the error term implied by the chosen pooled model; p-values via:# pf(Fvalue, df1, df2, lower.tail = FALSE)
5.3 Multiple Comparison (Critical Difference)
library(agricolae)dfe <-df.residual(fit)mse_pool <-sum(resid(fit)^2) / dfealpha <-0.05# default (0.01 optional)# LSD is the default; TUKEY and DMRT are the alternativesLSD.test(warpbreaks$breaks, warpbreaks$tension,DFerror = dfe, MSerror = mse_pool, alpha = alpha, p.adj ="none")$groupsHSD.test(warpbreaks$breaks, warpbreaks$tension,DFerror = dfe, MSerror = mse_pool, alpha = alpha)$groupsduncan.test(warpbreaks$breaks, warpbreaks$tension,DFerror = dfe, MSerror = mse_pool, alpha = alpha)$groups
5.4 Normalising Transformations
x <- warpbreaks$breakslog_x <-if (min(x) <=0) log10(x -min(x) +1) elselog10(x) # logarithmicsqrt_x <-if (any(x ==0)) sqrt(x +0.5) elsesqrt(x) # square rootp <- (x -min(x)) / (max(x) -min(x))asin_x <-asin(sqrt(p)) # arcsine (proportions)# Aitken (automatic, applied per environment only when Bartlett is significant):# x_within_environment / sqrt(MSE_of_that_environment)
5.5 Multivariate Selection Index and Correlation
library(scales)# illustrative multi-trait treatment means (rows = treatments, cols = traits)means <-aggregate(cbind(Sepal.Length, Sepal.Width, Petal.Length) ~ Species,data = iris, FUN = mean)X <-scale(means[, -1])pca <-prcomp(X, center =TRUE, scale. =TRUE) # requires >= 2 traitsindex <-rescale(pca$x[, 1], to =c(0, 1)) # PCA-based selection index in [0, 1]data.frame(Treatment = means$Species, Index =round(index, 3))cor(iris[, 1:4]) # trait correlation matrix
Explore the entire Pooled Analysis in CRD module in preview mode using our demo datasets. To submit suggestions or report a workflow issue, please use the discussion section below, or visit the official RAISINS website.
6 RAISINS Native Statistical Framework
RAISINS uses R for all its statistical computations. Every package used to generate major results is listed and demonstrated with examples, so results can be reproduced independently. These results are then organized and formatted on the RAISINS website along with visualisation to make them easier to use and interpret. RAISINS also has its own custom-built statistical tools for managing workflows, validating results, and generating reports. Details of these are not fully covered here, they’re shared with outside researchers only on request, and are subject to licensing terms.
7 Package References
R Core Team. (2025). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/
de Mendiburu, F. (2023). agricolae: Statistical Procedures for Agricultural Research (R package version 1.3-7). https://doi.org/10.32614/CRAN.package.agricolae
Fox, J., & Weisberg, S. (2019). car: Companion to Applied Regression (R package version 3.1-5). https://doi.org/10.32614/CRAN.package.car
Gross, J., & Ligges, U. (2015). nortest: Tests for Normality (R package version 1.0-4). https://doi.org/10.32614/CRAN.package.nortest
Wickham, H., Pedersen, T. L., & Seidel, D. (2025). scales: Scale Functions for Visualization (R package version 1.4.0). https://doi.org/10.32614/CRAN.package.scales