RAISINS
  • Home
  • Get Started!
    • Data Analysis
    • Analysis of Experiments
    • Non Parametric tests
    • Statistical Genetics
    • Social Sciences
    • Sample size Calculator
    • Econometrics
    • Custom Tools
  • Learn
    • Tutorials
    • Quick Videos
    • Webinars
    • Wine
  • Team
  • Resources
    • Citation Info
    • Discussion
  • Pricing Plans
  • RAISINS Agent
  • Feedback
  • Contact us

On this page

  • 1 What is a pooled two-factor RBD?
  • 2 The four pooled models
  • 3 How RAISINS decides whether your environments can be pooled
  • 4 Assumptions of the pooled analysis
  • 5 Getting to the module
    • 5.1 Computational Provenance & Reproducibility Record
  • 6 Preview mode and Quick Tour
  • 7 A working example
  • 8 How to prepare your data
    • 8.1 Preparing data in MS Excel
    • 8.2 Prepare using Create Data in RAISINS
    • 8.3 Download Model Datasets
    • 8.4 Creating a dataset using RA-One chat
  • 9 The Analysis Results tab
    • 9.1 The five presentation options
    • 9.2 What the module tells you before the ANOVA
  • 10 Transformation
  • 11 Analysis results
    • 11.1 The pooled ANOVA table
    • 11.2 Interpretation from Figure 8
  • 12 Basic plots
  • 13 Advanced plots and diagnostics
  • 14 Looking at all traits together: the PCA index
  • 15 Interpretation
  • 16 Chat with your data using RA-One
  • 17 FAQs
  • 18 View data
  • 19 Wrapping up

Pooled Two Factor RBD

Data Analysis

Two treatment factors, several environments, eight effects and one pooled error term. This tutorial explains how they fit together and works a complete pooled two-factor RBD analysis in RAISINS… Read more …

Authors
Affiliations

Arshida A K

Statoberry LLP

Dr. Pratheesh P Gopinath

Kerala Agricultural University

Published

September 11, 2026

Abstract

A pooled two-factor RBD takes a factorial experiment laid out in randomized complete blocks and repeats it over several locations, seasons or years, then analyses every environment together in one combined analysis of variance. The point is not simply to gain replication. Alongside the two main effects and their interaction, the pooled analysis asks a question no single-site trial can answer: does the recommendation hold everywhere, or does it change from place to place? This tutorial explains what the analysis involves - the four models you can choose between, the homogeneity check that licenses pooling, and the order in which the eight effects must be read - and then carries one dataset through the module from upload to conclusion: the ANOVA, mean separation, transformation, plots, principal component analysis and AI-assisted interpretation, without writing a line of code.

1 What is a pooled two-factor RBD?

Suppose you are testing two spacings, S1 and S2, in combination with two crop-management schedules, C1 and C2. That is a two-factor factorial: four treatment combinations, laid out in a randomized complete block design with four blocks to soak up the fertility gradient running across the field. One season at one farm gives you a clean answer for that farm. The trouble starts when you are asked to write the recommendation.

A recommendation is a claim about places you did not test. If the trial ran only at Location A, you have no evidence at all about whether S2 will still beat S1 at Location B, where the soil is heavier and the rain arrives three weeks later. The honest answer from a single site is that you do not know. So you repeat the whole factorial at a second location, and now you have two experiments - and the question of what to do with them.

The wrong thing to do is analyse each location separately and eyeball the two tables for agreement. Two separate analyses cannot tell you whether a difference between the locations is real or is just the noise you would expect from two independent experiments. A pooled analysis puts every observation into one analysis of variance, and in doing so it earns you a new set of terms that no single-site analysis contains: the effect of Location itself, and the interactions of each treatment factor with location.

Those interaction terms are the entire reason the design is worth the extra season. A significant FactorA × Location says that the spacing effect is not the same at the two sites - that the best spacing at A is not the best spacing at B, and a single blanket recommendation would be wrong. A non-significant one says the opposite, and much more usefully: the effect you measured is stable across the environments you sampled, and you may recommend it for both.

With two factors and one environment factor, the pooled analysis estimates eight effects in all: the three you would have had from a single site (FactorA, FactorB, FactorA × FactorB), the environment main effect (Location), the three environment interactions (A × Location, B × Location, Location × A × B), and the blocks, which are nested inside locations and appear as Location × Block.

In one sentence

A pooled two-factor RBD tests whether your two factors matter and whether what they do stays the same when you move to a different location, season or year.

2 The four pooled models

Before the module can compute a single F ratio it needs to know something you cannot read off the data: whether your treatments and your environments are fixed or random. The distinction is not statistical pedantry - it decides which mean square goes into the denominator of which F test, and therefore which effects come out significant.

A factor is fixed when its levels are the only ones you care about and you chose them deliberately. Two specific spacings that you intend to recommend are fixed. A factor is random when its levels are a sample from a larger population that you want to generalise to. Three locations drawn from a wider growing region, or three consecutive seasons standing in for “seasons in general”, are random.

Model Treatment Location / Season Use it when
Model 1 random random Both the treatments and the environments are samples from larger populations
Model 2 fixed random You chose the treatments deliberately; the sites stand in for a wider region
Model 3 random fixed The treatments are a sample; these specific sites are the ones you care about
Model 4 fixed fixed You chose both the treatments and the sites deliberately

Model 2 and Model 4 cover most agricultural work. Choose Model 2 when the locations are a sample meant to represent a region and you want a recommendation that generalises beyond them. Choose Model 4 when the trial was run at exactly the sites you intend to advise about - a research station and a farmer’s field you will both report on by name - and you are not extrapolating. The worked example in this tutorial uses Model 4.

The choice changes your p-values, not just your wording

Treating Location as random moves it into the denominator of the treatment F tests, which usually makes the treatment effects harder to declare significant - deliberately so, because you are asking the stronger question of whether the effect survives across the whole population of environments. If you pick the model to get the answer you want rather than the one your sampling justifies, the p-values are not meaningful. Decide before you look.

The module carries a Know more link beside the model selector that restates these four cases while you are choosing.

3 How RAISINS decides whether your environments can be pooled

Pooling rests on an assumption that is easy to state and easy to violate: that the experimental error is about the same size at every location. If one site was uniform and the other was a patchwork, their errors are not comparable, and averaging them produces a pooled error term that is too large for the good site and too small for the bad one. Every F test in the table would then be wrong, in opposite directions at the two sites.

So the module checks first. Before the ANOVA is computed it runs Bartlett’s test on the error variances across locations, separately for every response variable you selected. Bartlett’s test asks whether the error variances differ by more than chance would explain. Its null hypothesis is that they are homogeneous, so here - unusually - a non-significant result is the one you want.

Bartlett’s result What it means What the module does
p > 0.05 (non-significant) Error variances are homogeneous across environments Pools the error directly; analysis proceeds unchanged
p < 0.05 (significant) Error variances differ between environments Applies an Aitken transformation, weighting each environment by its own error, before pooling

The Aitken correction is not a fallback that quietly degrades your analysis - it is the standard remedy. Each environment’s data is rescaled by its own error variance so that, after transformation, the environments are on a comparable footing and the pooled error is legitimate. The module tells you plainly which path it took: it prints either “No Aitken transformation applied” or a note that the correction was made, together with the per-location MSE values it used.

Why you see MSE values listed per location

Underneath the Bartlett table the module lists the error mean square for each response at each location - for example, for Yield, A = 0.11 and B = 0.03. These are the numbers Bartlett’s test is comparing. Reading them gives you a feel for how different the two sites really were, independently of whether the test crossed a threshold. Two MSEs of 0.11 and 0.03 differ by roughly a factor of four, yet Bartlett’s test on this dataset returns p = 0.09 - not quite enough evidence, with four blocks, to reject homogeneity.

The check is performed per character, not once for the whole file. It is entirely normal for a dataset to pass on most traits and fail on one, and the module handles each response on its own terms.

4 Assumptions of the pooled analysis

The pooled two-factor RBD inherits the assumptions of an ordinary factorial RBD and adds one of its own.

Assumption What it means What to do if it fails
Homogeneity of error variance across environments Each location contributes error of comparable size Handled automatically - see Section 3. The Aitken correction is applied for you
Normality of residuals The residuals follow an approximately normal distribution Inspect the Q–Q plot in Advanced Plots (Section 13); consider a transformation (Section 10)
Homogeneity of variance across treatments Spread is similar for every treatment combination A log or square-root transformation usually helps (Section 10)
Additivity Block and treatment effects add rather than multiply A transformation often restores additivity
Independence Observations do not influence one another Design and field execution decide this; no post-hoc fix
Complete, balanced blocks Every treatment combination appears once in every block, at every location The module expects a balanced file; check it in View Data (Section 18)
Independence is the one you cannot repair

Every other assumption on this list has a remedy available after the data are collected. Independence does not. If plots were not randomized, if the same plant was measured twice and entered as two rows, or if whole blocks share an irrigation channel that treatments do not cross, no transformation will fix it. That is decided when you lay out the trial, not when you analyse it.

With the theory settled, let us open the module and run an analysis - continue to Section 5.

5 Getting to the module

Visit the RAISINS home page at www.raisins.live and go to Data Analysis. In this tutorial we use the Pooled Two Factor RBD module. Opening its tile brings you to the Welcome page in Figure 1.

Figure 1: Welcome page of the Pooled Two Factor RBD module

Two routes lead inside. Get Started is for users holding an individual licence; Institutional Login is for users whose institution has purchased access. If you have neither yet, Explore Subscription Plans and Preview mode are both on this page, and the latter needs no account at all (Section 6). The same page carries a Quick video, the View License and Version Info links, and a Contact Us button that reaches the team directly.

5.1 Computational Provenance & Reproducibility Record

CPRR (Computational Provenance & Reproducibility Record) provides a transparent and comprehensive record. Click on the CPRR icon carried on the module’s tile in the Data Analysis section to access it and know about the computational workflow performed during the analysis. The record for this module states the R version and the exact version of every package used, and names the specific function behind each reported result. CPRR lists every default parameter and decision rule applied by the module and provides fully runnable R code that reproduces each analytical step. Users can execute the code in R to independently reproduce and verify the results. It carries its own DOI.

To cite the platform itself in a paper, thesis, or report, use the RAISINS citation, available in APA, Harvard, and BibTeX formats at www.raisins.live/citation.html. That is the primary reference, and for most manuscripts it is all you need.

The CPRR for the Pooled Two Factor RBD is at www.raisins.live/module_record/pooled_2frbd.html.

How to use the two together

Cite the RAISINS paper as your primary reference for the platform. Add the CPRR as supporting documentation when a journal asks for details of the computing environment, or when you want your methods section to be precise about versions and functions rather than saying “analysis was carried out using an online tool.” For this module the CPRR is particularly worth citing, because it states exactly how the homogeneity test and the Aitken correction are implemented - the step most likely to differ between software packages, and the one a reviewer is most likely to query.

6 Preview mode and Quick Tour

Before subscribing, you can explore the entire module using Preview mode, reached from the Welcome page in Figure 1. In this mode the upload control is withdrawn - your own data cannot be submitted - and a demo dataset selector appears in its place, offering the same model datasets described in Section 8.3. Every other facility remains available: the analysis, the plots, the multivariate procedures and the RA-One assistant can all be exercised on the demonstration data before you commit to anything.

First-time users are also offered a Quick Tour, an interactive, step-by-step walkthrough that highlights each control and explains what it does, running from the upload panel through to the View Data diagnostic. You can retake the tour at any time from the Quick Tour tab in the top navigation bar.

7 A working example

Everything from here on uses one dataset, so that each screen you see is a step in a single continuous analysis rather than a disconnected illustration.

Element In this dataset
Environments 2 locations, A and B
Factor A 2 levels, S1 and S2
Factor B 2 levels, C1 and C2
Blocks 4 per location
Treatment combinations 4 per location; 8 combinations of treatment × location in all
Responses 7 - Yield plus Char1 to Char6
Rows 32 - that is 2 locations × 2 × 2 treatment combinations × 4 blocks

This is the smallest arrangement that still exercises every part of the design: two factors so there is a factorial interaction, two environments so there are environment interactions, and four blocks so the pooled error has enough degrees of freedom to be worth testing against. Seven responses are carried through together, which is the normal case - you rarely measure just one thing - and lets us show how the module handles a file where different traits behave differently.

8 How to prepare your data

The module accepts a CSV or Excel file in one specific shape, and there are four ways to arrive at it. Build it yourself in a spreadsheet (Section 8.1), let the Create Data tab generate a blank template (Section 8.2), download one of the ready-made model datasets (Section 8.3), or describe what you need to the RA-One assistant in plain English (Section 8.4). All four produce the same layout; pick whichever suits you.

8.1 Preparing data in MS Excel

The rule is one row per observation and one column per variable. Four identifier columns come first - the environment, the two treatment factors and the block - followed by one column for each response you measured. Figure 2 shows the working example laid out exactly this way.

Figure 2: The pooled two-factor RBD data file in MS Excel: four identifier columns followed by the responses
Column Holds In Figure 2
Location The environment: location, season or year A, B
FactorA Levels of the first treatment factor S1, S2
FactorB Levels of the second treatment factor C1, C2
Block The replication or block within each environment 1, 2, 3, 4
Yield, Char1 … Char6 One column per measured response Numeric values
The mistakes that actually cause failures

Blank rows or blank columns inside the data block; a trailing space after a level name, which makes S1 a different level from S1; text such as NA, - or missing typed into a numeric column; merged cells; and a block numbered differently at the two locations. Every one of these is silent in Excel and fatal in analysis. The View Data tab (Section 18) is built to catch them before you run anything.

Naming rules for columns and levels
  • Column names may contain letters, numbers, dots and underscores. Avoid spaces, and avoid starting a name with a digit.
  • Keep level names short and identical everywhere they appear - S1 throughout, never S1 in one place and s1 in another.
  • The block column must be complete: every treatment combination appears once in every block, at every location.
  • Response columns must be purely numeric. If a value is genuinely missing, leave the cell empty rather than typing a placeholder.
  • The order of the rows does not matter. The order of the columns does not matter either, because you map each one by name in the Analysis tab.

8.2 Prepare using Create Data in RAISINS

If you have not yet collected the data, the Create Data tab builds the empty file for you, correctly structured, so that you can carry it to the field and fill it in. Figure 3 shows it.

Figure 3: The CSV data file creator in the Create Data tab

You supply five numbers - how many locations, seasons or years; how many levels of Factor A; how many levels of Factor B; how many blocks; and how many characters you intend to measure - then press Create. The panel on the right fills with a complete template carrying one row for every combination, with an empty response column y1 waiting for your readings. You can type into it directly, or paste a block of values straight from Excel with Ctrl+V. Download CSV file saves it, and the saved file uploads into the Analysis tab unchanged.

8.3 Download Model Datasets

The Datasets tab carries worked example files you can download and run immediately - useful for learning the module, and for checking that a problem is in your data rather than in the app. Figure 4 shows the page.

Figure 4: Model datasets supplied with the module

Dataset 1 is the file used throughout this tutorial: FactorA at two levels (S1, S2) and FactorB at two levels (C1, C2) in four blocks, across locations A and B, with Yield and Char1 to Char6 recorded. Dataset 2 is a larger arrangement - FactorA at three levels (a1, a2, a3), FactorB at two (b1, b2), three blocks, two locations (L1, L2) and six responses named y1 to y6 - which is worth loading once you want to see how the tables grow when a factor has more than two levels.

8.4 Creating a dataset using RA-One chat

The fourth route is to ask. Open the RA-One tab and describe the design you want in ordinary English - “Create a Pooled 2FRBD data template for 2 locations, 3 × 2 factorial, 3 reps, 3 responses”. RA-One builds the grid in the conversation, as in Figure 5.

Figure 5: RA-One building a pooled 2FRBD data template from a plain-English request

What comes back is not a picture of a table but a working one. The header restates the design it understood - locations, Factor A, Factor B, replications and response columns - and each of those is an editable box, so if it read “3 responses” when you meant two, change the number and press Rebuild rather than starting the conversation again. Useful detail: it reports the resulting pooled error degrees of freedom as you adjust the numbers, which is a quick way to check that a planned design will have enough error df to be worth running before you commit a season to it. Add column appends another response, and Download CSV saves the file for upload.

Your file is ready. Continue to Section 9 to run the analysis.

9 The Analysis Results tab

This is the tab where the analysis is specified and run. The panel on the left takes your file and tells the module what each column is; the strip along the top governs how the results are presented. Figure 6 shows it with the working example loaded.

Figure 6: The Analysis Results tab: the upload and mapping panel on the left, presentation options along the top, and the pooling diagnostics below

Work down the left panel in order. Each selector is a dropdown listing the column names read from your file, so nothing has to be typed.

Control What to give it
Upload data file Browse to your CSV or Excel file; a blue Upload complete bar confirms it
Select pooled Model One of the four models from Section 2. The working example uses MODEL 4
Select Location/Season/Year The column identifying the environment - here Location
Select Factor A The first treatment factor - here FactorA
Select Factor B The second treatment factor - here FactorB
Select the blocks The replication or block column - here Block
Select variables Every response you want analysed. All seven are selected at once
Click for Transformation Optional; opens the panel described in Section 10
Run Analysis! Computes everything
Select every response at once

The variables selector is multi-select, and there is no penalty for choosing all of them. The module analyses each response separately and lays the results out character by character, so a single run gives you the whole experiment. You do not need to run the analysis seven times.

9.1 The five presentation options

The green strip along the top of Figure 6 carries five controls. They change how results are reported, not what is computed, so you can adjust them after a run and the tables update without re-uploading anything.

Option Choices What it does
Multiple comparison test LSD, TUKEY, DMRT The post-hoc test used to separate means once an effect is significant
P-adjustment None, Bonferroni (FWER), Holm-Bonferroni (FWER), Benjamini-Hochberg (FDR) Controls the error rate across many pairwise comparisons
Level of significance (α) 0.05 and other conventional levels The threshold against which p-values are judged
Digits after decimal A number How many decimal places the tables display
Select Font A list of fonts The typeface of the rendered tables
P-adjustment appears only with LSD

The P-adjustment selector is shown when the multiple comparison test is LSD, and is hidden for TUKEY and DMRT. That is deliberate rather than an oversight: Tukey’s HSD and Duncan’s multiple range test already control the error rate across the family of comparisons by their own construction, so layering a Bonferroni or Benjamini-Hochberg correction on top of them would penalise the same comparisons twice. LSD does not control family-wise error on its own, which is exactly why the option is offered there.

Which post-hoc test should I choose?
  • LSD (Fisher’s protected least significant difference) is the most liberal of the three. It is defensible when the ANOVA F test for that effect is already significant, and it is the module’s default.
  • TUKEY (HSD) controls the family-wise error rate across all pairwise comparisons. It is the conservative, widely accepted choice when you intend to compare every mean with every other.
  • DMRT (Duncan’s multiple range test) sits between the two, using a critical range that widens as the means being compared move further apart in rank.
  • Choose before you look at the output, not after. Running all three and reporting whichever produced the most letters is the multiple-comparison problem in a new costume.

9.2 What the module tells you before the ANOVA

Below the options strip, and above the results proper, the module prints a plain-language account of what it is about to do and what it found when it checked your data. For the working example it reports the model in force - “You have selected Model 4, which assumes both Treatment and Location/Season as fixed effects” - and then restates the design it read from your file: a pooled two-factor factorial in RCBD with 4 blocks, 2 treatments (S1 and S2), 2 levels of Location (A and B), and 8 treatment combinations evaluated in all, with LSD selected at α = 0.05.

Next comes the pooling diagnostic described in Section 3. The Bartlett χ² test results over pooling table gives a chi-square statistic and a p-value for every response. In Figure 6 every p-value is comfortably above 0.05 - Yield returns χ² = 2.93 with p = 0.09, and the six Char variables range from p = 0.25 to p = 0.68 - so the error variances are homogeneous and no correction is needed. Underneath, the per-location MSE values are listed for each response, and the module closes the section with two statements you should always read: “No Aitken transformation applied” and “Variance is homogeneous for all characters across all Location.”

Read those two lines every time

They are the audit trail for the most consequential decision the module makes on your behalf. If they instead report that an Aitken transformation was applied, nothing has gone wrong - the correction is the standard remedy - but your methods section must say so, because the analysis that follows is then a weighted one rather than a simple pooling.

10 Transformation

Ticking Click for Transformation in the left panel opens the panel shown in Figure 7. It offers three transformations, and each takes its own list of variables, so you can apply a log to one response, a square root to another, and leave the rest untouched in a single run.

Figure 7: The transformation panel: log, square-root and arcsin, each with its own variable selector
Transformation Use it for Typical case
Log Data whose variance grows with the mean; strongly right-skewed responses Counts of insects, spores or weeds; yields spanning an order of magnitude
Square-root Count data, particularly with small numbers and several zeros Number of tillers, pods, branches
Arcsin Proportions and percentages bounded between 0 and 1, or 0 and 100 Germination percentage, disease incidence
A transformation changes what the means mean

After a log transformation the analysis is carried out on the logged values, so the treatment means the ANOVA compares are means of logs - not the log of the mean, and not in the original units. Back-transforming a mean of logs gives a geometric mean, which is legitimate but is not the arithmetic average of your raw readings. Report the scale you analysed on, and say in your methods that a transformation was applied. RAISINS shows the transformed mean in parentheses alongside the original so that both are available to you.

Transformation is a remedy for the assumptions in Section 4, not a routine step. Run the analysis untransformed first, look at the Q–Q plot and the distribution plots in Section 13, and reach for a transformation only if they show a problem. If the residuals look reasonable, leave the data alone.

After choosing a transformation, or deciding against one, proceed to Section 11.

11 Analysis results

Pressing Run Analysis! produces two kinds of output. The Analysis Results tab carries the character-wise summary tables - one block of tables per response, giving the mean of every factor level and every interaction combination, with a letter grouping attached as a superscript wherever the effect was significant. The Indv. ANOVA tab carries the ANOVA table itself, one character at a time.

Read the ANOVA first. It tells you which effects are real; the mean tables then tell you what those effects are. Reading them the other way round invites you to interpret differences that the F test never licensed.

11.1 The pooled ANOVA table

Open the Indv. ANOVA tab and choose a response from the Select Character dropdown. Figure 8 shows the result for Yield.

Figure 8: The pooled ANOVA table for Yield, from the Indv. ANOVA tab

Every row is one of the eight effects introduced in Section 1, plus the block term and the pooled error. The columns are the degrees of freedom, the mean square, the F statistic and the p-value, and the footnotes below the table record the conventions: a single asterisk marks significance at the 5% level, a double asterisk at 1%, and NS marks a non-significant effect.

A quick arithmetic check you can do yourself

The degrees of freedom should account for every observation. Here: Location 1, FactorA 1, FactorB 1, Location × Block 6, A × B 1, A × Location 1, B × Location 1, Location × A × B 1, and pooled error 18. They sum to 31, which is one less than the 32 rows in the file - exactly as it should be. The block term carries 6 df because there are 2 locations each contributing 4 − 1 = 3. If your own table does not add up this way, the file is unbalanced and the View Data tab (Section 18) is the place to find out why.

11.2 Interpretation from Figure 8

For Yield, nothing in this experiment is significant. The two main effects are small - FactorA has a mean square of 0.06 against a pooled error of 0.08, giving F = 0.86, and FactorB gives F = 0.70 with p = 0.41. Neither factor shifted yield. The factorial interaction A × B returns F = 0.38 (p = 0.54), so the two factors do not modify one another either.

The environment terms tell the same story, and this is the part that matters for a recommendation. Location itself gives F = 0.10 (p = 0.76): the two sites yielded alike. A × Location gives F = 1.20 (p = 0.29) and B × Location F = 0.10 (p = 0.76), so neither factor behaved differently at the two sites. The three-way Location × A × B returns F = 0.09 (p = 0.77). The Location × Block term is the largest in the table, F = 2.49 with p = 0.06, which simply says that blocking was doing some work - there was a fertility gradient worth removing - and is not a treatment result.

A null result across environments is still a finding

It is tempting to read a table like this as a failed experiment. It is not. Non-significant environment interactions are precisely the evidence that licenses a general recommendation: whatever these treatments do, they do consistently at both sites. The correct conclusion for Yield here is that no treatment difference was detected, and that this absence was consistent across locations - which is a stronger and more useful statement than a single-site null would have been. What you may not say is that the treatments are equivalent; absence of evidence is not evidence of absence, and with four blocks the experiment simply may not have had the power to detect a small difference.

Yield is not the whole file. Working through the other six characters with the same dropdown turns up a different picture for Char2, where both the Location × FactorA interaction and the three-way Location × FactorA × FactorB interaction are significant. Those results, and the mean separation that follows from them, are read out in full in Section 15.

What the other numbers in the summary tables mean
  • CD (Critical Difference), sometimes reported as LSD: two means differ significantly if they differ by more than this. It is printed only when the corresponding F test was significant.
  • SE(m) - the standard error of a mean. A measure of the precision of a single treatment mean.
  • SE(d) - the standard error of the difference between two means. Roughly √2 times SE(m) in a balanced design.
  • CV(%) - the coefficient of variation, the pooled error expressed as a percentage of the grand mean. A rough guide to how well the trial was conducted; what counts as acceptable depends heavily on the crop and the trait.
  • Letter groupings appear as superscripts on the means. Two means that share a letter are not significantly different. They are shown only where the F test was significant, so a table with no letters is not an error.

12 Basic plots

The Basic Plots tab offers five standard displays, each generated by clicking its icon. Figure 9 shows the tab with a box plot drawn for Factor A.

Figure 9: The Basic Plots tab: five plot types, the settings panel, and the download controls
Plot Shows Good for
Boxplot Median, quartiles and outliers per level Spotting unequal spread and stray values
Violin Plot The full distribution’s shape per level Seeing bimodality a boxplot would hide
Mean Value Plot Treatment means with error bars The figure most often wanted for a paper
Connected Line Plot Means joined across levels Trends across an ordered factor
Bar Plot Means as bars, with letter groupings available Presentations and extension material

Every plot opens with a Plot Settings panel that controls display mode, titles, axis text, legend, colours and background, plot styling, statistical labels and download settings. A note at the top of the tab explains that you are viewing plots for a single character, and that the settings icon switches between single and multiple character views. Downloads are offered as PNG, JPEG, TIFF, PDF or SVG - choose TIFF or SVG when a journal asks for high resolution, and PNG for everything else.

Statistical Labels puts the letters on the figure

The Statistical Labels section of the settings panel adds the post-hoc letter groupings directly onto the bars or points, so a single figure carries both the means and their separation. That is usually the figure you want in a results section, because it saves the reader from having to hold a table and a chart in mind at once.

13 Advanced plots and diagnostics

The Advanced Plots tab carries twelve further displays. Some are presentation graphics; several are diagnostics that speak directly to the assumptions in Section 4. Figure 10 shows the tab with Interaction Plot I drawn for FactorA × FactorB on Yield.

Figure 10: The Advanced Plots tab: twelve plot types, with Interaction Plot I displayed
Plot What it is for
Interaction Plot, Interaction Plot II The visual counterpart of the interaction F test - parallel lines mean no interaction, crossing lines mean a strong one
Summary Plot A compact overview of all responses at once
Raincloud, Advanced Raincloud Distribution, individual points and summary statistics in one figure
Circular Plot Many treatment combinations arranged radially
QQ Plot Diagnostic - checks the normality assumption
Distribution Plot Diagnostic - the shape of each response
Pair Plot, Correlation Plot Relationships between the responses, not between treatments
3D Scatter Plot, 3D Scatter + Line Three variables at once

The interaction plots are worth dwelling on, because they are where a pooled analysis becomes legible. Choose which pair of factors to cross in Factors to plot and which response in Select Response Variable. For Yield in Figure 10 the lines fall in much the same direction in each panel, which is the visual form of the non-significant interactions read in Section 11.2. When you reach a response where an environment interaction is significant, this is the figure that shows you what it looks like.

Use the diagnostics before you trust the ANOVA

The QQ Plot and Distribution Plot exist to be looked at, not to be skipped. If the QQ plot’s points bend systematically away from the reference line, the normality assumption is strained and the p-values in Section 11 are approximate at best. That is the moment to consider a transformation from Section 10 - and the moment to do it is before you write the conclusions, not after.

14 Looking at all traits together: the PCA index

Seven separate ANOVAs answer seven separate questions. They do not answer the question a breeder or agronomist usually has, which is which treatment is best overall. The Multivariate tab addresses that with a principal component analysis, shown in Figure 11.

Figure 11: The PCA-based index score on the Multivariate tab, with the eigenvalue table

Principal component analysis reduces many correlated traits to a few independent components, and the module uses them to build a single index score by which treatment combinations can be ranked. Press Click here for PCA Index to compute it. The eigenvalue table reports, for each component, its eigenvalue, the percentage of total variance it explains, and the cumulative percentage.

For the working example the first component carries an eigenvalue of 2.49 and explains 35.53% of the variation; the second adds 26.53%, bringing the cumulative total to 62.06%; the third brings it to 81.39%. Four components reach 93.04%. The remaining three contribute almost nothing - PC7 has an eigenvalue of 0.01 - which is the usual pattern and the reason the technique is worth applying: seven traits are doing the work of about three independent dimensions.

Two conditions attached to the index

The module states both on the tab. PCA needs at least two characters to have any relationships to analyse, so the index is unavailable for a single-response file. And if an Aitken transformation was applied during pooling (Section 3), the PCA is computed on the transformed treatment means rather than the raw ones - which is correct, but is worth knowing when you report the index.

An index is a judgement, not a measurement

The index weights traits by how much variance they contribute, not by how much you care about them. A trait that varies a lot will dominate the first component whether or not it is the trait you are selecting for, and traits measured on very different scales can distort the result. Use the index to shortlist, then go back to the individual ANOVAs in Section 11 to justify the choice.

The table can be copied to the clipboard or exported as CSV, Excel or PDF using the buttons beneath it.

The numbers are in. Section 15 shows what the module makes of them.

15 Interpretation

The Interpretation tab writes the results out in prose. Tick the confirmation box - “I’m not a robot and I have checked that on running analysis there was no error reported” - and press Click here for interpretation. Figure 12 shows the output for the working example.

Figure 12: The Interpretation tab: an automatically written account of the analysis

The narrative opens by restating the design it analysed - a pooled two-factor factorial in RCBD with 4 blocks, FactorA at levels S1 and S2, FactorB at C1 and C2, across locations A and B under Model 4, giving 8 treatment combinations each replicated 4 times - and lists the seven characters studied. It then records the Bartlett check and confirms that no Aitken transformation was applied, before working through the effects one at a time.

For this dataset it reports no significant differences for Factor A, for Factor B, for Location, or for the Factor A × Factor B interaction, on any character, and notes that no pairwise comparison is performed where no significant difference was detected. Two findings survive, both on Char2. The Location × Factor A interaction is significant at p = 0.02: S2 at location B has the highest mean at 1.52 ± 0.33 and S2 at location A the lowest at 1.18 ± 0.35, with S2×B on par with S1×A, and S2×A on par with S1×A and S1×B. The three-way Location × Factor A × Factor B interaction is significant at p = 0.03, where S2×C1×B is highest at 1.70 ± 0.37 and S1×C1×B lowest at 1.08 ± 0.20.

The passage closes with the software citation - RAISINS (R & AI Solutions in INferential Statistics), citing Hisham et al. (2025) and R Core Team (2024) - followed by the references for the methods used. A Copy button places the whole text on the clipboard, and Stop halts generation if you launched it by mistake.

Read it, then check it against the tables

The interpretation is generated from the results the module just computed, so it will not contradict them. It is still your analysis and your paper. Two habits are worth keeping: confirm that every number quoted in the prose appears in the corresponding table, and satisfy yourself that the interpretation fits your experiment - the text can tell you that S2×C1×B had the highest mean for Char2, but only you know whether that combination is agronomically sensible or an artefact of one unusual block.

16 Chat with your data using RA-One

RA-One is the assistant reached from the RA-One tab in the top navigation bar, or from the small robot icon that floats at the bottom-right of every screen. You met it in Section 8.4 building a data template; once an analysis has run, it can also answer questions about that analysis in ordinary language.

The difference between RA-One and the Interpretation tab is the direction of travel. The Interpretation tab writes one comprehensive account of everything, whether or not you wanted all of it. RA-One waits for a question and answers that question - “why is the three-way interaction significant for Char2 but not for Yield?”, “which treatment combination should I recommend at location B?”, “what does the Bartlett p-value of 0.09 mean for my analysis?” Its answers are grounded in the results your run produced, not in general statistical advice.

The sidebar keeps your conversations, so a line of questioning you began yesterday is still there today, and New conversation starts a fresh thread when you move to a different dataset. When no analysis has been run, the panel notes that there is no analysis context yet - a reminder that RA-One is most useful after Section 11, when it has real output to reason about.

One assistant, four jobs

RA-One will build you a data template before the trial (Section 8.4), explain a control you are unsure about while you set the analysis up, interpret a table once the analysis has run, and help you word a result for a methods or results section. It is the same assistant in all four roles - the only thing that changes is what you ask it.

17 FAQs

The FAQs tab collects short answers to the questions this module attracts most often, shown in Figure 13.

Figure 13: The FAQs tab

Six entries are listed: how to prepare and upload a file, which model you need to choose, what Bartlett’s test is, the transformation algorithm used, how to master plots in RAISINS, and more on the PCA-based index score. The second and third are the ones worth reading before your first real analysis, because between them they cover the two decisions that shape every number in the output - the fixed-or-random question from Section 2, and the homogeneity check from Section 3.

18 View data

The View Data tab shows the file exactly as the module read it, colour-coded so that problems are visible at a glance. Figure 14 shows the working example passing its health check.

Figure 14: The View Data tab, with the uploaded file health check

The rule is simple, and the instructions panel states it: factor columns should be highlighted in yellow or green, and all numerical values should appear in green. If both conditions hold, your file is in good health. In Figure 14 the Location, FactorA and FactorB columns are yellow, and Block together with Yield and Char1 to Char6 are green - which is exactly right, because those are the numeric columns.

A yellow response column is the warning you must not ignore

If a column that should be numeric appears in yellow, the module has read at least one of its values as text. The instructions name the usual causes: a space between numbers, or an incorrect decimal point. Other frequent culprits are a comma used as a decimal separator, a stray character left over from copying, or a placeholder such as NA or a hyphen typed into a blank cell. Find the offending cell and fix it in the source file, then re-upload. Analysing a file in this state will either fail outright or silently drop rows.

The Show entries selector at the top left controls how many rows are displayed at a time, and the arrows beside each column heading sort by that column - a quick way to check that every block number appears the expected number of times at every location.

19 Wrapping up

The question a pooled two-factor RBD exists to answer is not do these treatments differ - a single-site factorial answers that. It is does the answer travel. Everything in this tutorial serves that question: the four models in Section 2, which decide whether you are generalising to a population of environments or describing the ones you tested; the Bartlett check in Section 3, which decides whether the environments may be combined at all; and above all the environment interaction terms in Section 11, which are where the answer actually lives.

Read them in that order and the analysis is straightforward. Significant environment interactions mean your recommendation has to be site-specific, and the mean tables tell you what to recommend where. Non-significant ones - the case in our working example for Yield - mean the treatment effects you measured were stable across the environments you sampled, and may be stated generally for that range of conditions.

If your design is not quite this one, RAISINS carries neighbouring modules for it. A single treatment factor pooled over environments belongs in Pooled RCBD; a factorial in completely randomized rather than blocked layout belongs in Pooled 2FCRD; and a main-plot/sub-plot arrangement carried over environments belongs in Pooled Split Plot (2,1). A single-environment two-factor factorial needs no pooling at all and belongs in the ordinary Two Factor RBD module.

Before you write it up

Record four things in your methods section, all of which the module has already told you: which of the four models you selected and why; the result of Bartlett’s test and whether an Aitken transformation was applied; the post-hoc test and significance level you used; and any transformation applied to a response. The CPRR in Section 5.1 supplies the version and function details that sit behind them.

If something in the module does not behave as this tutorial describes, or you would like a walkthrough with your own data, write to [email protected]. The team is happy to arrange a free online meeting to work through it with you.

Explore

  • Data analysis
  • Feedback

Policies

  • Privacy policy
  • Data policy
  • Refund policy

Contact

  • Contact us
  • Team
  • Statoberry LLP
Statoberry LLP
© 2026 Statoberry LLP. All rights reserved.
Making statistics sweet — www.raisins.live
RAISINS
Ask AI
Ask AI
RAISINS Logo Powered by RAISINS