Skip to content

Volcano plot ignores the t-test method chosen on the T-test page (uses the package default instead) #377

Description

@shoo99

Volcano plot ignores the t-test method chosen on the T-test page (uses the package default instead)

Package: MetaboAnalystR 4.3.0 (GitHub fd37171)
R: 4.3.1, aarch64-apple-darwin20
Also affects: the public web server (metaboanalyst.ca), Statistical Analysis [one factor] → Volcano plot


Summary

Volcano.Anal() decides its p-value method with

tt.method <- if (identical(fc.method, "limma")) "limma" else formals(Ttests.Anal)$tt.method

formals(Ttests.Anal)$tt.method is the static package default ("limma" since e028f54), not the
method the user actually chose. So when a user selects Student's t-test on the Significance Testing
page and then opens the Volcano plot in Classical mode, the volcano still reports moderated
(limma) p-values
.

The comment directly above that line states the intent:

# The p-value axis must use the same test as the T-test view, otherwise the same dataset
# reports two different p-values for one feature depending on which page it is read from.

That is exactly the situation the current implementation produces, because it reads the default rather
than the user's selection.

The on-screen description on the web Volcano page also disagrees with the behaviour:

Classical – Uses classical FC (raw data) on the x-axis and Student's t-test p-values on the y-axis


Where

R/stats_univariates.R

  • L717–718 — Volcano.Anal(..., fc.method = "classical")
  • L727 — the line above
  • L394–395 — Ttests.Anal(..., tt.method = "limma") (default changed in e028f54, 2026-07-27)

Reproducible example (synthetic data, no external files)

library(MetaboAnalystR)

set.seed(1)
n <- 4; p <- 200
m <- matrix(rlnorm(2 * n * p, 10, 1), nrow = 2 * n, ncol = p)
m[1:n, 1:20] <- m[1:n, 1:20] * 4                        # 20 truly changed features
rownames(m) <- c(paste0("A", 1:n), paste0("B", 1:n))
colnames(m) <- paste0("M", sprintf("%03d", 1:p))
write.csv(data.frame(Sample = rownames(m), Label = rep(c("A", "B"), each = n), m,
                     check.names = FALSE), "input.csv", row.names = FALSE)

mSet <- InitDataObjects("conc", "stat", FALSE, default.dpi = 72)
mSet <- Read.TextData(mSet, "input.csv", "rowu", "disc")
mSet <- SanityCheckData(mSet); mSet <- PreparePrenormData(mSet)
mSet <- Normalization(mSet, "MedianNorm", "LogNorm", "NULL", ratio = FALSE, ratioNum = 20)

cat("package default tt.method:", as.character(formals(Ttests.Anal)$tt.method), "\n")

# user explicitly asks for Student's t-test
mSet <- Ttests.Anal(mSet, nonpar = FALSE, threshp = 0.05, paired = FALSE, equal.var = TRUE,
                    pvalType = "raw", all_results = TRUE, tt.method = "student")
p.student <- mSet$analSet$tt$p.value

# volcano in Classical mode
mSet <- Volcano.Anal(mSet, paired = FALSE, fcthresh = 2.0, cmpType = 0, nonpar = FALSE,
                     threshp = 0.05, equal.var = TRUE, pval.type = "raw",
                     fc.method = "classical")
p.volcano <- 10^(-mSet$analSet$volcano$p.log)

# limma for reference
mSet <- Ttests.Anal(mSet, nonpar = FALSE, threshp = 0.05, paired = FALSE, equal.var = TRUE,
                    pvalType = "raw", all_results = TRUE, tt.method = "limma")
p.limma <- mSet$analSet$tt$p.value

k <- intersect(names(p.student), names(p.volcano))
cat("max |volcano - student| =", signif(max(abs(p.volcano[k] - p.student[k])), 3), "\n")
cat("max |volcano - limma  | =", signif(max(abs(p.volcano[k] - p.limma[k])),   3), "\n")

Output

package default tt.method: limma
Ttests.Anal(tt.method = 'student') -> method stored: classical | significant: 19
Volcano.Anal(fc.method = 'classical')  -> method stored: limma  | significant: 16
Ttests.Anal(tt.method = 'limma')   -> significant: 16

max |volcano - student| = 0.413
max |volcano - limma  | = 5.55e-17

The volcano's p-values are limma's to machine precision, although Student's t-test was requested.
mSet$analSet$tt$tt.method is also silently overwritten from classical back to limma.

Observed on the web server

Real dataset, 5 groups × n = 3, 384 features, median normalisation + log transform,
FC ≥ 2, raw p < 0.05.

Significance Testing page: Student → 85 significant, Limma → 95 significant (same comparison).

Volcano plot, same comparison, both modes — highest point on the y-axis:

comparison max −log10 p, Student max −log10 p, limma max −log10 p shown on Classical volcano
B / A 3.815 2.883 2.89
D / C 3.773 4.930 4.93
D / B 4.644 6.004 6.00
E / C 5.414 6.152 6.15

In all four the Classical volcano matches limma, not Student.

Because only the fold-change axis actually changes between the two modes, switching Classical ↔
Moderated moves the DE count by 0–3 features, which makes the problem easy to miss:

comparison Classical Moderated
B / A 22 ↑ / 28 ↓ = 50 25 ↑ / 28 ↓ = 53
D / C 16 ↑ / 12 ↓ = 28 15 ↑ / 12 ↓ = 27
D / B 7 ↑ / 7 ↓ = 14 7 ↑ / 7 ↓ = 14
E / C 15 ↑ / 25 ↓ = 40 15 ↑ / 25 ↓ = 40

With Student's t-test genuinely applied, the same four comparisons give 33 / 22 / 8 / 35.


Suggested fix

Ttests.Anal() already records the method it used (tt <- list(..., tt.method = tt.method, ...),
L544–546), so the user's actual choice is available on the mSet:

tt.method <- if (identical(fc.method, "limma")) {
               "limma"
             } else if (!is.null(mSetObj$analSet$tt$tt.method)) {
               mSetObj$analSet$tt$tt.method     # what the T-test page actually ran
             } else {
               formals(Ttests.Anal)$tt.method   # nothing run yet -> fall back to default
             }

nonpar and equal.var are already passed through by Volcano.Anal, so the
student / welch / wilcox variants are preserved.

If instead the intended behaviour is that Classical mode always means the ordinary t-test, then
"classical" could be hard-coded there — but in that case the two pages will still disagree whenever
the user picks limma on the T-test page, so reading the recorded method seems preferable.

Either way, it would help to state on the Volcano page which test the y-axis actually used.

Thanks for the package.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions