The package gdmmTMB provides functions for fitting and interrogating
Generalised Dissimilarity Mixed Models (GDMM) using maximum likelihood
estimation, Laplace approximation and (optionally) Bayesian
Bootstrapping.
The package can be installed from GitHub using
devtools::install_github:
devtools::install_github('dhercar/gdmmTMB')The package gdmmTMB was designed primarily for the analysis of
ecological datasets where a set of environmental variables is used to
explain variability in species composition.
For this example, we use two datasets available as part of the gdm
package. We start with two matrices:
sp: A community matrix (site x species)env: A matrix of environmental covariates.
library(tidyverse)
library(gdmmTMB)
sp <- gdm::southwest %>% select(site, species) %>%
group_by(site, species) %>%
summarise(n = 1) %>%
pivot_wider(id_cols = site, names_from = species,values_from = n, values_fill = 0) %>%
arrange(site)
env <- gdm::southwest %>% select(-c(species)) %>%
distinct() %>% arrange(site)We use bio6, bio15, and bio19 as explanatory variables. Monotonic
I-splines for dissimilarity gradients (diss_formula) are specified via
isp(x) and mono = TRUE. The resulting non-linear transformations
reflect the degree of directional community change along each
environmental gradient.
m <- gdmm(Y = sp[,-1],
X = env,
diss_formula = ~ isp(bio6) + isp(bio15) + isp(bio19),
family = 'binomial',
method = 'jaccard',
n_boot = 100, # should be at least 1000
bboot = TRUE,
mono = TRUE)## Running bayesian bootstrapping on 14 cores ...done
We specify a binomial error family to model Jaccard dissimilarities,
where the dissimilarity between observations
By specifying a binomial distribution, the Jaccard dissimilarity index
is modelled as the probability
The summary method in gdmmTMB provides general information about the
model, including estimated coefficients, credible or confidence
intervals, and common metrics used for model selection such as AIC and
BIC.
summary(m)## --------------------------------------------------------------------------------
## BBGDMM SUMMARY
## --------------------------------------------------------------------------------
##
## call:
##
## gdmm(Y = sp[, -1], X = env, diss_formula = ~isp(bio6) + isp(bio15) + isp(bio19), mono = TRUE, family = "binomial", method = "jaccard", bboot = TRUE, n_boot = 100)
##
## --------------------------------- Coeff. Table ---------------------------------
##
## Estimate Std.Err. 2.5% 50% 97.5% pseudo-Z
## diss: isp(bio6)1 3.217e-01 1.355e-01 8.045e-02 3.259e-01 6.025e-01 2.3740
## diss: isp(bio6)2 1.341e-02 4.014e-02 1.010e-08 1.672e-08 1.598e-01 0.3340
## diss: isp(bio6)3 1.750e-02 4.444e-02 1.000e-08 1.109e-08 1.448e-01 0.3937
## diss: isp(bio15)1 8.631e-01 1.443e-01 5.817e-01 8.552e-01 1.143e+00 5.9792
## diss: isp(bio15)2 1.210e-08 1.360e-08 1.000e-08 1.003e-08 1.341e-08 0.8896
## diss: isp(bio15)3 1.246e-08 1.990e-08 1.000e-08 1.000e-08 1.253e-08 0.6259
## diss: isp(bio19)1 1.876e+00 1.440e-01 1.609e+00 1.860e+00 2.143e+00 13.0284
## diss: isp(bio19)2 1.412e-08 3.178e-08 1.000e-08 1.000e-08 1.240e-08 0.4442
## diss: isp(bio19)3 1.627e-08 5.185e-08 1.000e-08 1.000e-08 2.131e-08 0.3139
## (Intercept) 8.789e-01 7.069e-02 7.537e-01 8.739e-01 1.012e+00 12.4331
## pseudo-Pval
## diss: isp(bio6)1 < 0.01 ***
## diss: isp(bio6)2 < 0.01 ***
## diss: isp(bio6)3 < 0.01 ***
## diss: isp(bio15)1 < 0.01 ***
## diss: isp(bio15)2 < 0.01 ***
## diss: isp(bio15)3 < 0.01 ***
## diss: isp(bio19)1 < 0.01 ***
## diss: isp(bio19)2 < 0.01 ***
## diss: isp(bio19)3 < 0.01 ***
## (Intercept) < 0.01 ***
##
## CAUTION: P-values for dissimilarity components (diss) should not be used for hypothesis testing when "mono = TRUE"
## ---
## signif. codes: '***' <0.001 '**' <0.01 '*' <0.05 '.' <0.1
## ---
##
## Marginal log-likelihood (average): -27848.53
## AIC: 55717.06 , AICc: 55717.11 , BIC: 55780.88
##
## --------------------------------------------------------------------------------
Predictors in ‘diss_formula’ are treated as environmental distances, and
their contribution to the expected dissimilarity depends on the distance
between samples after (optionally) transformations are applied. For
instance, isp(bio6) with mono = TRUE defines a non-linear, monotonic
change in community composition that generates increasing
dissimilarities with increasing distance along the gradient
(dissimilarity gradient).The shape of the resulting function can be
informative of the degree of directional community change that each
environmental variable is responsible for.
The function diss_gradient provides a list of dissimilarity gradients
and their confidence or credible intervals, which can be easily combined
and plotted:
f_x <- diss_gradient(m, CI_quant = c(0.5, 0.95))Note that the argument CI_quant specifies the coverage of the
confidence or credible interval, not the quantiles themselves (e.g.,
0.95 implies quantiles of 0.025 and 0.975). We have therefore specified
50% and 95% credible credible intervals.
library(ggplot2)
do.call(rbind, f_x) %>% ggplot(aes(fill = var)) +
theme_bw() +
geom_line(aes(x = x, y = f_x)) +
geom_ribbon(aes(x = x, ymax = `CI 2.5%`, ymin = `CI 97.5%`), alpha = 0.1) +
geom_ribbon(aes(x = x, ymax = `CI 25%`, ymin = `CI 75%`), alpha = 0.3) +
facet_wrap(~var, scales = 'free_x', strip.position = "bottom") +
ylab('f(x)') +
theme(strip.placement = 'outside',
panel.grid = element_blank(),
strip.background = element_blank(),
aspect.ratio = 1)The plot shows the relationship between each variable (
