%%{
init: {
'themeVariables': {
'lineColor': '#757575'
}
}
}%%
flowchart LR
n1["Environmental data"]
n2["Model"]
n3["Species data"]
n4["Range map"]
classDef commonStyle color:#000000,fill:#D9D9D9,stroke:#737373,stroke-width:1px
class n1,n3 commonStyle
n1 --> n2
style n2 fill:#016DD7,stroke:#016DD7,color:white
n3 --> n2
n2 --> n4
style n4 fill:#00D0B1,stroke:#00D0B1,color:white
linkStyle 0,1,2 stroke:#757575,stroke-width:0.5px,marker-width:0.5px
3 Species distribution models
3.1 Introduction
MPA Europe will produce a priority map of conservation areas in Europe, designed to protect the highest proportion of biodiversity. Consequently, a critical aspect of the project is understanding species distributions. Despite the growing number of available records in OBIS and other databases, significant gaps remain in our knowledge of species ranges. To address this, MPA Europe is utilizing species distribution models (SDMs) to generate high-resolution (~5 km grid) species range maps.
SDMs are statistical models used to predict the potential geographic distribution of a species based on the relationship between observed species occurrences and environmental variables. These models estimate areas where environmental conditions are suitable for the presence of a species, producing maps of predicted habitat suitability1. You can find more details about SDMs in the review of Elith and Leathwick (2009), and more broadly ecological niche modelling in Peterson et al. (2011).
In a simplified look, species data (in the form of information of where the species occur, and ideally where it is absent) and environmental data (e.g. temperature, salinity) are used as input in a model (see example here). A statistical model is an abstraction2 that help us to understand the relationship between variables. Many algorithms can be used to produce a species distribution model, depending on the available data. Once the model is fitted, it can then be used to predict the suitability in new locations according to the local environmental conditions.
Of course, this is just a simplification. Many steps are necessary to obtain the data, process it, fit the model, evaluate the models and finally make predictions. In this project we followed those steps:
- Obtain species list (all species occurring in the study area)
- Obtain occurrence data for the selected species
- Apply quality control procedures to the occurrence data
- Obtain environmental data
- Obtain ecological information (in our case, the mode of life of species, e.g. benthic or pelagic)
- Fit models
- Evaluate models
- Produce map predictions
- Post-process the predictions (ensembles, uncertainty, thermal ranges, diversity metrics and habitat maps)
Steps 1 to 3 and 5 are detailed in the next chapter, and steps 4 and 6 to 9 are summarized at the end of this one. But first we detail the statistical framework adopted in the WP3.
3.2 Framework used in this project
All modeling was done considering a point process framework. Spatial Point Process Models (PPMs) are used to model any type of events that arise as points in space (Renner et al. 2015). Those points have random number and location, but are related to an underlying process. It was only recently that PPMs were more intensively applied to SDMs (Renner et al. 2015; Warton and Shepherd 2010), although widely used in spatial analysis (examples of applications range from disease mapping to detection of landslides; e.g. Lombardo et al. (2018)).
The interest in PPMs for SDMs was mainly driven by the challenge of modelling presence-only data, when all the information available is the geographical locations where the species was recorded (Fithian and Hastie 2013). Usually, one would sample pseudo-absences (points that should reflect places where the species is absent), but the number and place of pseudo-absences can have a great influence on models (Barbet-Massin et al. 2012). Also there is no clear justification to pseudo-absences: the location was not sampled and thus it is impossible to really know if the species is absent. Instead, on a PPM framework, quadrature (or background) points are used only as a device to describe the available environmental conditions (Fithian and Hastie 2013). Points are sampled at random and the number of points can be chosen based on the accuracy of the likelihood estimate (Renner et al. 2015). In theory, all the points of the environment could be used, but this would increase computational time.
In the case of a PPM, and considering that the points does not present a strong bias, the resulting intensity of the point process can be used as a proxy to the suitability of the habitat (under the assumption that the species is more easily recorded where the habitat suitability is higher) (Renner et al. 2015). One important point to note is that the results of any PPM model (and in general, any presence-only model) should not be interpreted as probability of occurrence, but instead as a relative occurrence rate (following (Merow et al. 2013)).
PPMs are closely related to the widely used Maxent (Renner and Warton 2013). Indeed, even the pseudo-absence (or presence-background) modeling can approximate a Poisson point process model, under certain conditions (Renner et al. 2015; Warton and Shepherd 2010).
3.3 Modelling workflow
The workflow is operationalized in the mpaeu_sdm repository (model_fit.R and functions/model_species.R), while the underlying functions are provided by the obissdm package. Below is a summary of what was done for each species.
3.3.1 Environmental data
Environmental layers were obtained from Bio-ORACLE for the baseline period (2000-2020) and for five future scenarios (SSP1-2.6, SSP2-4.5, SSP3-7.0, SSP4-6.0 and SSP5-8.5), for two decades: 2040-2050 (dec50) and 2090-2100 (dec100). Both surface and mean-depth layers were used, according to the mode of life of the species. Additional layers (e.g. wave fetch, rugosity, distance to coast) were also included.
Species were split in four groups (defined in sdm_conf.yml): seabirds, mammals, photosynthesizers (Chromista and Plantae) and others. Each group has its own set of candidate variables (e.g. sea temperature, salinity, sea ice concentration, bathymetry, wave stress, oxygen, nitrate and light, depending on the group), which was defined after testing for collinearity (prepare_envdata.R). For each group, alternative sets of variables (hypotheses, e.g. basevars, coastal and complexity) were compared using a Random Forest, and the best one was used for the species. Species restricted to shallow areas (up to 50 m) do not use bathymetry and distance to coast.
3.3.2 Fitting
Models are fitted only for species with at least 30 occurrence points after quality control. The algorithms used were Maxent, Random Forest (down-sampled classification) and XGBoost, all under the point process framework described above. Quadrature points were sampled at 1% of the available environmental cells (with a minimum of 10,000). Before fitting, we tested for spatial autocorrelation in the records and, when present, thinned the points accordingly.
For each species, the environmental data used for fitting was limited to the marine realms where the species occurs (plus adjacent ones) and to the depth range of the records (with a buffer of 50 m). Predictions are made for the whole extent of the environmental layers, and masks are provided so that users can restrict them (see how to interpret SDM outputs).
3.3.3 Evaluation
Models were evaluated with cross-validation using spatial blocks. The main metric is the Continuous Boyce Index (CBI), as the data is presence-only. A model was considered of good quality (and thus retained) if its mean CBI was at least 0.3. Species for which no algorithm reached this threshold have no SDM available. Additional metrics (e.g. AUC and TSS) are provided, and where an independent dataset was available it was also used to evaluate the model. A post-evaluation was performed comparing the predictions with the thermal niche of the species, niche equivalency and hypervolume estimates.
3.3.4 Predictions and ensemble
Good models were used to predict the current and future scenarios. When more than one algorithm produced a good model, an ensemble (median) with an associated uncertainty layer was also produced. Variable importance and partial response curves are saved with each model. Several thresholds for binarization are also stored. Details on the files produced are in the data use chapter.
3.3.5 Post-processing
After the models were fitted, the following products were derived: bootstrapped predictions (to assess uncertainty; post_bootstrap.R), thermal range maps (model_thermal.R), and habitat maps of biogenic habitats generated by stacking the SDMs of habitat-forming species (post_habitat.R). Results were organized in a STAC catalogue (post_generate_stac.R).
Note that SDM is a broad term, covering many statistical approaches used to model the distribution of species according to environmental variables. Some methods are capable of detecting the probability of occurrence of a species, while others offer just a relative suitability score.↩︎
This is an important aspect: models are designed to help us understand and capture general trends, providing insights into natural phenomena or enabling predictions. However, models are not perfect; many fine-scale aspects that contribute to the true (and often unknown) distribution of species are not fully represented.↩︎