Key Expoitable Results (KERs)

Browse the complete collection of AtlantECO Knowledge Outputs (KOs) that constitute the project's Key Exploitable Results (KERs). Use the available filters to explore KOs and quickly find the tools, methodologies, data sets, research articles, policy briefs and other project outcomes that are most relevant to your interests.

AtlantECO-KER-AM-1

Global maps of the Ocean microbiome: Presence probability (0-1) of 53 Operational Taxonomic Units of the phylum Picozoa (pOTUs) in the epipelagic layer (0-200 m), projected under contemporary (2005-2017) environmental conditions, using an ensemble of hab…

This collection of global maps provides global projections of habitat suitability index (HSI) estimated for each of 53 different picozoa operational taxonomic units (pOTUs), as described in Huber et al. 2024. For each pOTU-level map, mean, max, min, stdev, and coefficient of variation of HSI values computed across several algorithms are provided. We developed the single pOTU HSI maps from 2,394 eDNA samples relative to the pico-size fraction (i.e. 0.2 to 5 μm), which were retrieved from the EukBank database. Picozoa samples were classified based on their depth into epipelagic (0-200 m), mesopelagic (200-1000 m) and bathypelagic (>1000 m) samples. To estimate the HSI, we used species distribution models based on a presence/absence design; among all retrieved samples of pico-eukaryotes, georeferenced samples where a specific pOTU was detected were treated as presences, while sampling locations where that pOTU was not detected were treated as (pseudo-) absences. For each pOTU, we contrasted average multi-year depth-specific environmental conditions at presence vs absence points using an ensemble of modelling algorithms, including Generalized Linear Models (GLM), Artificial Neural Network (ANN), Boosted Regression Trees (BRT) and Random Forest (RF). Models were calibrated using multi-year average conditions across epipelagic, mesopelagic and bathypelagic layers. HSI maps were projected over epipelagic conditions only, due to lack of sufficient data coverage for deeper layers. Environmental predictors were retrieved from the World Ocean Atlas 2018.
KER category analysis & modelling
Target user science
AtlantECO-KER-AM-1

Global maps of the Ocean microbiome: Taxonomic diversity of prokaryotes (bacteria and archaea) in the surface mixed layer, projected under contemporary (2005-2012) environmental conditions, based on the modelled habitat suitability of marker gene-based O…

This collection of global maps provides global marine prokaryotic diversity using a standardized ensemble pipeline that integrates metagenomic profiles with environmental predictors. Metagenomic samples from the Ocean Microbiomics Database were processed into domain- and class-level taxonomic profiles, and diversity indices (richness, Shannon, Chao1) were estimated from rarified mOTU counts. Multiple algorithms—Generalized Linear and Additive Models, Random Forests, Boosted Regression Trees, Support Vector Machines, and shallow neural networks—were trained and tuned using five-fold cross-validation. Models were evaluated using root mean square error and R², with only well-performing models (R² ≥ 0.25) retained. Ensemble predictions were derived as the mean across all successful models, while associated uncertainty was quantified as the standard deviation among them. The resulting dataset contains annual, global projections of domain- and class-level prokaryotic diversity indices (richness, Shannon, Chao1). Each netCDF file provides both the ensemble average and standard deviation, offering spatially explicit estimates alongside model-based uncertainty.
KER category analysis & modelling
Target user science
AtlantECO-KER-AM-1

Global maps of the Ocean microbiome: Presence probability (0-1) of 20 Operational Protein Units (OPUs) in the epipelagic layer (0-200 m), projected under contemporary (2005-2017) environmental conditions, using an ensemble of habitat suitability models.

This collection of global maps is based on protein data from 1,379 metagenomes (from AtlantECO-BASEv2), we define Operational Protein Units (OPUs) through a sensitive heuristic clustering strategy that captures both known and unknown amino acid sequences, including those associated with "functional dark matter" often missed by traditional analyses. Based on abundance and occurrence, 20 representative OPUs were selected for the modelling process. These OPUs were assigned to different functional pathways. For each of 20 OPUs we ran species distribution models to predict their global distribution. We modeled using four models- Generalized Linear Models (GLM), Artificial Neural Network (ANN), Boosted Regression Trees (BRT), Random Forest (RF), resulting in 80 maps. With this, with the ensemble approach, we used the mean of the suitability among the four models, resulting in the potential distribution maps. The OPUs were modeled to sunlit ocean regions ( 200 meters). The models showed different OPUs distribution patterns: widely, polar, no polar, tropical, temperate and sub polar distributions. These resulting maps are stored in the folder “SDM OPUs”. We used species distribution models to predict the global distribution of 20 OPUs well represented across all the ocean basins.
KER category analysis & modelling
Target user science
AtlantECO-KER-AM-1

Global maps of the Ocean microbiome: Abundance of 17 groups of autotrophs in the surface mixed layer, available as monthly climatologies projected under contemporary (2005-2012) environmental conditions, using an ensemble of habitat suitability models.

This collection of global maps provides a global ensemble modeling framework allying cell abundance dataset derived from empirical quantification of phytoplankton groups in metagenomic samples with global environmental variables. Cell abundances of phytoplankton groups were estimated from the quantification of the psbO-gene in metagenomics samples from the Tara Oceans and Tara Pacific expeditions. Phytoplankton communities were resolved into 16 taxonomic groups. Separate models were trained for each of 17 responses (16 groups + total phytoplankton), using only samples with complete predictors and targets. Hyperparameters (tree number, minimum leaf size) were tuned by Bayesian optimization. Optimal configurations were determined via 5-fold cross-validation, balancing predictive accuracy and data use. Performance was quantified using RMSE and R². Final models were retrained with all valid data and applied to the global prediction grid. The ensemble models relied on a suite of key environmental predictors, sea surface temperature (SST), sea surface salinity (SSS), chlorophyll-a, iron, nitrate, phosphate, and silicate. Model predictions were made over monthly global environmental grids, with phytoplankton maps generated at 1° × 1° resolution. We computed Mahalanobis distances between environmental conditions at each grid point and the training dataset to identify regions outside the empirical environmental envelope and quantify prediction confidence and extrapolation risk. In parallel, we conducted sensitivity analyses by varying subsets of the data, quantifying how training data composition influences global predictions.
KER category analysis & modelling
Target user science
AtlantECO-KER-AM-1

Global maps of the Ocean microbiome: Taxonomic and functional diversity of copepods in surface waters (0-10 m), available as monthly climatologies projected under contemporary (2012–2031) and future, end-of-century (2081–2100) environmental conditions, u…

This collection of global maps provides global gridded estimates of the functional diversity (FD) of marine copepod assemblages. FD indices were estimated by combining species distribution models (SDMs) with species-level functional traits (i.e., body size, trophic group, feeding mode, myelination and spawning mode) for > 300 copepod species. The dataset includes two complementary products: a contemporary baseline representing average copepod FD patterns for the present-day ocean and a future projection representing expected end-of-century (2081–2100) average changes in FD under a high-emission climate scenario (RCP8.5). Both products are provided as monthly climatologies at 1° × 1° spatial resolution for the surface ocean (0–10 m). Multiple FD indices were computed to cover multiple aspects of functional diversity including: functional richness, evenness, divergence, dispersion, and trait dissimilarity. These data allow researchers to investigate how copepod trait composition is structured across the global ocean today, and how it may be reshuffled in response to anthropogenic climate change. This dataset offers a valuable resource to assess potential impacts of changing zooplankton communities on marine productivity, carbon cycling, and climate feedbacks.
KER category analysis & modelling
Target user science
AtlantECO-KER-AM-1

Global maps of the Ocean microbiome: Taxonomic and functional diversity of prokaryotes in the epipelagic layer (0-200 m), projected under contemporary (2005-2017) environmental conditions, based on the modelled habitat suitability of 1,210 marker gene-ba…

This collection of global maps is based on metagenomic data collected during multiple oceanographic expeditions, including Tara Oceans, Malaspina, GO-SHIP, bioGEOTRACES, and OSD2014. Sample metadata, including geographic coordinates and sampling depth, were curated and merged with the environmental raster stack using bilinear extraction (terra::extract). Only samples with complete environmental metadata and valid coordinates were retained, and with samples from the 0.22-3 µm size fraction. The metagenomic data is from AtlantECO-BASEv2, a database developed within the AtlantECO project. All the metagenomic data were analyzed using the version 5.0 of MGnify’s pipeline. Taxonomic profiles were generated by running mOTUs on MGnify. For each sample (n = 1,210), tables containing mOTUs taxonomic assignments were downloaded and compiled into a single table. Diversity was then calculated from this dataset using the Shannon and Simpson indices in R (vegan package). We conducted our analysis separately, using either data from all expeditions pooled together or using only GO-SHIP data, as this was the only dataset covering the latitudinal gradient. We focused on the sunlit zone due to its ecological relevance for primary production and microbial functional diversity. To predict Simpson and Shannon diversity across global marine expeditions, we applied a species distribution modeling (SDM) framework using four algorithms within R, using the package “h2o” (v3.42.0.2). All models were trained to predict Simpson or Shannon functional/taxonomic diversity based on a selected set of non-collinear environmental predictors: mixed layer depth (MLD), chlorophyll (Chl), salinity (Sal), silicate (Si), temperature (T), and the excess of nitrates over phosphates (N*) according to Redfield ratio [NO3−] − 16[PO43−]. Models were trained on the final dataset using 5-fold cross-validation, for four algorithms (Artificial Neural Network, Random Forest, Generalized Linear Model, Gradient Boosting Machine). Predictions were projected using each trained model. Extracted predictions at sampling locations were compared against observed Simpson/Shannon diversity values using Pearson correlation, RMSE, and R². An ensemble prediction was calculated as the pixelwise mean across the four models.
KER category analysis & modelling
Target user science