Search USGSSearch

Geology topics

Simon Nemer Topp

Publications and source records attributed to Simon Nemer Topp.

6 recordsLinked to original sources

Deep learning of estuary salinity dynamics is physically accurate at a fraction of hydrodynamic model computational cost

Salinity dynamics in the Delaware Bay estuary are a critical water quality concern as elevated salinity can damage infrastructure and threaten drinking water supplies. Current state-of-the-art modeling approaches use hydrodynamic models, which can produce accurate results but are limited by significant computational costs. We developed a machine learning (ML) model to predict the 250 mg L −1 Cl − isochlor, also known as the “salt front,” using daily river discharge, meteorological drivers, and tidal water level data. We use the ML model to predict the location of the salt front, measured in river miles (RM) along the Delaware River, during the period 2001–2020, and we compare predictions of the ML model to the hydrodynamic Coupled Ocean–Atmosphere-Wave-Sediment Transport (COAWST) model. The ML model predicts the location of the salt front with greater accuracy (root mean squared error [RMSE] = 2.52 RM) than the COAWST model does (RMSE = 5.36); however, the ML model struggles to predict extreme events. Furthermore, we use functional performance and expected gradients, tools from information theory and explainable artificial intelligence, to show that the ML model learns physically realistic relationships between the salt front location and drivers (particularly discharge and tidal water level). These results demonstrate how an ML modeling approach can provide predictive and functional accuracy at a significantly reduced computational cost compared to process-based models. In addition, these results provide support for using ML models in operational forecasting, scenario testing, management decisions, hindcasting, and resulting opportunities to understand past behavior and develop hypotheses.

Limnology and Oceanography

National-scale remotely sensed lake trophic state from 1984 through 2020

Lake trophic state is a key ecosystem property that integrates a lake’s physical, chemical, and biological processes. Despite the importance of trophic state as a gauge of lake water quality, standardized and machine-readable observations are uncommon. Remote sensing presents an opportunity to detect and analyze lake trophic state with reproducible, robust methods across time and space. We used Landsat surface reflectance data to create the first compendium of annual lake trophic state for 55,662 lakes of at least 10 ha in area throughout the contiguous United States from 1984 through 2020. The dataset was constructed with FAIR data principles (Findable, Accessible, Interoperable, and Reproducible) in mind, where data are publicly available, relational keys from parent datasets are retained, and all data wrangling and modeling routines are scripted for future reuse. Together, this resource offers critical data to address basic and applied research questions about lake water quality at a suite of spatial and temporal scales.

Scientific Data

Train, inform, borrow, or combine? Approaches to process-guided deep learning for groundwater-influenced stream temperature prediction

Although groundwater discharge is a critical stream temperature control process, it is not explicitly represented in many stream temperature models, an omission that may reduce predictive accuracy, hinder management of aquatic habitat, and decrease user confidence. We assessed the performance of a previously-described process-guided deep learning model of stream temperature in the Delaware River Basin (USA). We found lower accuracy (root mean square error [RMSE] of 1.71 versus 1.35°C) and stronger seasonal bias (absolute mean monthly bias of 1.06 vs. 0.68°C) for reaches primarily influenced by deep groundwater as compared to atmospheric conditions. We then tested four approaches for improving groundwater process representation: (a) a custom loss function leveraging the unique patterns of air and water temperature coupling characteristic of different temperature drivers, (b) inclusion of additional groundwater-relevant catchment attributes, (c) incorporation of additional process model outputs, and (d) a composite model. The custom loss function and the additional attributes significantly improved the predictive accuracy in groundwater-dominated reaches (RMSE of 1.37 and 1.26°C) and reduced the seasonal bias (absolute mean monthly bias of 0.44 and 0.48°C), but neither approach could identify holdout groundwater reaches. Variable importance analysis indicates the custom loss function nudges the model to use the existing inputs more efficiently, whereas with the added features the model relies on a broader suite of inputs. This analysis is a substantial step toward more accurately representing groundwater discharge processes in stream temperature models and will improve predictive accuracy and inform habitat management.

Delaware River Basin

Stream temperature prediction in a shifting environment: The influence of deep learning architecture

Stream temperature is a fundamental control on ecosystem health. Recent efforts incorporating process guidance into deep learning models for predicting stream temperature have been shown to outperform existing statistical and physical models. This performance is in part because deep learning architectures can actively learn spatiotemporal relationships that govern how water and energy propagate through a river network. However, exploration of how spatiotemporal awareness and process guidance influence a model's generalizability under shifting environmental conditions such as climate change is limited. Here, we use Explainable Artificial Intelligence (XAI) to interrogate how differing deep learning architectures affect a model's learned spatial and temporal dependencies, and how those learned dependencies affect a model's ability to maintain high accuracy when applied to unseen environmental conditions. Using the Delaware River Basin in the northeastern United States as a test case, we compare two spatiotemporally aware process-guided deep learning models for predicting stream temperature (a recurrent graph convolution network—RGCN, and a temporal convolution graph model—Graph WaveNet). Both models achieve equally high predictive performance when testing data are well represented in the training data (test root mean squared errors of 1.64°C and 1.65°C); however, Graph WaveNet significantly outperforms RGCN in 4 out of 5 experiments where test partitions represent different types of unseen environmental conditions. XAI results show that the architecture of Graph WaveNet leads to learned spatial relationships with greater fidelity to physical processes, and that this fidelity improves the generalizability of the model when applied to shifting and/or unseen environmental conditions.

Delaware River Basin

Daily surface temperatures for 185,549 lakes in the conterminous United States estimated using deep learning (1980–2020)

The dataset described here includes estimates of historical (1980–2020) daily surface water temperature, lake metadata, and daily weather conditions for lakes bigger than 4 ha in the conterminous United States ( n = 185,549), and also in situ temperature observations for a subset of lakes ( n = 12,227). Estimates were generated using a long short-term memory deep learning model and compared to existing process-based and linear regression models. Model training was optimized for prediction on unmonitored lakes through cross-validation that held out lakes to assess generalizability and estimate error. On the held-out lakes with in situ observations, median lake-specific error was 1.24°C, and the overall root mean squared error was 1.61°C. This dataset increases the number of lakes with daily temperature predictions when compared to existing datasets, as well as substantially improves predictive accuracy compared to a prior empirical model and a debiased process-based approach (2.01°C and 1.79°C median error, respectively).

Limnology & Oceanography: Letters

The AEMON-J “Hacking Limnology” workshop series & virtual summit: Incorporating data science and open science in aquatic research

Following the 2020 “Virtual Summit: Incorporating Data Science and Open Science in Aquatic Research” (DSOS; Meyer and Zwart 2020 ), a grassroots group of scientists convened the 2nd Virtual DSOS Summit on 22–23 July 2021. DSOS combined forces with the Aquatic Ecosystem MOdeling Network - Junior (AEMON-J; https://github.com/aemon-j) to host a 4-d “Hacking Limnology” Workshop Series prior to the summit (13–16 July 2021). The aim was to focus more deeply on skill development and networking among early career researchers (ECRs), both of which are key to growing a workforce of data-intensive aquatic scientists (López Moreira M et al. in press; Meyer et al. 2021 a ). To support ECRs further, we hosted a virtual job board, where participants could note if they were either looking for employment or hiring for a position. Like the 2020 summit, there was high enthusiasm for both the summit and the workshops. In total, 686 people from over 50 countries registered for the AEMON-J Workshop Series and the DSOS Summit. Countries with the highest number of registrants included the United States (41%), Nigeria (20%), Canada (6%), Brazil (6%), and Germany (5%) (Fig. 1). To increase accessibility, there were no registration costs for the workshops and summit, and we centralized introductory training materials, coding scripts, and presentation recordings in one community website (https://aquaticdatasciopensci.github.io/; Fig. 2), which we hope will continue to support the AEMON-J and DSOS communities over time.

Limnology and Oceanography Bulletin