Search USGSSearch

Geology topics

Matthias Schmid

Publications and source records attributed to Matthias Schmid.

5 recordsLinked to original sources

Achieving interpretable machine learning by functional decomposition of black-box models into explainable predictor effects

Machine learning (ML) models are often based on complex black-box architectures that are difficult to interpret. This interpretability problem can hinder the use of ML in fields like medicine, ecology, and insurance, and has boosted research in interpretable machine learning (IML). Here, we propose a novel approach for the functional decomposition of black-box predictions, which is a core concept of IML. This approach replaces the prediction function with a surrogate model consisting of simpler subfunctions, providing insights into the direction and strength of the main feature contributions and their interactions. Our method is based on a concept termed “stacked orthogonality”, which ensures that the main effects capture as much functional behavior as possible. To compute the subfunctions, we combine neural additive modeling with an efficient post-hoc orthogonalization procedure. Our method yielded plausible results in an analysis of stream biological condition in the Chesapeake Bay watershed (United States).

Chesapeake Bay watershed

Explainable machine learning improves interpretability in the predictive modeling of biological stream conditions in the Chesapeake Bay Watershed, USA

Anthropogenic alterations have resulted in widespread degradation of stream conditions. To aid in stream restoration and management, baseline estimates of conditions and improved explanation of factors driving their degradation are needed. We used random forests to model biological conditions using a benthic macroinvertebrate index of biotic integrity for small, non-tidal streams (upstream area ≤200 km 2 ) in the Chesapeake Bay watershed (CBW) of the mid-Atlantic coast of North America. We utilized several global and local model interpretation tools to improve average and site-specific model inferences, respectively. The model was used to predict condition for 95,867 individual catchments for eight periods (2001, 2004, 2006, 2008, 2011, 2013, 2016, 2019). Predicted conditions were classified as Poor, FairGood, or Uncertain to align with management needs and individual reach lengths and catchment areas were summed by condition class for the CBW for each period. Global permutation and local Shapley importance values indicated percent of forest, development, and agriculture in upstream catchments had strong impacts on predictions. Development and agriculture negatively influenced stream condition for model average (partial dependence [PD] and accumulated local effect [ALE] plots) and local (individual condition expectation and Shapley value plots) levels. Friedman's H-statistic indicated large overall interactions for these three land covers, and bivariate global plots (PD and ALE) supported interactions among agriculture and development. Total stream length and catchment area predicted in FairGood conditions decreased then increased over the 19-years (length/area: 66.6/65.4% in 2001, 66.3/65.2% in 2011, and 66.6/65.4% in 2019). Examination of individual catchment predictions between 2001 and 2019 showed those predicted to have the largest decreases in condition had large increases in development; whereas catchments predicted to exhibit the largest increases in condition showed moderate increases in forest cover. Use of global and local interpretative methods together with watershed-wide and individual catchment predictions support conservation practitioners that need to identify widespread and localized patterns, especially acknowledging that management actions typically take place at individual-reach scales.

Chesapeake Bay Watershed

Techniques to improve ecological interpretability of black box machine learning models

Statistical modeling of ecological data is often faced with a large number of variables as well as possible nonlinear relationships and higher-order interaction effects. Gradient boosted trees (GBT) have been successful in addressing these issues and have shown a good predictive performance in modeling nonlinear relationships, in particular in classification settings with a categorical response variable. They also tend to be robust against outliers. However, their black-box nature makes it difficult to interpret these models. We introduce several recently developed statistical tools to the environmental research community in order to advance interpretation of these black-box models. To analyze the properties of the tools, we applied gradient boosted trees to investigate biological health of streams within the contiguous USA, as measured by a benthic macroinvertebrate biotic index. Based on these data and a simulation study, we demonstrate the advantages and limitations of partial dependence plots (PDP), individual conditional expectation (ICE) curves and accumulated local effects (ALE) in their ability to identify covariate–response relationships. Additionally, interaction effects were quantified according to interaction strength (IAS) and Friedman’s H 2 "> H 2 statistic. Interpretable machine learning techniques are useful tools to open the black-box of gradient boosted trees in the environmental sciences. This finding is supported by our case study on the effect of impervious surface on the benthic condition, which agrees with previous results in the literature. Overall, the most important variables were ecoregion, bed stability, watershed area, riparian vegetation and catchment slope. These variables were also present in most identified interaction effects. In conclusion, graphical tools (PDP, ICE, ALE) enable visualization and easier interpretation of GBT but should be supported by analytical statistical measures. Future methodological research is needed to investigate the properties of interaction tests. Supplementary materials accompanying this paper appear on-line.

Journal of Agricultural, Biological, and Environme

A random forest approach for bounded outcome variables

Random forests have become an established tool for classication and regres- sion, in particular in high-dimensional settings and in the presence of non-additive predictor-response relationships. For bounded outcome variables restricted to the unit interval, however, classical modeling approaches based on mean squared error loss may severely suer as they do not account for heteroscedasticity in the data. To address this issue, we propose a random forest approach for relating a beta dis- tributed outcome to a set of explanatory variables. Our approach explicitly makes use of the likelihood function of the beta distribution for the selection of splits dur- ing the tree-building procedure. In each iteration of the tree-building algorithm it chooses one explanatory variable in combination with a split point that maximizes the log-likelihood function of the beta distribution with the parameter estimates de- rived from the nodes of the currently built tree. Results of several simulation studies and an application using data from the U.S.A. National Lakes Assessment Survey demonstrate the properties and usefulness of the method, in particular when com- pared to random forest approaches based on mean squared error loss and parametric regression models.

Journal of Computational and Graphical Statistics

Developing and testing temperature models for regulated systems: a case study on the Upper Delaware River

Water temperature is an important driver of many processes in riverine ecosystems. If reservoirs are present, their releases can greatly influence downstream water temperatures. Models are important tools in understanding the influence these releases may have on the thermal regimes of downstream rivers. In this study, we developed and tested a suite of models to predict river temperature at a location downstream of two reservoirs in the Upper Delaware River (USA), a section of river that is managed to support a world-class coldwater fishery. Three empirical models were tested, including a Generalized Least Squares Model with a cosine trend (GLScos), AutoRegressive Integrated Moving Average (ARIMA), and Artificial Neural Network (ANN). We also tested one mechanistic Heat Flux Model (HFM) that was based on energy gain and loss. Predictor variables used in model development included climate data (e.g., solar radiation, wind speed, etc.) collected from a nearby weather station and temperature and hydrologic data from upstream U.S. Geological Survey gages. Models were developed with a training dataset that consisted of data from 2008 to 2011; they were then independently validated with a test dataset from 2012. Model accuracy was evaluated using root mean square error (RMSE), Nash Sutcliffe efficiency (NSE), percent bias (PBIAS), and index of agreement (d) statistics. Model forecast success was evaluated using baseline-modified prime index of agreement (md) at the one, three, and five day predictions. All five models accurately predicted daily mean river temperature across the entire training dataset (RMSE = 0.58–1.311, NSE = 0.99–0.97, d = 0.98–0.99); ARIMA was most accurate (RMSE = 0.57, NSE = 0.99), but each model, other than ARIMA, showed short periods of under- or over-predicting observed warmer temperatures. For the training dataset, all models besides ARIMA had overestimation bias (PBIAS = −0.10 to −1.30). Validation analyses showed all models performed well; the HFM model was the most accurate compared other models (RMSE = 0.92, both NSE = 0.98, d = 0.99) and the ARIMA model was least accurate (RMSE = 2.06, NSE = 0.92, d = 0.98); however, all models had an overestimation bias (PBIAS = −4.1 to −10.20). Aside from the one day forecast ARIMA model (md = 0.53), all models forecasted fairly well at the one, three, and five day forecasts (md = 0.77–0.96). Overall, we were successful in developing models predicting daily mean temperature across a broad range of temperatures. These models, specifically the GLScos, ANN, and HFM, may serve as important tools for predicting conditions and managing thermal releases in regulated river systems such as the Delaware River. Further model development may be important in customizing predictions for particular biological or ecological needs, or for particular temporal or spatial scales.

Delaware, New York, Pennsylvania