Search USGSSearch

USGS · 70219219

Evaluation of six methods for correcting bias in estimates from ensemble tree machine learning regression model

Abstract

Ensemble-tree machine learning (ML) regression models can be prone to systematic bias: small values are overestimated and large values are underestimated. Additional bias can be introduced if the dependent variable is a transform of the original data. Six methods were evaluated for their ability to correct systematic and introduced bias. Method performance was evaluated using four case studies of groundwater quality: the units of the dependent variable were pH in two and log-concentration in the others. When performance metrics (bias and RMSE for both points and the CDF) were computed using the same units as those in the ML model, empirical distribution matching (EDM) provided the best results. When the metrics were computed using retransformed concentration, EDM and a method incorporating Duan's smearing estimate were both effective. A method based on the Z-score transform approximates EDM if the correlation coefficient between rank-ordered ML estimates and rank-ordered observations approaches one.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kenneth Belitz, Paul E. Stackelberg. 2021. Evaluation of six methods for correcting bias in estimates from ensemble tree machine learning regression model. https://doi.org/10.1016/j.envsoft.2021.105006

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Multiple machine-learning estimation of groundwater levels and trends for the regional Mississippi River Valley alluvial aquifer

The Mississippi River Valley alluvial aquifer provides irrigation, public, and domestic water supplies across the south-central United States. Declining groundwater levels require improved characterization of changing conditions. Traditional potentiometric-surface mapping does not use all available water-level data or quantify uncertainty. To address these limitations, we developed a data-driven multiple machine-learning (MML) framework delivered through two open-source R packages. The covMRVAgen1 software assembles covariates to 155,960 monthly groundwater levels from 57,695 wells; the mmlMRVAgen1 software trains Cubist and Random Forest models, blends them, and makes 1-kilometer gridded predictions of monthly potentiometric surfaces for the period January 1980–December 2022. The MML approach provides a methodological foundation for region-scale spatiotemporal groundwater prediction and uncertainty quantification, generating 90-percent prediction limits with appropriate empirical coverage. Model performance is acceptable, with a root-mean-square error of about 4.2 feet, standard deviation of 24.82 feet, and a normalized Nash–Sutcliffe efficiency of 0.973.

Arkansas, Illinois, Louisiana, Mississippi, Missou

Leveraging machine learning to automate regression model evaluations for large multi-site water-quality trend studies

Large multi-site trend studies provide an opportunity to evaluate progress of waterbodies towards water-quality goals across broad geographic areas. Such studies often aggregate the results of site-specific models and thus contend with evaluating each model for appropriate fit and statistical assumptions. We explored the use of four traditional machine learning models (logistic regression, linear and quadratic discriminant analysis, and k-nearest neighbors) to perform these checks and estimate probabilities that an analyst would publish or reject a site-specific trend model from a multi-site study. We trained these “model-checking models” (MCMs) using a national study of over 6000 trend models and tested the MCMs using a smaller set of novel trend models. Although the MCMs did not perform well enough to bypass analyst review entirely, we found incorporating an MCM into a larger evaluation workflow can reduce the number of trend models needing an analyst review by more than half.

Environmental Modeling and Software

A machine learning approach to predicting equilibrium ripple wavelength

Sand ripples are geomorphic features on the seafloor that affect bottom boundary layer dynamics including wave attenuation and sediment transport. We present a new equilibrium ripple predictor using a machine learning approach that outputs a probability distribution of wave-generated equilibrium wavelengths and statistics including an estimate of ripple height, the most probable ripple wavelength, and sediment and flow parameterizations. The Bayesian Optimal Model System (BOMS) is an ensemble machine learning system that combines two machine learning algorithms and two deterministic empirical ripple predictors with a Bayesian meta-learner to produce probabilistic wave-generated equilibrium ripple wavelength estimates in sandy locations. A ten-fold cross validation of BOMS resulted in an adjusted R-squared value of 0.93 and an average root mean square error (RMSE) of 8.0 cm. During both cross validation and testing on three unique field datasets, BOMS provided more accurate wavelength predictions than each individual base model and other common ripple predictors.

Environmental Modeling and Software