Search USGSSearch

Geology topics

Bernard T. Nolan

Publications and source records attributed to Bernard T. Nolan.

At least 19 recordsLinked to original sources

Machine learning predictions of mean ages of shallow well samples in the Great Lakes Basin, USA

The travel time or “age” of groundwater affects catchment responses to hydrologic changes , geochemical reactions, and time lags between management actions and responses at down-gradient streams and wells. Use of atmospheric tracers has facilitated the characterization of groundwater ages, but most wells lack such measurements. This study applied machine learning to predict ages in wells across a large region around the Great Lakes Basin using well, chemistry, and landscape characteristics. For a dataset of age tracers in 961 samples, the travel time from the land surface to the sample location was estimated for each sample using parametric functions. The mean travel times were then modeled using a gradient boosting machine (GBM) algorithm with cross validation tuning of model metaparameters. The GBM approach was able to closely match estimated ages for the training data (RMSE = 0.26 natural-log scale years) and provided a reasonable match to testing data (RMSE = 0.84). Of the variables tested, well characteristics (e.g. depth), land use, hydrologic indicators (e.g. topographic wetness index), and water chemistry (e.g. nitrate, fluoride, and pH), substantially affected the predictions of age. GBM prediction was applied to 14,335 groundwater samples with median sample depth of 5.4 m, indicating for the Great Lakes Basin a broad distribution of ages among wells with a median of 32.9 years. Lag times of decades are likely for these wells to respond to changing solute fluxes near land surface. While depth variables most strongly affected predicted mean ages, chemical constituents exhibited smooth trends with age, consistent with prevailing conceptual models of evolving sources and geochemistry flowpaths. The results provide proof of concept for use of readily available variables of well, landscape, and chemical characteristics to improve groundwater age estimates across large regions.

Great Lakes basin

Machine learning predictions of nitrate in groundwater used for drinking supply in the conterminous United States

Groundwater is an important source of drinking water supplies in the conterminous United State (CONUS), and presence of high nitrate concentrations may limit usability of groundwater in some areas because of the potential negative health effects. Prediction of locations of high nitrate groundwater is needed to focus mitigation and relief efforts. A three-dimensional extreme gradient boosting (XGB) machine learning model was developed to predict the distribution of nitrate. Nitrate was predicted at a 1 km resolution for two drinking water zones, each of variable depth, one for domestic supply and one for public supply. The model used measured nitrate concentrations from 12,082 wells and included predictor variables representing well characteristics, hydrologic conditions, soil type, geology, land use, climate, and nitrogen inputs. Predictor variables derived from empirical or numerical process-based models were also included to integrate information on controlling processes and conditions. The model provided accurate estimates at national and regional scales: the training (R 2 of 0.83) and hold-out (R 2 of 0.49) data fits compared favorably to previous studies. Predicted nitrate concentrations were less than 1 mg/L across most of the CONUS. Nationally, well depth, soil and climate characteristics, and the absence of developed land use were among the most influential explanatory factors. Only 1% of the area in either water supply zone had predicted nitrate concentrations greater than 10 mg/L; however, about 1.4 M people depend on groundwater for their drinking supplies in those areas. Predicted high concentrations of nitrate were most prevalent in the central CONUS. In areas of predicted high nitrate concentration, applied manure, farm fertilizer , and agricultural land use were influential predictor variables. This work represents the first application of XGB to a three-dimensional national-scale groundwater quality model and provides a significant milestone in the efforts to document nitrate in groundwater across the CONUS.

Science of the Total Environment

Modeling groundwater nitrate exposure in private wells of North Carolina for the Agricultural Health Study

Unregulated private wells in the United States are susceptible to many groundwater contaminants. Ingestion of nitrate, the most common anthropogenic private well contaminant in the United States, can lead to the endogenous formation of N-nitroso-compounds, which are known human carcinogens. In this study, we expand upon previous efforts to model private well groundwater nitrate concentration in North Carolina by developing multiple machine learning models and testing against out-of-sample prediction. Our purpose was to develop exposure estimates in unmonitored areas for use in the Agricultural Health Study (AHS) cohort. Using approximately 22,000 private well nitrate measurements in North Carolina, we trained and tested continuous models including a censored maximum likelihood-based linear model, random forest, gradient boosted machine, support vector machine, neural networks, and kriging. Continuous nitrate models had low predictive performance (R 2 < 0.33), so multiple random forest classification models were also trained and tested. The final classification approach predicted <1 mg/L, 1–5 mg/L, and ≥5 mg/L using a random forest model with 58 variables and maximizing the Cohen's kappa statistic. The final model had an overall accuracy of 0.75 and high specificity for the higher two categories and high sensitivity for the lowest category. The results will be used for the categorical prediction of private well nitrate for AHS cohort participants that reside in North Carolina.

North Carolina

Maps showing predicted probabilities for selected dissolved oxygen and dissolved manganese threshold events in depth zones used by the domestic and public drinking water supply wells, Central Valley, California

The purpose of the prediction grids for selected redox constituents—dissolved oxygen and dissolved manganese—are intended to provide an understanding of groundwater-quality conditions at the domestic and public-supply drinking water depths. The chemical quality of groundwater and the fate of many contaminants is influenced by redox processes in all aquifers, and understanding the redox conditions horizontally and vertically is critical in evaluating groundwater quality. The redox condition of groundwater—whether oxic (oxygen present) or anoxic (oxygen absent)—strongly influences the oxidation state of a chemical in groundwater. The anoxic dissolved oxygen thresholds of <0.5 milligram per liter (mg/L), <1.0 mg/L, and <2.0 mg/L were selected to apply broadly to regional groundwater-quality investigations. Although the presence of dissolved manganese in groundwater indicates strongly reducing (anoxic) groundwater conditions, it is also considered a “nuisance” constituent in drinking water, making drinking water undesirable with respect to taste, staining, or scaling. Three dissolved manganese thresholds, <50 micrograms per liter (µg/L), <150 µg/L, and <300 µg/L, were selected to create predicted probabilities of exceedances in depth zones used by domestic and public-supply water wells. The 50 µg/L event threshold represents the secondary maximum contaminant level (SMCL) benchmark for manganese (U.S. Environmental Protection Agency, 2017; California Division of Drinking Water, 2014), whereas the 300 µg/L event threshold represents the U.S. Geological Survey (USGS) health-based screening level (HBSL) benchmark, used to put measured concentrations of drinking-water contaminants into a human-health context (Toccalino and others, 2014). The 150 µg/L event threshold represents one-half the USGS HBSL. The resultant dissolved oxygen and dissolved manganese prediction grids may be of interest to water-resource managers, water-quality researchers, and groundwater modelers concerned with the occurrence of natural and anthropogenic contaminants related to anoxic conditions. Prediction grids for selected redox constituents and thresholds were created by the USGS National Water-Quality Assessment (NAWQA) modeling and mapping team.

California

Regional variability of nitrate fluxes in the unsaturated zone and groundwater, Wisconsin, USA

Process-based modeling of regional NO3− fluxes to groundwater is critical for understanding and managing water quality, but the complexity of NO3− reactive transport processes make implementation a challenge. This study introduces a regional vertical flux method (VFM) for efficient estimation of reactive transport of NO3− in the vadose zone and groundwater. The regional VFM was applied to 443 well samples in central-eastern Wisconsin. Chemical measurements included O2, NO3−, N2 from denitrification, and atmospheric tracers of groundwater age including carbon-14, chlorofluorocarbons, tritium, and tritiogenic helium. VFM results were consistent with observed chemistry, and calibrated parameters were in-line with estimates from previous studies. Results indicated that (1) unsaturated zone travel times were a substantial portion of the transit time to wells and streams (2) since 1945 fractions of applied N leached to groundwater have increased for manure-N, possibly due to increased injection of liquid manure, and decreased for fertilizer-N, and (3) under current practices and conditions, approximately 60% of the shallow aquifer will eventually be affected by downward migration of NO3−, with denitrification protecting the remaining 40%. Recharge variability strongly affected the unsaturated zone lag times and the eventual depth of the NO3− front. Principal components regression demonstrated that VFM parameters and predictions were significantly correlated with hydrogeochemical landscape features. The diverse and sometimes conflicting aspects of N management (e.g. limiting N volatilization versus limiting N losses to groundwater) warrant continued development of large-scale holistic strategies to manage water quality and quantity.

Wisconsin

Metamodeling and mapping of nitrate flux in the unsaturated zone and groundwater, Wisconsin, USA

Nitrate contamination of groundwater in agricultural areas poses a major challenge to the sustainability of water resources. Aquifer vulnerability models are useful tools that can help resource managers identify areas of concern, but quantifying nitrogen (N) inputs in such models is challenging, especially at large spatial scales. We sought to improve regional nitrate (NO 3 − ) input functions by characterizing unsaturated zone NO 3 − transport to groundwater through use of surrogate, machine-learning metamodels of a process-based N flux model. The metamodels used boosted regression trees (BRTs) to relate mappable landscape variables to parameters and outputs of a previous “vertical flux method” (VFM) applied at sampled wells in the Fox, Wolf, and Peshtigo (FWP) river basins in northeastern Wisconsin. In this context, the metamodels upscaled the VFM results throughout the region, and the VFM parameters and outputs are the metamodel response variables. The study area encompassed the domain of a detailed numerical model that provided additional predictor variables, including groundwater recharge, to the metamodels. We used a statistical learning framework to test a range of model complexities to identify suitable hyperparameters of the six BRT metamodels corresponding to each response variable of interest: NO 3 − source concentration factor (which determines the local NO 3 − input concentration); unsaturated zone travel time; NO 3 − concentration at the water table in 1980, 2000, and 2020 (three separate metamodels); and NO 3 − “extinction depth”, the eventual steady state depth of the NO 3 − front. The final metamodels were trained to 129 wells within the active numerical flow model area, and considered 58 mappable predictor variables compiled in a geographic information system (GIS). These metamodels had training and cross-validation testing R 2 values of 0.52 – 0.86 and 0.22 – 0.38, respectively, and predictions were compiled as maps of the above response variables. Testing performance was reasonable, considering that we limited the metamodel predictor variables to mappable factors as opposed to using all available VFM input variables. Relationships between metamodel predictor variables and mapped outputs were generally consistent with expectations, e.g. with greater source concentrations and NO 3 − at the groundwater table in areas of intensive crop use and well drained soils. Shorter unsaturated zone travel times in poorly drained areas likely indicated preferential flow through clay soils, and a tendency for fine grained deposits to collocate with areas of shallower water table. Numerical estimates of groundwater recharge were important in the metamodels and may have been a proxy for N input and redox conditions in the northern FWP, which had shallow predicted NO 3 − extinction depth. The metamodel results provide proof-of-concept for regional characterization of unsaturated zone NO 3 − transport processes in a statistical framework based on readily mappable GIS input variables.

Wisconsin

Estimating the high-arsenic domestic-well population in the conterminous United States

Arsenic concentrations from 20 450 domestic wells in the U.S. were used to develop a logistic regression model of the probability of having arsenic >10 μg/L (“high arsenic”), which is presented at the county, state, and national scales. Variables representing geologic sources, geochemical, hydrologic, and physical features were among the significant predictors of high arsenic. For U.S. Census blocks, the mean probability of arsenic >10 μg/L was multiplied by the population using domestic wells to estimate the potential high-arsenic domestic-well population. Approximately 44.1 M people in the U.S. use water from domestic wells. The population in the conterminous U.S. using water from domestic wells with predicted arsenic concentration >10 μg/L is 2.1 M people (95% CI is 1.5 to 2.9 M). Although areas of the U.S. were underrepresented with arsenic data, predictive variables available in national data sets were used to estimate high arsenic in unsampled areas. Additionally, by predicting to all of the conterminous U.S., we identify areas of high and low potential exposure in areas of limited arsenic data. These areas may be viewed as potential areas to investigate further or to compare to more detailed local information. Linking predictive modeling to private well use information nationally, despite the uncertainty, is beneficial for broad screening of the population at risk from elevated arsenic in drinking water from private wells.

Environmental Science & Technology

Predicted pH at the domestic and public supply drinking water depths, Central Valley, California

This scientific investigations map is a product of the U.S. Geological Survey (USGS) National Water-Quality Assessment (NAWQA) project modeling and mapping team. The prediction grids depicted in this map are of continuous pH and are intended to provide an understanding of groundwater-quality conditions at the domestic and public supply drinking water zones in the groundwater of the Central Valley of California. The chemical quality of groundwater and the fate of many contaminants is often influenced by pH in all aquifers. These grids are of interest to water-resource managers, water-quality researchers, and groundwater modelers concerned with the occurrence of natural and anthropogenic contaminants related to pH. In this work, the median well depth categorized as domestic supply was 30 meters below land surface, and the median well depth categorized as public supply is 100 meters below land surface. Prediction grids were created using prediction modeling methods, specifically boosted regression trees (BRT) with a Gaussian error distribution within a statistical learning framework within the computing framework of R ( http://www.r-project.org/ ). The statistical learning framework seeks to maximize the predictive performance of machine learning methods through model tuning by cross validation. The response variable was measured pH from 1,337 wells and was compiled from two sources: USGS National Water Information System (NWIS) database (all data are publicly available from the USGS: http://waterdata.usgs.gov/ca/nwis/nwis ) and the California State Water Resources Control Board Division of Drinking Water (SWRCB-DDW) database (water quality data are publicly available from the SWRCB: http://www.waterboards.ca.gov/gama/geotracker_gama.shtml ). Only wells with measured pH and well depth data were selected, and for wells with multiple records, only the most recent sample in the period 1993–2014 was used. A total of 1,003 wells (training dataset) were used to train the BRT model, and 334 wells (hold-out dataset) were used to validate the prediction model. The training r-squared was 0.70, and the root-mean-square error (RMSE) in standard pH units was 0.26. The hold-out r-squared was 0.43, and RMSE in standard pH units was 0.37. Predictor variables consisting of more than 60 variables from 7 sources were assembled to develop a model that incorporates regional-scale soil properties, soil chemistry, land use, aquifer textures, and aquifer hydrology. Previously developed Central Valley model outputs of textures (Central Valley Textural Model, CVTM; Faunt and others, 2010) and MODFLOW-simulated vertical water fluxes and predicted depth to water table (Central Valley Hydrologic Model, CVHM; Faunt, 2009) were used to represent aquifer textures and groundwater hydraulics, respectively. In this work, wells were attributed to predictor variable values in ArcGIS using a 500-meter buffer. Faunt, C.C., ed., 2009, Groundwater availability in the Central Valley aquifer, California: U.S. Geological Survey Professional Paper 1776, 225 p., accessed at https://pubs.usgs.gov/pp/1766/ . Faunt, C.C., Belitz, K., and Hanson, R.T., 2010, Development of a three-dimensional model of sedimentary texture in valley-fill deposits of Central Valley, California, USA: Hydrogeology Journal, v. 18, no. 3, p. 625–649, https://doi.org/10.1007/s10040-009-0539-7 .

California

Prediction and visualization of redox conditions in the groundwater of Central Valley, California

Regional-scale, three-dimensional continuous probability models, were constructed for aspects of redox conditions in the groundwater system of the Central Valley, California. These models yield grids depicting the probability that groundwater in a particular location will have dissolved oxygen (DO) concentrations less than selected threshold values representing anoxic groundwater conditions, or will have dissolved manganese (Mn) concentrations greater than selected threshold values representing secondary drinking water-quality contaminant levels (SMCL) and health-based screening levels (HBSL). The probability models were constrained by the alluvial boundary of the Central Valley to a depth of approximately 300 m. Probability distribution grids can be extracted from the 3-D models at any desired depth, and are of interest to water-resource managers, water-quality researchers, and groundwater modelers concerned with the occurrence of natural and anthropogenic contaminants related to anoxic conditions. Models were constructed using a Boosted Regression Trees (BRT) machine learning technique that produces many trees as part of an additive model and has the ability to handle many variables, automatically incorporate interactions, and is resistant to collinearity. Machine learning methods for statistical prediction are becoming increasing popular in that they do not require assumptions associated with traditional hypothesis testing. Models were constructed using measured dissolved oxygen and manganese concentrations sampled from 2767 wells within the alluvial boundary of the Central Valley, and over 60 explanatory variables representing regional-scale soil properties, soil chemistry, land use, aquifer textures, and aquifer hydrologic properties. Models were trained on a USGS dataset of 932 wells, and evaluated on an independent hold-out dataset of 1835 wells from the California Division of Drinking Water. We used cross-validation to assess the predictive performance of models of varying complexity, as a basis for selecting final models. Trained models were applied to cross-validation testing data and a separate hold-out dataset to evaluate model predictive performance by emphasizing three model metrics of fit: Kappa; accuracy; and the area under the receiver operator characteristic curve (ROC). The final trained models were used for mapping predictions at discrete depths to a depth of 304.8 m. Trained DO and Mn models had accuracies of 86–100%, Kappa values of 0.69–0.99, and ROC values of 0.92–1.0. Model accuracies for cross-validation testing datasets were 82–95% and ROC values were 0.87–0.91, indicating good predictive performance. Kappas for the cross-validation testing dataset were 0.30–0.69, indicating fair to substantial agreement between testing observations and model predictions. Hold-out data were available for the manganese model only and indicated accuracies of 89–97%, ROC values of 0.73–0.75, and Kappa values of 0.06–0.30. The predictive performance of both the DO and Mn models was reasonable, considering all three of these fit metrics and the low percentages of low-DO and high-Mn events in the data.

California

A hybrid machine learning model to predict and visualize nitrate concentration throughout the Central Valley aquifer, California, USA

Intense demand for water in the Central Valley of California and related increases in groundwater nitrate concentration threaten the sustainability of the groundwater resource. To assess contamination risk in the region, we developed a hybrid, non-linear, machine learning model within a statistical learning framework to predict nitrate contamination of groundwater to depths of approximately 500 m below ground surface. A database of 145 predictor variables representing well characteristics, historical and current field and landscape-scale nitrogen mass balances, historical and current land use, oxidation/reduction conditions, groundwater flow, climate, soil characteristics, depth to groundwater, and groundwater age were assigned to over 6000 private supply and public supply wells measured previously for nitrate and located throughout the study area. The boosted regression tree (BRT) method was used to screen and rank variables to predict nitrate concentration at the depths of domestic and public well supplies. The novel approach included as predictor variables outputs from existing physically based models of the Central Valley. The top five most important predictor variables included two oxidation/reduction variables (probability of manganese concentration to exceed 50 ppb and probability of dissolved oxygen concentration to be below 0.5 ppm), field-scale adjusted unsaturated zone nitrogen input for the 1975 time period, average difference between precipitation and evapotranspiration during the years 1971–2000, and 1992 total landscape nitrogen input. Twenty-five variables were selected for the final model for log-transformed nitrate. In general, increasing probability of anoxic conditions and increasing precipitation relative to potential evapotranspiration had a corresponding decrease in nitrate concentration predictions. Conversely, increasing 1975 unsaturated zone nitrogen leaching flux and 1992 total landscape nitrogen input had an increasing relative impact on nitrate predictions. Three-dimensional visualization indicates that nitrate predictions depend on the probability of anoxic conditions and other factors, and that nitrate predictions generally decreased with increasing groundwater age.

California

Predicting arsenic in drinking water wells of the Central Valley, California

Probabilities of arsenic in groundwater at depths used for domestic and public supply in the Central Valley of California are predicted using weak-learner ensemble models (boosted regression trees, BRT) and more traditional linear models (logistic regression, LR). Both methods captured major processes that affect arsenic concentrations, such as the chemical evolution of groundwater, redox differences, and the influence of aquifer geochemistry. Inferred flow-path length was the most important variable but near-surface-aquifer geochemical data also were significant. A unique feature of this study was that previously predicted nitrate concentrations in three dimensions were themselves predictive of arsenic and indicated an important redox effect at >10 μg/L, indicating low arsenic where nitrate was high. Additionally, a variable representing three-dimensional aquifer texture from the Central Valley Hydrologic Model was an important predictor, indicating high arsenic associated with fine-grained aquifer sediment. BRT outperformed LR at the 5 μg/L threshold in all five predictive performance measures and at 10 μg/L in four out of five measures. BRT yielded higher prediction sensitivity (39%) than LR (18%) at the 10 μg/L threshold–a useful outcome because a major objective of the modeling was to improve our ability to predict high arsenic areas.

California

Evaluating the sources of water to wells: Three techniques for metamodeling of a groundwater flow model

For decision support, the insights and predictive power of numerical process models can be hampered by insufficient expertise and computational resources required to evaluate system response to new stresses. An alternative is to emulate the process model with a statistical &ldquo;metamodel.&rdquo; Built on a dataset of collocated numerical model input and output, a groundwater flow model was emulated using a Bayesian Network, an Artificial neural network, and a Gradient Boosted Regression Tree. The response of interest was surface water depletion expressed as the source of water-to-wells. The results have application for managing allocation of groundwater. Each technique was tuned using cross validation and further evaluated using a held-out dataset. A numerical MODFLOW-USG model of the Lake Michigan Basin, USA, was used for the evaluation. The performance and interpretability of each technique was compared pointing to advantages of each technique. The metamodel can extend to unmodeled areas.

Illinois, Indiana, Michigan, Ohio, Wisconsin

Assessing the relationship between groundwater nitrate and animal feeding operations in Iowa (USA)

Nitrate-nitrogen is a common contaminant of drinking water in many agricultural areas of the United States of America (USA). Ingested nitrate from contaminated drinking water has been linked to an increased risk of several cancers, specific birth defects, and other diseases. In this research, we assessed the relationship between animal feeding operations (AFOs) and groundwater nitrate in private wells in Iowa. We characterized AFOs by swine and total animal units and type (open, confined, or mixed), and we evaluated the number and spatial intensities of AFOs in proximity to private wells. The types of AFO indicate the extent to which a facility is enclosed by a roof. Using linear regression models, we found significant positive associations between the total number of AFOs within 2 km of a well (p trend < 0.001), number of open AFOs within 5 km of a well (p trend < 0.001), and number of mixed AFOs within 30 km of a well (p trend < 0.001) and the log nitrate concentration. Additionally, we found significant increases in log nitrate in the top quartiles for AFO spatial intensity, open AFO spatial intensity, and mixed AFO spatial intensity compared to the bottom quartile (0.171 log(mg/L), 0.319 log(mg/L), and 0.541 log(mg/L), respectively; all p < 0.001). We also explored the spatial distribution of nitrate-nitrogen in drinking wells and found significant spatial clustering of high-nitrate wells (> 5 mg/L) compared with low-nitrate (&le; 5 mg/L) wells ( p = 0.001). A generalized additive model for high-nitrate status identified statistically significant areas of risk for high levels of nitrate. Adjustment for some AFO predictor variables explained a portion of the elevated nitrate risk. These results support a relationship between animal feeding operations and groundwater nitrate concentrations and differences in nitrate loss from confined AFOs vs. open or mixed types.

Iowa

Corn stover harvest increases herbicide movement to subsurface drains: RZWQM simulations

BACKGROUND Crop residue removal for bioenergy production can alter soil hydrologic properties and the movement of agrochemicals to subsurface drains. The Root Zone Water Quality Model (RZWQM), previously calibrated using measured flow and atrazine concentrations in drainage from a 0.4 ha chisel-tilled plot, was used to investigate effects of 50 and 100% corn ( Zea mays L.) stover harvest and the accompanying reductions in soil crust hydraulic conductivity and total macroporosity on transport of atrazine, metolachlor, and metolachlor oxanilic acid (OXA). RESULTS The model accurately simulated field-measured metolachlor transport in drainage. A 3-yr simulation indicated that 50% residue removal decreased subsurface drainage by 31% and increased atrazine and metolachlor transport in drainage 4 to 5-fold when surface crust conductivity and macroporosity were reduced by 25%. Based on its measured sorption coefficient, ~ 2-fold reductions in OXA losses were simulated with residue removal. CONCLUSION RZWQM indicated that if corn stover harvest reduces crust conductivity and soil macroporosity, losses of atrazine and metolachlor in subsurface drainage will increase due to reduced sorption related to more water moving through fewer macropores. Losses of the metolachlor degradation product OXA will decrease due to the more rapid movement of the parent compound into the soil.

Pest Management Science

A statistical learning framework for groundwater nitrate models of the Central Valley, California, USA

We used a statistical learning framework to evaluate the ability of three machine-learning methods to predict nitrate concentration in shallow groundwater of the Central Valley, California: boosted regression trees (BRT), artificial neural networks (ANN), and Bayesian networks (BN). Machine learning methods can learn complex patterns in the data but because of overfitting may not generalize well to new data. The statistical learning framework involves cross-validation (CV) training and testing data and a separate hold-out data set for model evaluation, with the goal of optimizing predictive performance by controlling for model overfit. The order of prediction performance according to both CV testing R 2 and that for the hold-out data set was BRT > BN > ANN. For each method we identified two models based on CV testing results: that with maximum testing R 2 and a version with R 2 within one standard error of the maximum (the 1SE model). The former yielded CV training R 2 values of 0.94&ndash;1.0. Cross-validation testing R 2 values indicate predictive performance, and these were 0.22&ndash;0.39 for the maximum R 2 models and 0.19&ndash;0.36 for the 1SE models. Evaluation with hold-out data suggested that the 1SE BRT and ANN models predicted better for an independent data set compared with the maximum R 2 versions, which is relevant to extrapolation by mapping. Scatterplots of predicted vs. observed hold-out data obtained for final models helped identify prediction bias, which was fairly pronounced for ANN and BN. Lastly, the models were compared with multiple linear regression (MLR) and a previous random forest regression (RFR) model. Whereas BRT results were comparable to RFR, MLR had low hold-out R 2 (0.07) and explained less than half the variation in the training data. Spatial patterns of predictions by the final, 1SE BRT model agreed reasonably well with previously observed patterns of nitrate occurrence in groundwater of the Central Valley.

California

Modeling groundwater nitrate concentrations in private wells in Iowa

Contamination of drinking water by nitrate is a growing problem in many agricultural areas of the country. Ingested nitrate can lead to the endogenous formation of N-nitroso compounds, potent carcinogens. We developed a predictive model for nitrate concentrations in private wells in Iowa. Using 34,084 measurements of nitrate in private wells, we trained and tested random forest models to predict log nitrate levels by systematically assessing the predictive performance of 179 variables in 36 thematic groups (well depth, distance to sinkholes, location, land use, soil characteristics, nitrogen inputs, meteorology, and other factors). The final model contained 66 variables in 17 groups. Some of the most important variables were well depth, slope length within 1 km of the well, year of sample, and distance to nearest animal feeding operation. The correlation between observed and estimated nitrate concentrations was excellent in the training set (r-square = 0.77) and was acceptable in the testing set (r-square = 0.38). The random forest model had substantially better predictive performance than a traditional linear regression model or a regression tree. Our model will be used to investigate the association between nitrate levels in drinking water and cancer risk in the Iowa participants of the Agricultural Health Study cohort.

Iowa

Simulating maize yield and bomass with spatial variability of soil field capacity

Spatial variability in field soil properties is a challenge for system modelers who use single representative values, such as means, for model inputs, rather than their distributions. In this study, the root zone water quality model (RZWQM2) was first calibrated for 4 yr of maize ( Zea mays L.) data at six irrigation levels in northern Colorado and then used to study spatial variability of soil field capacity (FC) estimated in 96 plots on maize yield and biomass. The best results were obtained when the crop parameters were fitted along with FCs, with a root mean squared error (RMSE) of 354 kg ha &ndash;1 for yield and 1202 kg ha &ndash;1 for biomass. When running the model using each of the 96 sets of field-estimated FC values, instead of calibrating FCs, the average simulated yield and biomass from the 96 runs were close to measured values with a RMSE of 376 kg ha &ndash;1 for yield and 1504 kg ha &ndash;1 for biomass. When an average of the 96 FC values for each soil layer was used, simulated yield and biomass were also acceptable with a RMSE of 438 kg ha &ndash;1 for yield and 1627 kg ha &ndash;1 for biomass. Therefore, when there are large numbers of FC measurements, an average value might be sufficient for model inputs. However, when the ranges of FC measurements were known for each soil layer, a sampled distribution of FCs using the Latin hypercube sampling (LHS) might be used for model inputs.

Agronomy Journal