Search USGSSearch

SEARCH · Search USGS

Results for “Computational Statistics and Data Analysis”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Updating data inputs, assessing trends, and evaluating a method to estimate probable high groundwater levels in selected areas of Massachusetts

A method to estimate the probable high groundwater level in Massachusetts, excluding Cape Cod and the islands, was developed in 1981. The method uses a groundwater measurement from a test site, groundwater measurements from an index well, and a distribution of high groundwater levels from wells in similar geologic and topographic settings. The U.S. Geological Survey, in cooperation with the Massachusetts Department of Environmental Protection, conducted an update to the Frimpter method for estimating the probable high groundwater levels in Massachusetts. The study evaluated the potential changes to the method resulting from four decades of additional groundwater-level data and the expansion of the network of wells for monitoring groundwater levels. The differences and potential benefits of daily, as opposed to monthly, measurements in the application of the method were examined because of the increased availability of high-frequency (subdaily) groundwater-level data. The study also considered long-term trends in groundwater levels that may alter the accuracy of the method. Finally, the accuracy of the estimated high groundwater levels was evaluated, and improved implementation guidance was prepared. For this study, groundwater levels in 153 wells in Massachusetts and surrounding States with records with lengths of 16 to 78 years were analyzed. The highest recorded groundwater levels ranged from 1.2 feet (ft) above land surface (flooded conditions) to 45.8 ft below land surface, with a median of 4.6 ft below land surface. The maximum annual groundwater-level range was 1.4 to 17.9 ft, with a median of 5.5 ft. The within-month variation, maximum annual groundwater-level range, and highest recorded groundwater level were computed using daily mean groundwater-level values from 28 wells with continuous records. The use of daily data resulted in larger maximum annual groundwater-level ranges (0.02 to 2.94 ft larger, with a median of 0.58 ft larger) and shallower highest-recorded groundwater levels (0.0 to 1.60 ft shallower, with a median of 0.18 ft shallower) than computations based on monthly measurements in the same wells. Statistical tests showed moderate to strong evidence of trends in measurements of both high and low groundwater levels within most of the periods during which water levels were analyzed. High groundwater levels rose beneath the land surface at most sites during four of the six periods used for analysis (1966–2015, 1986–2015, 1991–2010, and 1981–2010). Low groundwater levels also increased at many sites during most of the periods evaluated, but this trend was less widespread than the similar trends in high groundwater levels, and the trend was to deeper low groundwater levels at more sites than the trend to deeper high groundwater levels. There was no clear trend in annual groundwater-level ranges at most sites during the six periods analyzed. In general, the Frimpter method predicted shallower (higher) high groundwater levels than were observed but correctly classified sites according to their suitabilities for unmounded septic systems. The mean error of the predictions (difference between the estimated and observed groundwater levels) ranged from −3.23 ft to −1.40 ft for various approaches to estimating the groundwater-level range and selecting an index well. The method correctly classified 83 to 86 percent of monitoring-well sites according to their suitability for an unmounded septic system for many approaches to estimating the annual groundwater-level range and selecting an index well. The approach selected for estimating the annual groundwater-level range and selecting an index well will depend upon the importance of an accurate estimate of the high groundwater level as compared to the importance of an estimated high groundwater level that is less likely to be exceeded.

Connecticut, Massachusetts, New Hampshire, Rhode I

Streamflow model of Wisconsin River for estimating flood frequency and volume

A set of daily streamflow-routing models are used to simulate streamflow at 10 sites along the Wisconsin River for water years 1915-76, to determine the effects the reservoir system has on flood discharges. Streamflow is simulated under the following two conditions: (1) No reservoirs are in the system and (2) all of the present reservoirs are in place and operated with current rules. At Wisconsin Dells, 20 miles upstream from Portage, daily streamflow hydrographs are estimated for the 10-, 50-, 100-, and 500-year floods. These were determined from statistical analysis of the simulated daily streamflows for the condition of all reservoirs in place. The reservoirs have a significant impact on floods. The mean annual flood peak at Wisconsin Dells is lowered about 20% from 43,000 cubic feet per second for the simulated, unregulated condition to 34,000 cubic feet per second for the simulated, regulated condition. The 100-year flood peak at Wisconsin Dells is reduced about 10% (92,000 to 82,000 cubic feet per second) between the simulated, unregulated and simulated, regulated conditions. The 100-year flood peak at Wisconsin Dells, computed from the simulated, regulated streamflow data for the period 1915-76, is 82,000 cubic feet per second, including the effects of all the reservoirs in the river system, as they are currently operated. It also includes the effects of Lakes Du Bay, Petenwell, and Castle Rock which are significant for spring floods but are insignificant for summer or fall floods because they are normally maintained nearly full in the summer and fall and have very little storage for floodwaters. (USGS)

Wisconsin

Evaluation of thermograph data for California streams

Statistical analysis of water-temperature data from California streams indicates that, for most purposes, long-term operation of thermographs (automatic water-temperature recording instruments) does not provide a more useful record than either short-term operation of such instruments or periodic measurements. Harmonic analyses were made of thermograph records 5 to 14 years in length from 82 stations. More than 80 percent of the annual variation in water temperature is explained by the harmonic function for 77 of the 82 stations. Harmonic coefficients based on 8 years of thermograph record at 12 stations varied only slightly from coefficients computed using two equally split 4-year records. At five stations where both thermograph and periodic (10 to 23 measurements per year) data were collected concurrently, harmonic coefficients for periodic data were defined nearly as well as those for thermograph data. Results of this analysis indicate that, except where detailed surveillance of water temperatures is required or where there is a chance of temporal change, thermograph operations can be reduced substantially without affecting the usefulness of temperature records.

California

Water levels in the Calumet aquifer and their relation to surface-water levels in northern Lake County, Indiana, 1985-92

The U.S. Geological Survey made 2,328 water-level measurements at a total of 96 ground-water and surface-water sites in northern Lake County, Indiana, from August 1985 through September 1992. This report lists and summarizes the significance of the measurements. Northern Lake County is on the southern shore of Lake Michigan and includes the cities of East Chicago, Gary, Hammond, and Whiting. The study area is underlain by the unconfined Calumet aquifer and receives about 36 inches of precipitation per year. The U.S. Geological Survey investigated ground-water levels and flow in the Calumet aquifer and the effect of Lake Michigan levels on ground-water and surface-water levels throughout the study area. Summary statistics of the water-level data were computed for each site. Ground-water levels annually reach a maximum in June or July and a minimum in September or October. Measured groundwater fluctuations in the Calumet aquifer during the study period ranged from 0.40 to 5.01 feet, and the mean ground-water fluctuation was about 2.3 feet The largest surface-water fluctuations were affected by record setting Lake Michigan levels. Midmonth daily averages for the data-collection period show that Lake Michigan fluctuated 4.14 feet Water-level fluctuations on the Grand Calumet River were from 1.06 to 2.45 feet. Analysis of water-level data indicates that the 1988 drought did not substantially affect water levels in the Calumet aquifer, but the deficit in precipitation reversed vertical flow gradients in ground water at three paired deep and shallow wells. High water levels in Lake Michigan during 1985-87 created long-term backwater effects on the Grand Calumet River as far as 11.0 miles upstream from Lake Michigan. Analysis of water-level data from the data-collection network indicates that the water table normally slopes toward streams, ditches, sewers, the Indiana Harbor Canal, and Lake Michigan. The slope of the water table toward the Grand Calumet River is greatest in the winter and can decrease to being almost horizontal in the summer. Wells near streams respond quickly to nearby surface-water-level changes. Water-table maps indicate that sewers and dewatering systems are lowering ground-water levels in large areas. Ditches, the Grand Calumet River, and the Indiana Harbor Canal connect the Lake Michigan water level to large parts of the study area. The surface-water stage in the Indiana Harbor Canal, which functions as a ditch, can equal Lake Michigan's stage up to 3.75 miles inland from the lakeshore. Human activity, the stage of Lake Michigan, and the storage capacity of the Calumet aquifer combine to reduce vertical changes in the water table in the study area.

Indiana

Identifying trends in sediment discharge from alterations in upstream land use

Environmental monitoring is a primary reason for collecting sediment data. One emphasis of this monitoring is identification of trends in suspended sediment discharge. A stochastic equation was used to generate time series of annual suspended sediment discharges using statistics from gaging stations with drainage areas between 1606 and 1 805 230 km2. Annual sediment discharge was increased linearly to yield a given increase at the end of a fixed period and trend statistics were computed for each simulation series using Kendal's tau (at 0.05 significance level). A parameter was calculated from two factors that control trend detection time: (a) the magnitude of change in sediment discharge, and (b) the natural variability of sediment discharge. In this analysis the detection of a trend at most stations is well over 100 years for a 20% increase in sediment discharge. Further research is needed to assess the sensitivity of detecting trends at sediment stations.

Effects of scale on interpretation and management

Methods for determining magnitude and frequency of floods in California, based on data through water year 2006

Methods for estimating the magnitude and frequency of floods in California that are not substantially affected by regulation or diversions have been updated. Annual peak-flow data through water year 2006 were analyzed for 771 streamflow-gaging stations (streamgages) in California having 10 or more years of data. Flood-frequency estimates were computed for the streamgages by using the expected moments algorithm to fit a Pearson Type III distribution to logarithms of annual peak flows for each streamgage. Low-outlier and historic information were incorporated into the flood-frequency analysis, and a generalized Grubbs-Beck test was used to detect multiple potentially influential low outliers. Special methods for fitting the distribution were developed for streamgages in the desert region in southeastern California. Additionally, basin characteristics for the streamgages were computed by using a geographical information system. Regional regression analysis, using generalized least squares regression, was used to develop a set of equations for estimating flows with 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities for ungaged basins in California that are outside of the southeastern desert region. Flood-frequency estimates and basin characteristics for 630 streamgages were combined to form the final database used in the regional regression analysis. Five hydrologic regions were developed for the area of California outside of the desert region. The final regional regression equations are functions of drainage area and mean annual precipitation for four of the five regions. In one region, the Sierra Nevada region, the final equations are functions of drainage area, mean basin elevation, and mean annual precipitation. Average standard errors of prediction for the regression equations in all five regions range from 42.7 to 161.9 percent. For the desert region of California, an analysis of 33 streamgages was used to develop regional estimates of all three parameters (mean, standard deviation, and skew) of the log-Pearson Type III distribution. The regional estimates were then used to develop a set of equations for estimating flows with 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities for ungaged basins. The final regional regression equations are functions of drainage area. Average standard errors of prediction for these regression equations range from 214.2 to 856.2 percent. Annual peak-flow data through water year 2006 were analyzed for eight streamgages in California having 10 or more years of data considered to be affected by urbanization. Flood-frequency estimates were computed for the urban streamgages by fitting a Pearson Type III distribution to logarithms of annual peak flows for each streamgage. Regression analysis could not be used to develop flood-frequency estimation equations for urban streams because of the limited number of sites. Flood-frequency estimates for the eight urban sites were graphically compared to flood-frequency estimates for 630 non-urban sites. The regression equations developed from this study will be incorporated into the U.S. Geological Survey (USGS) StreamStats program. The StreamStats program is a Web-based application that provides streamflow statistics and basin characteristics for USGS streamgages and ungaged sites of interest. StreamStats can also compute basin characteristics and provide estimates of streamflow statistics for ungaged sites when users select the location of a site along any stream in California.

California

Comparison of machine learning approaches used to identify the drivers of Bakken oil well productivity

Geologists and petroleum engineers have struggled to identify the mechanisms that drive productivity in horizontal hydraulically fractured oil wells. The machine learning algorithms of Random Forest (RF), gradient boosting trees (GBT) and extreme gradient boosting (XGBoost) were applied to a dataset containing 7311 horizontal hydraulically fractured wells drilled into the middle member of the Bakken Formation from 2010 through 2017. The initial goal is to use these data‐driven machine learning algorithms to identify the most important explanatory predictors of well productivity within nine subareas and the composite area. Predictor variables representing initial gas production, the initial 180‐day water cut, and vertical depth vary spatially and are identified with geologically favorable areas. Well‐completion predictors include the well lateral length, number of fracture stages, volume of proppant per stage, and the volume of injected fluids per stage. The performance of methods is compared based on a common test sample. The analysis then examines the comparative predictive performance of the three algorithms for 1330 wells that had initiated production after the initial 7311 well sample had been producing. The computations of predictor importance identified the initial 180‐day water cut and the 30‐day initial gas production predictors as having a dominant influence in most subareas and for the composite area. The relative importance of well completion predictor variables, that is, the number of fracture stages per well, volume of injected proppant per stage, volume of injected fluids per stage, and lateral length, varied considerably across the subareas. For the common test or holdout sample, the models calibrated with the XGBoost algorithm had superior predictive power. The predictive power of all the algorithms trained on the data from the original sample suffered some loss when tested with a sample of wells that had started production after the end of that period. Implications of the empirical findings and strategies to mitigate loss of predictive power are discussed in the concluding section.

Statistical Analysis and Data Mining

Stochastic analyses to identify wellfield withdrawal effects on surface-water and groundwater in Miami-Dade County, Florida

Several stochastic analyses were conducted in Miami-Dade County, Florida, to evaluate the effects of wellfield withdrawal on aquifer water levels, canal stage, and canal flow. Multiyear data for withdrawals at four water-supply wellfields, water levels at the S-121 canal control structure and groundwater head at a nearby monitoring well were used to determine the interrelation between wellfield withdrawals and water levels in the canal and aquifer. A spectral analysis was performed first on the wellfield withdrawals, showing similar patterns of fluctuations, but no well-defined seasonality. In order to compare water-level response with withdrawals at each wellfield, the intercorrelation effects between wellfields was removed through a ‘causal chain’ approach where the inter-wellfield correlation is used to isolate the wellfield/water-level correlation. Most computed correlations have magnitudes less than 5 percent, but with statistical significance above 90 percent. Results indicate that withdrawals from the wellfields most distant from the canal had no significant correlation to the canal levels. However the highest correlation was not at the wellfield closest to the canal, but at the two wellfields at the intermediate distance that have higher withdrawal rates. The hydraulic interconnectivity of the canal with the rest of the canal network, covering the study area, allows the canal equalizes with all connected canals. This explains why proximity to a particular canal location does not appear to be as important a factor as the withdrawal rate. Groundwater levels are more highly correlated to a wellfield on the same side of the canal, and to pumping wells in the same wellfield on the same side of the canal. This indicates that canals are an effective barrier and source/sink for the groundwater. Further nonlinear correlation analysis indicates that high withdrawal rates disproportionally affect water levels and are the predominant effect on the canal.

Florida

UCODE_2005 and six other computer codes for universal sensitivity analysis, calibration, and uncertainty evaluation constructed using the JUPITER API

This report documents the computer codes UCODE_2005 and six post-processors. Together the codes can be used with existing process models to perform sensitivity analysis, data needs assessment, calibration, prediction, and uncertainty analysis. Any process model or set of models can be used; the only requirements are that models have numerical (ASCII or text only) input and output files, that the numbers in these files have sufficient significant digits, that all required models can be run from a single batch file or script, and that simulated values are continuous functions of the parameter values. Process models can include pre-processors and post-processors as well as one or more models related to the processes of interest (physical, chemical, and so on), making UCODE_2005 extremely powerful. An estimated parameter can be a quantity that appears in the input files of the process model(s), or a quantity used in an equation that produces a value that appears in the input files. In the latter situation, the equation is user-defined. UCODE_2005 can compare observations and simulated equivalents. The simulated equivalents can be any simulated value written in the process-model output files or can be calculated from simulated values with user-defined equations. The quantities can be model results, or dependent variables. For example, for ground-water models they can be heads, flows, concentrations, and so on. Prior, or direct, information on estimated parameters also can be considered. Statistics are calculated to quantify the comparison of observations and simulated equivalents, including a weighted least-squares objective function. In addition, UCODE_2005 can be used fruitfully in model calibration through its sensitivity analysis capabilities and its ability to estimate parameter values that result in the best possible fit to the observations. Parameters are estimated using nonlinear regression: a weighted least-squares objective function is minimized with respect to the parameter values using a modified Gauss-Newton method or a double-dogleg technique. Sensitivities needed for the method can be read from files produced by process models that can calculate sensitivities, such as MODFLOW-2000, or can be calculated by UCODE_2005 using a more general, but less accurate, forward- or central-difference perturbation technique. Problems resulting from inaccurate sensitivities and solutions related to the perturbation techniques are discussed in the report. Statistics are calculated and printed for use in (1) diagnosing inadequate data and identifying parameters that probably cannot be estimated; (2) evaluating estimated parameter values; and (3) evaluating how well the model represents the simulated processes. Results from UCODE_2005 and codes RESIDUAL_ANALYSIS and RESIDUAL_ANALYSIS_ADV can be used to evaluate how accurately the model represents the processes it simulates. Results from LINEAR_UNCERTAINTY can be used to quantify the uncertainty of model simulated values if the model is sufficiently linear. Results from MODEL_LINEARITY and MODEL_LINEARITY_ADV can be used to evaluate model linearity and, thereby, the accuracy of the LINEAR_UNCERTAINTY results. UCODE_2005 can also be used to calculate nonlinear confidence and predictions intervals, which quantify the uncertainty of model simulated values when the model is not linear. CORFAC_PLUS can be used to produce factors that allow intervals to account for model intrinsic nonlinearity and small-scale variations in system characteristics that are not explicitly accounted for in the model or the observation weighting. The six post-processing programs are independent of UCODE_2005 and can use the results of other programs that produce the required data-exchange files. UCODE_2005 and the other six codes are intended for use on any computer operating system. The programs consist of algorithms programmed in Fortran 90/95, which efficiently performs numerical calculations. The model runs required to obtain perturbation sensitivities can be performed using multiple processors. The programs are constructed in a modular fashion using JUPITER API conventions and modules. For example, the data-exchange files and input blocks are JUPITER API conventions and many of those used by UCODE_2005 are read or written by JUPITER API modules. UCODE-2005 includes capabilities likely to be required by many applications (programs) constructed using the JUPITER API, and can be used as a starting point for such programs.

Techniques and Methods

Assessing the role of climate and resource management on groundwater dependent ecosystem changes in arid environments with the Landsat archive

Groundwater dependent ecosystems (GDEs) rely on near-surface groundwater. These systems are receiving more attention with rising air temperature, prolonged drought, and where groundwater pumping captures natural groundwater discharge for anthropogenic use. Phreatophyte shrublands, meadows, and riparian areas are GDEs that provide critical habitat for many sensitive species, especially in arid and semi-arid environments. While GDEs are vital for ecosystem services and function, their long-term (i.e. ~ 30 years) spatial and temporal variability is poorly understood with respect to local and regional scale climate, groundwater, and rangeland management. In this work, we compute time series of NDVI derived from sensors of the Landsat TM, ETM +, and OLI lineage for assessing GDEs in a variety of land and water management contexts. Changes in vegetation vigor based on climate, groundwater availability, and land management in arid landscapes are detectable with Landsat. However, the effective quantification of these ecosystem changes can be undermined if changes in spectral bandwidths between different Landsat sensors introduce biases in derived vegetation indices, and if climate, and land and water management histories are not well understood. The objective of this work is to 1) use the Landsat 8 under-fly dataset to quantify differences in spectral reflectance and NDVI between Landsat 7 ETM + and Landsat 8 OLI for a range of vegetation communities in arid and semiarid regions of the southwestern United States, and 2) demonstrate the value of 30-year historical vegetation index and climate datasets for assessing GDEs. Specific study areas were chosen to represent a range of GDEs and environmental conditions important for three scenarios: baseline monitoring of vegetation and climate, riparian restoration, and groundwater level changes. Google's Earth Engine cloud computing and environmental monitoring platform is used to rapidly access and analyze the Landsat archive along with downscaled North American Land Data Assimilation System gridded meteorological data, which are used for both atmospheric correction and correlation analysis. Results from the cross-sensor comparison indicate a benefit from the application of a consistent atmospheric correction method, and that NDVI derived from Landsat 7 and 8 are very similar within the study area. Results from continuous Landsat time series analysis clearly illustrate that there are strong correlations between changes in vegetation vigor, precipitation, evaporative demand, depth to groundwater, and riparian restoration. Trends in summer NDVI associated with riparian restoration and groundwater level changes were found to be statistically significant, and interannual summer NDVI was found to be moderately correlated to interannual water-year precipitation for baseline study sites. Results clearly highlight the complementary relationship between water-year PPT, NDVI, and evaporative demand, and are consistent with regional vegetation index and complementary relationship studies. This work is supporting land and water managers for evaluation of GDEs with respect to climate, groundwater, and resource management.

Remote Sensing of Environment

Analysis of potential errors in real-time streamflow data and methods of data verification by digital computer

The magnitude, frequency, and types of errors inherent in real-time streamflow data are presented in part I. It was found that real-time data are generally less accurate than are historical data, primarily because real-time data are often used before errors can be detected and corrections applied. Various methods of verifying real-time streamflow data are outlined in part II. Relatively large errors (those greater than 20-30 percent) can be detected readily by use of well-designed verification programs for a digital computer, and smaller errors can be detected only by discharge measurements and field observations. The capability to substitute a simulated discharge value for missing or erroneous data is incorporated in some of the verification routines described. The routines represent concepts ranging from basic statistical comparisons to complex watershed modeling and provide a selection from which real-time data users can choose a suitable level of verification.

Open-File Report

Rejoinder: Sifting through model space

Observational data sets generated by complex processes are common in ecology. Traditionally these have been very challenging to analyze because of the limitations of available statistical tools. This seems to be changing, and these are exciting times to be involved with ecological statistics, not just because of the neo-Bayesian revival but also because of the proliferation of computationally intensive methods in general. It is now possible to fit much richer models to observational data than in the relatively recent past, which in turn has stimulated much interest in how to evaluate and compare such models. In such an immature, vibrant, and rapidly growing field, not everyone is going to agree on the best way to do things. This is reflected in the contrast of opinions offered by the discussants. Each offers a thoughtful and thought-provoking critique of our work that reflects the current thinking in a non-negligible segment of the ecological data analysis community. We want to thank them for their insights.

Ecology

Observations of flocs in an estuary and implications for computation of settling velocity

The settling velocity ( w s ) in estuarine environments can impact whether a region is eroding or accreting sediment on the bed, yet determining this rate can be an indirect process requiring a number of assumptions. Accurate determination of w s is especially needed for numerical models to reproduce observed sediment concentrations at the appropriate timescale. We collected information on suspended sediment flocculation at a channel site (13 m deep) and a shallows site (4 m deep) within South San Francisco Estuary, alongside timeseries of flow, wave statistics, turbulent shear, and bottle samples analyzed for both w s and particle size. Using the measurements of floc size and settling velocity, we performed a sensitivity analysis on the unknown parameters in the general explicit formula for settling velocity. The collected particle size distribution data show that multiple classes of flocs are present; these are characterized as flocculi, microflocs, and macroflocs. We show that w s of flocculi is closest to w s for the full distribution. The determined parameter values lead to near-bed mass-weighted settling velocities (standard deviation) of 1.18 (0.55) and 0.22 (0.15) mm/s at the channel and shallows sites, respectively. Modeling efforts can use this work to help select an appropriate sediment model and parameter values.

California

A summary of ground-water pumpage in the Central Valley, California, 1961-77

In the Central Valley of California, a great agricultural economy has been developed in a semiarid environment. This economy is supported by imported surface water and 9 to 15 million acre-feet per year of ground water. Estimates of ground-water pumpage computed from power consumption have been compiled and summarized. Under ideal conditions, the accuracy of the methods used is about 3 percent. This level of accuracy is not sustained over the entire study area. When pumpage for the entire area is mapped, the estimates seem to be consistent areally and through time. A multiple linear-regression model was used to synthesize data for the years 1961 through 1977, when power data were not available. The model used a relation between ground-water pumpage and climatic indexes to develop a full suite of pumpage data to be used as input to a digital ground-water model, one of the products of the Central Valley Aquifer Project. Statistical analysis of well-perforation data from drillers ' logs and water-temperature data was used to determine the percentage of pumpage that was withdrawn from each of two horizontal layers. (USGS)

Water-Resources Investigations Report

Analyses of meteorological and hydrological records support Tribal members’ accounts of changing climate on the Fort Apache Reservation, east–central Arizona

The Fort Apache Reservation in east–central Arizona, home to the White Mountain Apache Tribe of the Fort Apache Reservation, Arizona, contains several climate zones because of the large variation in surface elevation within the reservation. This study was carried out in cooperation with the White Mountain Apache Tribe of the Fort Apache Reservation, Arizona, to raise awareness of how the changing climate affects the Fort Apache Reservation. This report documents the evaluation of existing multidecadal meteorological and hydrological datasets for the Fort Apache Reservation, used to evaluate the effects of a changing climate on the reservation. In this evaluation, near-surface air temperature, snow depth, snow water equivalent, precipitation, and streamflow datasets were analyzed for monotonic trends indicative of changing climatic conditions during specified periods of time. The results of these trend analyses were then compared with the Tribal community's memories of the changing climate. Trend analysis of near-surface air temperatures from a U.S. Historical Climatological Network station on the Fort Apache Reservation at Whiteriver, Arizona, indicated that mean annual air temperatures have increased by an average of 2.48 degrees Fahrenheit from 1980 to 2023. Records from the same station also indicated that average monthly maximum temperatures recorded for March increased by 5.39 degrees Fahrenheit for the same time period. Annual precipitation at the five precipitation stations used in this study decreased greatly from the 1980s to 2023. The largest total decrease was 10.07 inches, or 34.7 percent. However, only one of the two precipitation stations with longer term data available prior to 1980 had a significant negative trend when data from the entire period of record, from 1901 to 2023, were analyzed. Trend analyses show a decrease in the annual maximum snow water equivalent and an earlier disappearance of the snowpack at two Natural Resources Conservation Service snow telemetry stations in the mountainous region just east of the Fort Apache Reservation from 1981 to 2023. Based on the trend analyses, the average annual maximum snow water equivalent has decreased by more than 40 percent at both stations, and the average date when the snowpack was fully melted at the stations in the spring has moved earlier in time from late April to early April or late March. However, a statistically significant trend was not determined for the early April snow water equivalent measured at a nearby Natural Resources Conservation Service snow course across its period of record, indicating that the history of mountain snowpack in this area is not fully understood. Analysis of snowfall data from a National Oceanic and Atmospheric Administration Cooperative Observer Program network station on the Fort Apache Reservation at McNary 2N, AZ (station 025412) indicated that, on average, the measured total annual snowfall at the station decreased 42.4 percent from 1935 to 2023. Streamflow data from six U.S. Geological Survey streamgages on the Fort Apache Reservation were analyzed for trends. For most streamflow gages, statistically significant trends were not determined for tested parameters when the entire streamflow period of record was used for stations with records going back to at least the 1960s. However, when the data from 1980 to 2023 was tested, most of the streamflow parameters had statistically significant negative trends. All six streamgages showed a decrease in average annual runoff of at least 50 percent from 1980 to 2023; one streamgage showed an 81.8 percent decrease. A similar statistical finding was observed in the analysis of the annual spring snowmelt peak from one of the six streamgages used in the study and located in an area receiving measurable amounts of snowmelt runoff. When data from the entire period of record (1958–2023) was used, no trend in streamflow was determined; however, a significant negative trend was determined from 1980 to 2023, indicating a decrease in average annual springtime runoff of 62.6 percent. Statistical analysis on the timing of the annual spring snowmelt peak at the same streamgage indicated the snowmelt peak is happening on average about 12 days earlier now (2023) than it did in the past. The trend results for the timing of the annual spring snowmelt peak were the same and statistically significant for both periods tested (1958–2023 and 1980–2023). Two of the streamflow records from the Fort Apache Reservation were compared to the Palmer Hydrological Drought Index computed for Arizona Climate Division 4 (East Central) by the National Centers for Environmental Information. The comparison showed that the streamflow records generally tracked the Palmer Hydrological Drought Index. In interviews, Tribal community members living on the Fort Apache Reservation described the changes in climate that they observed during their lifetimes. Common themes reported were that air temperatures have become warmer, and the weather is less predictable with changes in seasonal patterns. Drier conditions, lower snowfall, shorter winters, and lower river levels were also reported. These community member observations align with the results of this study.

Arizona

Agricultural cropland extent and areas of South Asia derived using Landsat satellite 30-m time-series big-data using random forest machine learning algorithms on the Google Earth Engine cloud

The South Asia (India, Pakistan, Bangladesh, Nepal, Sri Lanka and Bhutan) has a staggering 900 million people (~43% of the population) who face food insecurity or severe food insecurity as per United Nations, Food and Agriculture Organization’s (FAO) the Food Insecurity Experience Scale (FIES). The existing coarse-resolution (>250-m) cropland maps lack precision in geo-location of individual farms and have low map accuracies. This also results in uncertainties in cropland areas calculated from such products. Thereby, the overarching goal of this study was to develop high spatial resolution (30-m or better) baseline cropland extent product of South Asia for the year 2015 using Landsat satellite time-series big-data and machine learning algorithms (MLAs) on the Google Earth Engine (GEE) cloud computing platform. To eliminate the impact of clouds, ten time-composited Landsat bands (blue, green, red, NIR, SWIR1, SWIR2, Thermal, EVI, NDVI, NDWI) were derived for each of the 3 time-periods over 12 months (monsoon: Julian days 151-300; winter: Julian days 301-365 plus 1-60; and summer: Julian days 61-150), taking the every 8-day data from Landsat-8 and 7 for the years 2013-2015, for a total of 30-bands plus global digital elevation model (GDEM) derived slope band. This 31-band mega-file big data-cube was composed for each of the 5 agro-ecological zones (AEZ’s) of South Asia and formed a baseline data for image classification and analysis. Knowledge-base for the Random Forest (RF) MLAs were developed using spatially well spread-out reference training data (N=2179) in 5 AEZs. Classification was performed on GEE for each of the 5 AEZs using well-established knowledge-based and RF MLAs on the cloud. Map accuracies were measured using independent validation data (N=1185). The survey showed that the South Asia cropland product had a producer’s accuracy of 89.9% (errors of omissions of 10.1%), user’s accuracy of 95.3% (errors of commission of 4.7%) and an overall accuracy of 88.7%. The National and sub-national (districts) areas computed from this cropland extent product explained 80-96% variability when compared with the National statistics of the South Asian Countries. The full resolution imagery can be viewed at full-resolution, by zooming-in to any location in South Asia or the world, at www.croplands.org and the cropland products of South Asia downloaded from The Land Processes Distributed Active Archive Center (LP DAAC) of National Aeronautics and Space Administration (NASA) and the United States Geological Survey (USGS): https://lpdaac.usgs.gov/products/gfsad30saafgircev001/

GIScience and Remote Sensing

Statistical Comparisons of watershed scale response to climate change in selected basins across the United States

In an earlier global climate-change study, air temperature and precipitation data for the entire twenty-first century simulated from five general circulation models were used as input to precalibrated watershed models for 14 selected basins across the United States. Simulated daily streamflow and energy output from the watershed models were used to compute a range of statistics. With a side-by-side comparison of the statistical analyses for the 14 basins, regional climatic and hydrologic trends over the twenty-first century could be qualitatively identified. Low-flow statistics (95% exceedance, 7-day mean annual minimum, and summer mean monthly streamflow) decreased for almost all basins. Annual maximum daily streamflow also decreased in all the basins, except for all four basins in California and the Pacific Northwest. An analysis of the supply of available energy and water for the basins indicated that ratios of evaporation to precipitation and potential evapotranspiration to precipitation for most of the basins will increase. Probability density functions (PDFs) were developed to assess the uncertainty and multimodality in the impact of climate change on mean annual streamflow variability. Kolmogorov?Smirnov tests showed significant differences between the beginning and ending twenty-first-century PDFs for most of the basins, with the exception of four basins that are located in the western United States. Almost none of the basin PDFs were normally distributed, and two basins in the upper Midwest had PDFs that were extremely dispersed and skewed.

Earth Interactions

Methods for estimating the magnitude and frequency of peak streamflows for unregulated streams in Oklahoma developed by using streamflow data through 2017

The U.S. Geological Survey (USGS), in cooperation with the Oklahoma Department of Transportation, updated peak-streamflow regression equations for estimating flows with annual exceedance probabilities from 50 to 0.2 percent for the State of Oklahoma. These regression equations incorporate basin characteristics to estimate peak-streamflow magnitude and frequency throughout the State by use of a generalized least-squares regression analysis. The most statistically significant independent variables required to estimate peak-streamflow magnitude and frequency for unregulated streams in Oklahoma are contributing drainage area, mean-annual precipitation, and main-channel slope. The regression equations are applicable for stream basins with drainage areas less than 2,510 square miles that are not affected by regulation. The standard model error ranged from 31.28 to 49.32 percent for the different annual exceedance probabilities that were computed. Annual-maximum peak flows observed at 212 USGS streamgages through water year 2017 were used for the regression analysis, excluding the Oklahoma Panhandle region. The USGS StreamStats web application was used to obtain the independent variables required for the peak-streamflow regression equations. Limitations on the use of the regression equations and the reliability of regression estimates for natural unregulated streams are described. Log-Pearson Type III analysis information, basin and climate characteristics, and the peak-streamflow frequency estimates for the 212 streamgages in and near Oklahoma are provided in this report. This report contains descriptions of the methods that can be used to estimate peak streamflows at ungaged sites by using estimates from streamgages on unregulated streams. For ungaged sites on urban streams and streams regulated by small floodwater-retarding structures, an adjustment of the statewide regression equations for natural unregulated streams can be used to estimate peak-streamflow magnitude and frequency.

Arkansas, Kansas, Missouri, Oklahoma, Texas