Search USGSSearch

SEARCH · Search USGS

Results for “Computational Statistics and Data Analysis”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Singularity and Nonnormality in the Classification of Compositional Data

Geologists may want to classify compositional data and express the classification as a map. Regionalized classification is a tool that can be used for this purpose, but it incorporates discriminant analysis, which requires the computation and inversion of a covariance matrix. Covariance matrices of compositional data always will be singular (noninvertible) because of the unit-sum constraint. Fortunately, discriminant analyses can be calculated using a pseudo-inverse of the singular covariance matrix; this is done automatically by some statistical packages such as SAS. Granulometric data from the Darss Sill region of the Baltic Sea is used to explore how the pseudo-inversion procedure influences discriminant analysis results, comparing the algorithm used by SAS to the more conventional Moore-Penrose algorithm. Logratio transforms have been recommended to overcome problems associated with analysis of compositional data, including singularity. A regionalized classification of the Darss Sill data after logratio transformation is different only slightly from one based on raw granulometric data, suggesting that closure problems do not influence severely regionalized classification of compositional data.

Mathematical Geology

Positional accuracy assessment of lidar point cloud from NAIP/3DEP pilot project

The Leica Geosystems CountryMapper hybrid system has the potential to collect data that satisfy the U.S. Geological Survey (USGS) National Geospatial Program (NGP) and 3D Elevation Program (3DEP) and the U.S. Department of Agriculture (USDA) National Agriculture Imagery Program (NAIP) requirements in a single collection. This research will help 3DEP determine if this sensor has the potential to meet current and future 3DEP topographic lidar collection requirements. We performed an accuracy analysis and assessment on the lidar point cloud produced from CountryMapper. The boresighting calibration and co-registration by georeferencing correction based on ground control points are assumed to be performed by the data provider. The scope of the accuracy assessment is to apply the following variety of ways to measure the accuracy of the delivered point cloud to obtain the error statistics. Intraswath uncertainty from a flat surface was computed to evaluate the point cloud precision. Intraswath difference between opposite scan directions and the interswath overlap difference were evaluated to find boresighting or any systematic errors. Absolute vertical accuracy over vegetated and non-vegetated areas were also assessed. Both horizontal and vertical absolute errors were assessed using the 3D absolute error analysis methodology of comparing conjugate points derived from geometric features. A three-plane feature makes a single unique intersection point. Intersection points were computed from ground-based lidar and airborne lidar point clouds for comparison. The difference between two intersection points form one error vector. The geometric feature-based error analysis was applied to intraswath, interswath, and absolute error analysis. The CountryMapper pilot data appear to satisfy the accuracy requirements suggested by the USGS lidar specification, based upon the error analysis results. The focus of this research was to demonstrate various conventional accuracy measures and novel 3D accuracy techniques using two different error computation methods on the CountryMapper airborne lidar point cloud.

Remote Sensing

A satellite-based digital data system for low-frequency geophysical data

A reliable method for collection, display, and analysis of low-frequency geophysical data from isolated sites, which can be throughout North and South America and the Pacific Rim, has been developed for use with the Geostationary Operational Environmental Satellite (GOES) system. Geophysical data primarily intended for earthquake hazard and crustal deformation monitoring are digitized with either 12-bit or 16-bit resolution and transmitted every 10 min through a satellite link to a bank of UNIX-based computers in Menlo Park, California. There the data are available for analysis and display within a few seconds of their transmit time. This system provides real-time monitoring of crustal deformation parameters such as tilt, strain, fault displacement, local magnetic field, crustal geochemistry, and water levels, as well as meteorological and other parameters, along faults in California and Alaska, and in volcanic regions in the western United States, Rabaul, and other locations in the New Britain region of the South Pacific. Various mathematical, statistical, and graphical algorithms process the incoming data to detect changes in crustal deformation and fault slip that may indicate the first stages of catastrophic fault failure. Alert trigger levels based on physical models, signal resolution, and previous history have been defined for particular instrument types. Computer-driven remote paging and mail systems are used to notify appropriate personnel when alarm status is reached. The system supports continuous historical records of low-frequency geophysical data, software for extensive analysis of these data, and programs for modeling fault rupture with and without seismic radiation, as well as providing an environment for real-time attempts at earthquake prediction.

Bulletin of the Seismological Society of America

Using the Lomb-Scargle method for wave statistics from gappy time series

Sandwich Town Neck Beach in Sandwich, MA, has experienced substantial erosion and has been the subject of efforts by the town and private landowners to limit the sand loss. Erosion has been particularly dramatic in the past five years with the loss of dwellings. Sandwich's nourishment efforts presented a unique opportunity for scientists at the U.S. Geological Survey Woods Hole Coastal and Marine Science Center to monitor beach morphology and to test new technologies and techniques such as geo-referenced drone imaging. Two bottom lander deployments were performed in Cape Cod Bay at a location that was key to model the fate of waves at Sandwich Town Neck Beach and to support the study of beach morphological evolution. The study period was after the town nourished the beach and during a time when several intense winter storms reshaped the beach and removed much of the nourished sand. A TRDI Workhorse Sentinel V ADCP was used for both deployments. For wave bursts, the instruments collected 2048 samples at 2 Hz every hour. The first deployment during the winter of 2016 returned good quality data. The second deployment during the following winter had gaps throughout the time series from a wiring problem in the external battery pack. The timing of the gaps was random, the duration approximately 100 s. While most of the bursts started at the top of each hour, many had 1-3 gaps within. Time series data with random gaps are problematic for computing spectral density, and thus, wave statistics. This kind of situation is familiar in other scientific disciplines such as astrophysics [1], where techniques exist to find stationary signals in sparse data. One of these methods is the Lomb-Scargle technique for computing periodograms. The most useful feature of the Lomb-Scargle (LS) method is that it allows the spectral analysis of incomplete records, without having to manipulate the record to extrapolate from or replace missing data. We compared the effectiveness of LS against common methods of averaging Fourier transforms such as a simple un-windowed Fast Fourier transform (FFT), Welch's method, and TRDI's Wavesmon software; methods that are commonly used in oceanography for non-gappy data. Synthetic data series that have been artificially modified to introduce gaps were used to evaluate the performance of each method. The LS approach was able to recover spectral density even with several 100-s gaps present. The method was applied here to the gappy and non-gappy data from both Sandwich deployments, and wave statistics were obtained and compared to the wave-buoy data. LS was used to process data that contains gaps that was rejected by Wavesmon, which was approximately 39% of the dataset. Significant wave height and peak period from LS compared well with buoy data. Mean period computed on gappy data using LS produced values biased low, compared with other methods when gaps were filled with the mean value. The LS technique has potential to uncover low-frequency signals such as infragravity waves from gappy records where the non-gappy segments are not long enough to resolve them. It has potential to unlock new information from older data sets.

Massachusetts

Development and application of a mark-recapture model incorporating predicted sex and transitory behaviour

We developed an extension of Cormack-Jolly-Seber models to handle a complex mark-recapture problem in which (a) the sex of birds cannot be determined prior to first moult, but can be predicted on the basis of body measurements, and (b) a significant portion of captured birds appear to be transients (i.e. are captured once but leave the area or otherwise become ' untrappable'). We applied this methodology to a data set of 4184 serins (Serinus serinus) trapped in northeastern Spain during 1985-96, in order to investigate age-, sex-, and time-specific variation in survival rates. Using this approach, we were able to successfully incorporate the majority of ringings of serins. Had we eliminated birds not previously captured (as has been advocated to avoid the problem of transience) we would have reduced our sample sizes by >2000 releases. In addition, we were able to include 1610 releases of birds of unknown (but predicted) sex; these data contributed to the precision of our estimates and the power of statistical tests. We discuss problems with data structure, encoding of the algorithms to compute parameter estimates, model selection, identifiability of parameters, and goodness-of-fit, and make recommendations for the design and analysis of future studies facing similar problems.

Bird Study

Problems with sampling desert tortoises: A simulation analysis based on field data

The desert tortoise (Gopherus agassizii) was listed as a U.S. threatened species in 1990 based largely on population declines inferred from mark-recapture surveys of 2.59-km2 (1-mi2) plots. Since then, several census methods have been proposed and tested, but all methods still pose logistical or statistical difficulties. We conducted computer simulations using actual tortoise location data from 2 1-mi2 plot surveys in southern California, USA, to identify strengths and weaknesses of current sampling strategies. We considered tortoise population estimates based on these plots as "truth" and then tested various sampling methods based on sampling smaller plots or transect lines passing through the mile squares. Data were analyzed using Schnabel's mark-recapture estimate and program CAPTURE. Experimental subsampling with replacement of the 1-mi2 data using 1-km2 and 0.25-km2 plot boundaries produced data sets of smaller plot sizes, which we compared to estimates from the 1-mi 2 plots. We also tested distance sampling by saturating a 1-mi 2 site with computer simulated transect lines, once again evaluating bias in density estimates. Subsampling estimates from 1-km2 plots did not differ significantly from the estimates derived at 1-mi2. The 0.25-km2 subsamples significantly overestimated population sizes, chiefly because too few recaptures were made. Distance sampling simulations were biased 80% of the time and had high coefficient of variation to density ratios. Furthermore, a prospective power analysis suggested limited ability to detect population declines as high as 50%. We concluded that poor performance and bias of both sampling procedures was driven by insufficient sample size, suggesting that all efforts must be directed to increasing numbers found in order to produce reliable results. Our results suggest that present methods may not be capable of accurately estimating desert tortoise populations.

California

Streamflow characteristics and trends in New Jersey, water years 1897-2003

Streamflow statistics were computed for 111 continuous-record streamflow-gaging stations with 20 or more years of continuous record and for 500 low-flow partial-record stations, including 66 gaging stations with less than 20 years of continuous record. Daily mean streamflow data from water year 1897 through water year 2001 were used for the computations at the gaging stations. (The water year is the 12-month period, October 1 through September 30, designated by the calendar year in which it ends). The characteristics presented for the long-term continuous-record stations are daily streamflow, harmonic mean flow, flow frequency, daily flow durations, trend analysis, and streamflow variability. Low-flow statistics for gaging stations with less than 20 years of record and for partial-record stations were estimated by correlating base-flow measurements with daily mean flows at long-term (more than 20 years) continuous-record stations. Instantaneous streamflow measurements through water year 2003 were used to estimate low-flow statistics at the partial-record stations. The characteristics presented for partial-record stations are mean annual flow; harmonic mean flow; and annual and winter low-flow frequency. The annual 1-, 7-, and 30-day low- and high-flow data sets were tested for trends. The results of trend tests for high flows indicate relations between upward trends for high flows and stream regulation, and high flows and development in the basin. The relation between development and low-flow trends does not appear to be as strong as for development and high-flow trends. Monthly, seasonal, and annual precipitation data for selected long-term meteorological stations also were tested for trends to analyze the effects of climate. A significant upward trend in precipitation in northern New Jersey, Climate Division 1 was identified. For Climate Division 2, no general increase in average precipitation was observed. Trend test results indicate that high flows at undeveloped, unregulated sites have not been affected by the increase in average precipitation. The ratio of instantaneous peak flow to 3-day mean flow, ratios of flow duration, ratios of high-flow/low-flow frequency, and coefficient of variation were used to define streamflow variability. Streamflow variability was significantly greater among the group of gaging stations located outside the Coastal Plain than among the group of gaging stations located in the Coastal Plain.

Scientific Investigations Report

Revised recommended methods for analyzing crater size-frequency distributions

Impact crater populations crucially help us to understand solar system dynamics, planetary surface histories, and surface modification processes. A single previous effort to standardize how crater data are displayed in graphs, tables, and archives, was in a 1978 NASA report by the Crater Analysis Techniques Working Group, published in 1979 in Icarus . The report had a significant lasting effect, but later decades brought major advances in statistical and computer sciences while the crater field has remained fairly stagnant. In this new work, we revisit the fundamental techniques for displaying and analyzing crater population data and demonstrate better statistical methods that can be used. Specifically, we address (1) how crater size-frequency distributions (SFDs) are constructed, (2) how error bars are assigned to SFDs, and (3) how SFDs are fit to power laws and other models. We show how the new methods yield results similar to those of previous techniques in that the SFDs have familiar shapes but better account for multiple sources of uncertainty. We also recommend graphic, display, and archiving methods that reflect computers' capabilities and fulfill NASA's current requirements for Data Management Plans.

Meteoritics and Planetary Science

Methods for estimating magnitude and frequency of floods in Arizona, developed with unregulated and rural peak-flow data through water year 2010

Flooding is among the worst natural disasters responsible for loss of life and property in Arizona, underscoring the importance of accurate estimation of flood magnitude for proper structural design and floodplain mapping. Twenty-four years of additional peak-flow data have been recorded since the last comprehensive regional flood frequency analysis conducted in Arizona. Periodically, flood frequency estimates and regional regression equations must be revised to maintain the accurate estimation of flood frequency and magnitude. Annual peak-flow data collected through water year 2010 were compiled from 448 unregulated streamflow-gaging stations, hereafter referred to as streamgages, in Arizona having a minimum of 10 years of record. Flood frequency estimates were first computed with station (or at-site) skew using the Expected Moments Algorithm with a multiple Grubbs-Beck test to identify multiple potentially influential low flows to fit a Pearson Type III distribution. Next, a multiple step Bayesian least-squares-regression approach was used to determine a new statewide regional skew of −0.09. No basin characteristics analyzed were statistically significant in explaining the variation in skew and as a result, the constant model was chosen as the best regional skew model for the Arizona study area. The mean square error used in Bulletin 17B (B17B) of the Interagency Advisory Committee on Water Data is used to describe the precision of the regional skew. The constant model had a mean square error equal to 0.08, which corresponds to an effective record length of 85 years. This is a marked improvement over a previous Arizona regional skew analysis, with a reported mean square error of 0.31, for a corresponding effective record length of around 17 years. Thus the new regional model had almost five times the information content (as measured by effective record length) of that calculated in USGS Water Supply Paper 2433, published in 1997, or the value of 0.302 reported in the B17B generalized skew map. The flood frequency estimates were recalculated using a weighted skew of the station and regional skew. Station flood frequency estimates for each streamgage are presented for the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities. Geographical information systems were used to compute basin characteristic information for each streamgage for the purpose of developing regional equations to estimate flood statistics at ungaged basins. Five hydrologic flood regions in Arizona were defined in a multivariate regionalization process based on mean basin elevation, mean annual precipitation, and soil permeability. A regional generalized least-squares-regression analysis was used to develop five sets of equations from 344 nonredundant streamgages, corresponding to five regions, for estimating the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities at ungaged basins in Arizona. The regression equations developed for these five regions were based on one or more of the statistically significant explanatory variables: drainage area, mean basin elevation, and mean annual precipitation. Average standard errors of prediction for the regression regions for the five regions ranged from 27 to 122 percent and the pseudo-coefficients of determination (pseudo-R 2 ), a measure of the proportion of peak-flow variation that is explained by the basin characteristics, ranged from 68 to 98 percent. Regression equations for Central Highlands (region 4) had the lowest model error and the greatest pseudo-R 2 metrics. The equations for Colorado Plateau (region 2) regression equations generally had greater model error and lower pseudo-R 2 metrics. The improvement of regional regression equation model error and pseudo-R 2 metrics was related to higher numbers of streamgages, longer period of record, and even spatial coverage within a region. The regional regression equations were integrated into the U.S. Geological Survey’s StreamStats program. The StreamStats program is a national map-based web application that allows the public to easily access published flood frequency and basin characteristic statistics. The interactive web application allows a user to select a point within a watershed (gaged or ungaged) and retrieve flood-frequency estimates derived from the current regional regression equations and geographic information system data within the selected basin. StreamStats provides users with an efficient and accurate means for retrieving the most up to date flood frequency and basin characteristic data. StreamStats is intended to provide consistent statistics, minimize user error, and reduce the need for large datasets and costly geographic information system software.

Arizona

Statistical frequency analysis of flood records

The U.S. Geological Survey, like other Federal agencies, uses Hydrology Subcommittee Bulletin 17 for guidance in statistical frequency analysis of flood records. This paper describes the formal statistical and computational aspects of the Bulletin 17 methodology. The methodology includes provisions for dealing with high and low out-liers, historic peaks, and other anomalous flood data. If these options are inadequate, alternative procedures may be used if properly documented.

Conference Paper

Changes in low-flow frequency from 1976-2006 at selected streamgages in New York, excluding Long Island

Many Federal, State, and local agencies use low-flow data to establish water-use policy and help determine the total maximum daily loads and effluent limits of point and nonpoint sources of contamination of surface water during periods of decreased streamflow. Low-flow magnitude and frequency are used often by water-supply planners, reservoir managers, and hydroelectric facilities to manage water availability for supply and power generation. Low-flow statistics for eight selected U.S. Geological Survey streamgages in New York State were calculated for the period from 1976 through 2006 and for the entire period of continuous streamflow record. The 7-day, 2-year and 10-year low flows were computed and compared with those low flows published in the1979 U.S. Geological Survey report, Low-flow frequency analysis of streams in New York, Bulletin 74. Observed changes in low-flow frequency at each gage were then examined and compared to changes in precipitation and land use to determine whether a relation between similar patterns could be identified. A statewide U.S. Geological Survey study has not been done to develop equations for estimating low flows on rural unregulated streams in New York. Currently (2010) only one regional study developed for parts of the lower Hudson River Basin in 1986 is available to assist in estimating low flows on rural streams with unregulated streamflow in New York. Low-flow statistics published in the 1979 report need to be updated by using additional data collected since 1976 to determine current low-flow conditions across New York State. At-site low-flow statistics were updated for eight streamgages in New York by using continuous daily streamflow data through 2006 for the future development of a statewide research study. Selection of the eight streamgages used in this study identified a major deficiency in the number of available unregulated long-term U.S. Geological Survey streamgages needed for the development of regional low-flow equations in New York. A limited analysis of the changes in land use for the contributing drainage areas for each streamgage, changes in precipitation, and trends in the annual 7-day minimum flow also are presented. The 7-day, 2-year low flow showed increases of 14 to 35 percent and the 7-day 10-year low flow showed zero to 19 percent increases at rural streamgages with unregulated streamflows when statistics were computed by using data from 1976 through 2006 and compared with published data in Bulletin 74. When the entire period of record was used to compute low flow frequencies, the 7-day, 2-year low flows increased from about 6 to 15 percent whereas the 7-day 10-year low flows showed zero to 5 percent increases. Streamgages affected by urbanization and regulation for water supply showed the most significant changes in the 7-day, 2-year and 10-year low-flow frequencies. These streamgages are included to help identify the effects of urbanization and regulation on streamflow at these locations. The 7-day 10-year low flow increased by 65 percent at the U.S. Geological Survey streamgage Hackensack River at West Nyack, N.Y., and increased 120 percent at the U.S. Geological Survey streamgage Neversink River at Godeffroy, N.Y., when statistics were computed by using data from 1976 through 2006 and compared with the statistics for the regulated period computed in Bulletin 74.

New York

A theory for modeling ground-water flow in heterogeneous media

Construction of a ground-water model for a field area is not a straightforward process. Data are virtually never complete or detailed enough to allow substitution into the model equations and direct computation of the results of interest. Formal model calibration through optimization, statistical, and geostatistical methods is being applied to an increasing extent to deal with this problem and provide for quantitative evaluation and uncertainty analysis of the model. However, these approaches are hampered by two pervasive problems: 1) nonlinearity of the solution of the model equations with respect to some of the model (or hydrogeologic) input variables (termed in this report system characteristics) and 2) detailed and generally unknown spatial variability (heterogeneity) of some of the system characteristics such as log hydraulic conductivity, specific storage, recharge and discharge, and boundary conditions. A theory is developed in this report to address these problems. The theory allows construction and analysis of a ground-water model of flow (and, by extension, transport) in heterogeneous media using a small number of lumped or smoothed system characteristics (termed parameters). The theory fully addresses both nonlinearity and heterogeneity in such a way that the parameters are not assumed to be effective values. The ground-water flow system is assumed to be adequately characterized by a set of spatially and temporally distributed discrete values, ?, of the system characteristics. This set contains both small-scale variability that cannot be described in a model and large-scale variability that can. The spatial and temporal variability in ? are accounted for by imagining ? to be generated by a stochastic process wherein ? is normally distributed, although normality is not essential. Because ? has too large a dimension to be estimated using the data normally available, for modeling purposes ? is replaced by a smoothed or lumped approximation y?. (where y is a spatial and temporal interpolation matrix). Set y?. has the same form as the expected value of ?, y 'line' ? , where 'line' ? is the set of drift parameters of the stochastic process; ?. is a best-fit vector to ?. A model function f(?), such as a computed hydraulic head or flux, is assumed to accurately represent an actual field quantity, but the same function written using y?., f(y?.), contains error from lumping or smoothing of ? using y?.. Thus, the replacement of ? by y?. yields nonzero mean model errors of the form E(f(?)-f(y?.)) throughout the model and covariances between model errors at points throughout the model. These nonzero means and covariances are evaluated through third and fifth-order accuracy, respectively, using Taylor series expansions. They can have a significant effect on construction and interpretation of a model that is calibrated by estimating ?.. Vector ?.. is estimated as 'hat' ? using weighted nonlinear least squares techniques to fit a set of model functions f(y'hat' ?) to a. corresponding set of observations of f(?), Y. These observations are assumed to be corrupted by zero-mean, normally distributed observation errors, although, as for ?, normality is not essential. An analytical approximation of the nonlinear least squares solution is obtained using Taylor series expansions and perturbation techniques that assume model and observation errors to be small. This solution is used to evaluate biases and other results to second-order accuracy in the errors. The correct weight matrix to use in the analysis is shown to be the inverse of the second-moment matrix E(Y-f(y?.))(Y-f(y?.))', but the weight matrix is assumed to be arbitrary in most developments. The best diagonal approximation is the inverse of the matrix of diagonal elements of E(Y-f(y?.))(Y-f(y?.))', and a method of estimating this diagonal matrix when it is unknown is developed using a special objective function to compute 'hat' ?. When considered to be an estimate of f

Professional Paper

Estimating concentrations of fine-grained and total suspended sediment from close-range remote sensing imagery

Fluvial sediment, a vital surface water resource, is hazardous in excess. Suspended sediment, the most prevalent source of impairment of river systems, can adversely affect flood control, navigation, fisheries and aquatic ecosystems, recreation, and water supply (e.g., Rasmussen et al., 2009; Qu, 2014). Monitoring programs typically focus on suspended-sediment concentration (SSC) and discharge (SSQ). These time-series data are used to study changes to basin hydrology, geomorphology, and ecology caused by disturbances. The U.S. Geological Survey (USGS) has traditionally used physical sediment sample-based methods (Edwards and Glysson, 1999; Nolan et al., 2005; Gray et al., 2008) to compute SSC and SSQ from continuous streamflow data using a sediment transport-curve (e.g., Walling, 1977) or hydrologic interpretation (Porterfield, 1972). Accuracy of these data is typically constrained by the resources required to collect and analyze intermittent physical samples. Quantifying SSC using continuous instream turbidity is rapidly becoming common practice among sediment monitoring programs. Estimations of SSC and SSQ are modeled from linear regression analysis of concurrent turbidity and physical samples. Sediment-surrogate technologies such as turbidity promise near real-time information, increased accuracy, and reduced cost compared to traditional physical sample-based methods (Walling, 1977; Uhrich and Bragg, 2003; Gray and Gartner, 2009; Rasmussen et al., 2009; Landers et al., 2012; Landers and Sturm, 2013; Uhrich et al., 2014). Statistical comparisons among SSQ computation methods show that turbidity-SSC regression models can have much less uncertainty than streamflow-based sediment transport-curves or hydrologic interpretation (Walling, 1977; Lewis, 1996; Glysson et al., 2001; Lee et al., 2008). However, computation of SSC and SSQ records from continuous instream turbidity data is not without challenges; some of these include environmental fouling, calibration, and data range among sensors. Of greatest interest to many programs is a hysteresis in the relationship between turbidity and SSC, attributed to temporal variation of particle size distribution (Landers and Sturm, 2013; Uhrich et al., 2014). This phenomenon causes increased uncertainty in regression-estimated values of SSC, due to changes in nephelometric reflectance off the varying grain sizes in suspension (Uhrich et al., 2014). Here, we assess the feasibility and application of close-range remote sensing to quantify SSC and particle size distribution of a disturbed, and highly-turbid, river system. We use a consumer-grade digital camera to acquire imagery of the river surface and a depth-integrating sampler to collect concurrent suspended-sediment samples. We then develop two empirical linear regression models to relate image spectral information to concentrations of fine sediment (clay to silt) and total suspended sediment. Before presenting our regression model development, we briefly summarize each data-acquisition method.

Conference Paper

Peak-flow and low-flow magnitude estimates at defined frequencies and durations for nontidal streams in Delaware

Reliable estimates of the magnitude of peak flows in streams are required for the economical and safe design of transportation and water conveyance structures. In addition, reliable estimates of the magnitude of low flows at defined frequencies and durations are needed for meeting regulatory requirements, quantifying base flows in streams and rivers, and evaluating time of travel and dilution of toxic spills. This report, in cooperation with the Delaware Department of Transportation and the Delaware Geological Survey, presents methods for estimating the magnitude of peak flows and low flows at defined frequencies and durations on nontidal streams in Delaware, at locations both monitored by streamflow-gage sites and ungaged. Methods are presented for estimating (1) the magnitude of peak flows for return periods ranging from 2 to 500 years (50-percent to 0.2-percent annual-exceedance probability), and (2) the magnitude of low flows as applied to 7-, 14-, and 30-consecutive day low-flow periods with recurrence intervals of 2, 10, and 20 years (50-, 10-, and 5-percent annual non-exceedance probabilities). These methods are applicable to watersheds that exhibit a full range of development conditions in Delaware. The report also describes StreamStats, a web application that allows users to easily obtain peak-flow and low-flow magnitude estimates for user-selected locations in Delaware. Peak-flow and low-flow magnitude estimates for ungaged sites are obtained using statistical regression analysis through a process known as regionalization, where information from a group of streamflow-gage sites within a region forms the basis for estimates for ungaged sites within the same region. Ninety-four streamflow-gage sites in and near Delaware with at least 10 years of nonregulated annual peak-flow data were used for the peak-flow regression analysis, a subset of the 121 sites for which peak-flow estimates were computed. These sites included both continuous-record streamflow-gage sites as well as partial record sites. Forty-five streamflow-gage sites with at least 10 years of nonregulated low-flow data available were used for the low-flow regression analyses, a subset of the 68 sites for which low-flow estimates were computed. Estimates for gaged sites are obtained by combining (1) the station peak-flow statistics (mean, standard deviation, and skew) and peak-flow estimates using the recent Bulletin 17C guidelines that incorporate the Expected Moments Algorithm with (2) regional estimates of peak-flow magnitude derived from regional regression equations and regional skew derived from sites with records greater than or equal to 35 years. Example peak-flow estimate calculations using the methods presented in the report are given for (1) ungaged sites, (2) gaged sites, (3) sites upstream or downstream from a gaged location, and (4) sites between gaged locations. Estimates for low-flow gaged sites are obtained by combining (1) the station low-flow statistics (mean, standard deviation, and skew) and low-flow estimates with (2) regional estimates of low-flow magnitude derived from regional regression equations. Example low-flow estimate calculations using the methods presented in the report are given for (1) ungaged sites, (2) gaged sites, (3) sites upstream or downstream from a gaged location, and (4) sites between gaged locations. A total of 54 sites in the Coastal Plain region were used to develop peak-flow regressions for the region and 40 sites were used for the Piedmont region. Similarly, 24 sites were used for low-flow regression equation development in the Coastal Plain, with 21 in the Piedmont. Peak and low-flow site inclusion in the Coastal Plain tended to be more restricted with tidal influence and ranges of basin characteristics, including drainage area, limiting regression equation development and application. Regional regression equations for peak flows and low flows, as applicable to ungaged sites in the Piedmont and Coastal Plain Physiographic Provinces in Delaware, are presented. Peak-flow regression equations used variables that quantified drainage area, basin slope, percent area with well-drained soils, percent area with poorly drained soils, impervious area, and percent area of surface water storage in estimating peak-flow estimates, whereas low-flow regression equations used only drainage area and percent poorly drained soils in the estimation of low flows. Average standard errors for peak-flow regressions tended to be lower than those for low- flow regressions, with lower errors in the Piedmont region for both peak- and low-flow regressions. For peak-flow estimates, a sensitivity analysis of Piedmont regression equation estimates to changes in impervious area is also presented. Additional topics associated with the analyses performed during the study are discussed, including (1) the availability and description of 32 basin and climatic characteristics considered during the development of the regional regression equations; (2) the treatment of increasing trends in the annual peak-flow series identified at 18 gaged sites and inclusion in or exclusion from the regional analysis; (3) regional skew analysis and determination of regression regions; (4) sample adjustments and removal of sites owing to regulation and redundancy; and (5) a brief comparison of peak- and low-flow estimates at gages used in previous studies.

Delaware

Sediment concentrations and loads upstream from and through John Redmond Reservoir, east-central Kansas, 2010–19

Streambank erosion and reservoir sedimentation are primary concerns of resource managers in Kansas and throughout many regions of the United States and negatively affect flood control, water supply, and recreation. The Cottonwood and upper Neosho Rivers drain into John Redmond Reservoir, and since reservoir completion in 1964, there has been substantial conservation-pool sedimentation and storage loss in John Redmond Reservoir, causing storage capacity losses more rapidly than most other Federal reservoirs in Kansas. The U.S. Geological Survey (USGS), in cooperation with the Kansas Water Office, has monitored water quality (temperature, specific conductance, and turbidity) on the Cottonwood River (upstream from the reservoir) and Neosho River (upstream and downstream from the reservoir) since 2007 with additional sites added in 2009. The purpose of this report is to quantify suspended-sediment concentrations, loads, and yields entering and exiting John Redmond Reservoir during January 1, 2010, through December 31, 2019. Three water-quality monitoring sites were upstream from the reservoir (Cottonwood River near Plymouth, Kansas [USGS site 07182250; hereinafter referred to as “Cottonwood”]; Neosho River at Burlingame Road near Emporia, Kans. [USGS site 07179750; hereinafter referred to as “Burlingame”]; and Neosho River at Neosho Rapids, Kans. [USGS site 07182390; hereinafter referred to as “Neosho Rapids”]), and one water-quality monitoring site was downstream from the reservoir (Neosho River at Burlington, Kans. [USGS site 07182510; hereinafter referred to as “Burlington”]). The Neosho Rapids streamgage is downstream from the confluence of the Cottonwood and upper Neosho Rivers and has a contributing drainage area accounting for 91 percent of the total contributing drainage area to John Redmond Reservoir. Continuously measured streamflow, water quality, and discrete water-quality data were used to develop updated regression models to compute suspended-sediment concentrations, loads, and yields upstream and downstream from John Redmond Reservoir in east-central Kansas. Several turbidity sensors were deployed during the analysis period, and there are no established relations between the sensors; therefore, individual models for each sensor were developed. Model statistics for the turbidity and suspended-sediment concentration linear regression models were better (based on the coefficient of determination, root mean square error, and model standard percentage error) than the streamflow and suspended-sediment concentration linear regression models, indicating better model performance. Computed concentrations, loads, and yields do not account for the ungaged 9 percent of the drainage basin downstream from the Neosho Rapids streamgage. Mean daily suspended-sediment loads upstream from the reservoir were largest at Neosho Rapids (2,250 tons), second largest at Cottonwood (2,180 tons), and smallest at Burlingame (624 tons). Streamflow at Burlington was predominately regulated by reservoir releases, and mean daily suspended-sediment loads were smaller (286 tons) than at upstream sites. Among the upstream sites, Cottonwood had the largest mean daily suspended-sediment concentration (179 milligrams per liter [mg/L]), followed by Neosho Rapids (162 mg/L), and Burlingame (108 mg/L). Burlington had the smallest mean daily suspended-sediment concentration of all sites (46 mg/L). Annual reservoir trapping efficiency ranged from 82 to 94 percent, and the largest sediment mass trapped was during 2019 (2,230,000 tons). Reservoir storage decreased an estimated 7,750 acre-feet during 2010 and 2014–19. Using the mean trapping efficiency to estimate suspended-sediment loads during years with missing data (2011–13), the total estimated reservoir storage lost to sedimentation for the analysis period (2010–19) was 8,690 acre-feet, about 17 percent of the remaining storage space reported in 2007. The mean annual sedimentation rate during the analysis period (747 acre-feet per year) was about 85 percent larger than the design sedimentation rate (404 acre-feet per year) originally projected during construction. Different reservoir outflow management strategies, including operating near normal capacity as opposed to higher flood pool levels, could reduce the total reservoir storage lost by 3 percent (about 261 acre-feet), which is equal to 14 percent of the total sediment removed during the dredging operation in 2016. During the study period, about 56 percent of the total suspended-sediment load was transported during streamflows greater than the National Weather Service flood action stage at the upstream sites (0.1–5 percent of the record; Cottonwood mean: 48 percent; Burlingame mean: 40 percent; Neosho Rapids mean: 78 percent). Disproportionately large sediment loads were delivered during short periods of time, and localized efforts of stream erosion protection (streambank stabilization, riparian buffers) were likely to be overwhelmed. Precipitation frequency and intensity are projected to continue to increase in this region; therefore, future sediment reduction strategies that account for extreme episodic events may be beneficial. Changes to reservoir outflow management could also minimize sediment accumulation while still preserving flood control. Continued investigation of sediment reduction measures is necessary for future mitigation with the understanding that sedimentation rate is largely driven by high flows. Results from this study can be used to calibrate sediment models, explore sediment reduction strategies, highlight the importance of continued water-quality monitoring to determine effectiveness and changes in sediment transport, and assess the ability of John Redmond Reservoir to support designated uses into the future.

Kansas

Analysis of flood-magnitude and flood-frequency data for streamflow-gaging stations in the Delaware and North Branch Susquehanna River Basins in Pennsylvania

The Delaware and North Branch Susquehanna River Basins in Pennsylvania experienced severe flooding as a result of intense rainfall during June 2006. The height of the flood waters on the rivers and tributaries approached or exceeded the peak of record at many locations. Updated flood-magnitude and flood-frequency data for streamflow-gaging stations on tributaries in the Delaware and North Branch Susquehanna River Basins were analyzed using data through the 2006 water year to determine if there were any major differences in the flood-discharge data. Flood frequencies for return intervals of 2, 5, 10, 50, 100, and 500 years (Q2, Q5, Q10, Q50, Q100, and Q500) were determined from annual maximum series (AMS) data from continuous-record gaging stations (stations) and were compared to flood discharges obtained from previously published Flood Insurance Studies (FIS) and to flood frequencies using partial-duration series (PDS) data. A Wilcoxon signed-rank test was performed to determine any statistically significant differences between flood frequencies computed from updated AMS station data and those obtained from FIS. Percentage differences between flood frequencies computed from updated AMS station data and those obtained from FIS also were determined for the 10, 50, 100, and 500 return intervals. A Mann-Kendall trend test was performed to determine statistically significant trends in the updated AMS peak-flow data for the period of record at the 41 stations. In addition to AMS station data, PDS data were used to determine flood-frequency discharges. The AMS and PDS flood-frequency data were compared to determine any differences between the two data sets. An analysis also was performed on AMS-derived flood frequencies for four stations to evaluate the possible effects of flood-control reservoirs on peak flows. Additionally, flood frequencies for three stations were evaluated to determine possible effects of urbanization on peak flows. The results of the Wilcoxon signed-rank test showed a significant difference at the 95-percent confidence level between the Q100 computed from AMS station data and the Q100 determined from previously published FIS for 97 sites. The flood-frequency discharges computed from AMS station data were consistently larger than the flood discharges from the FIS; mean percentage difference between the two data sets ranged from 14 percent for the Q100 to 20 percent for the Q50. The results of the Mann-Kendall test showed that 8 stations exhibited a positive trend (i.e., increasing annual maximum peaks over time) over their respective periods of record at the 95-percent confidence level, and an additional 7 stations indicated a positive trend, for a total of 15 stations, at a confidence level of greater than or equal to 90 percent. The Q2, Q5, Q10, Q50, and Q100 determined from AMS and PDS data for each station were compared by percentage. The flood magnitudes for the 2-year return period were 16 percent higher when partial-duration peaks were incorporated into the analyses, as opposed to using only the annual maximum peaks. The discharges then tended to converge around the 5-year return period, with a mean collective difference of only 1 percent. At the 10-, 50-, and 100-year return periods, the flood magnitudes based on annual maximum peaks were, on average, 6 percent higher compared to corresponding flood magnitudes based on partial-duration peaks. Possible effects on flood peaks from flood-control reservoirs and urban development within the basin also were examined. Annual maximum peak-flow data from four stations were divided into pre- and post-regulation periods. Comparisons were made between the Q100 determined from AMS station data for the periods of record pre- and post-regulation. Two stations showed a nearly 60- and 20-percent reduction in the 100-year discharges; the other two stations showed negligible differences in discharges. Three stations within urban basins were compared to 38 stations

Pennsylvania

Statistical methods in water resources

Preface This book began as class notes for a course we teach on applied statistical methods to hydrologists of the Water Resources Division, U. S. Geological Survey (USGS). It reflects our attempts to teach statistical methods which are appropriate for analysis of water resources data. As interest in this course has grown outside of the USGS, incentive grew to develop the material into a textbook. The topics covered are those we feel are of greatest usefulness to the practicing water resources scientist. Yet all topics can be directly applied to many other types of environmental data. This book is not a stand-alone text on statistics, or a text on statistical hydrology. For example, in addition to this material we use a textbook on introductory statistics in the USGS training course. As a consequence, discussions of topics such as probability theory required in a general statistics textbook will not be found here. Derivations of most equations are not presented. Important tables included in all general statistics texts, such as quantiles of the normal distribution, are not found here. Neither are details of how statistical distributions should be fitted to flood data -- these are adequately covered in numerous books on statistical hydrology. We have instead chosen to emphasize topics not always found in introductory statistics textbooks, and often not adequately covered in statistical textbooks for scientists and engineers. Tables included here, for example, are those found more often in books on nonparametric statistics than in books likely to have been used in college courses for engineers. This book points the environmental and water resources scientist to robust and nonparametric statistics, and to exploratory data analysis. We believe that the characteristics of environmental (and perhaps most other 'real') data drive analysis methods towards use of robust and nonparametric methods. Exercises are included at the end of chapters. In our course, students compute each type of analysis (t-test, regression, etc.) the first time by hand. We choose the smaller, simpler examples for hand computation. In this way the mechanics of the process are fully understood, and computer software is seen as less mysterious. We wish to acknowledge and thank several other scientists at the U. S. Geological Survey for contributing ideas to this book. In particular, we thank those who have served as the other instructors at the USGS training course. Ed Gilroy has critiqued and improved much of the material found in this book. Tim Cohn has contributed in several areas, particularly to the sections on bias correction in regression, and methods for data below the reporting limit. Richard Alexander has added to the trend analysis chapter, and Charles Crawford has contributed ideas for regression and ANOVA. Their work has undoubtedly made its way into this book without adequate recognition. Professor Ken Potter (University of Wisconsin) and Dr. Gary Tasker (USGS) reviewed the manuscript, spending long hours with no reward except the knowledge that they have improved the work of others. For that we are very grateful. We also thank Madeline Sabin, who carefully typed original drafts of the class notes on which the book is based. As always, the responsibility for all errors and slanted thinking are ours alone.

Techniques of Water-Resources Investigations

StreamStats—A quarter century of delivering web-based geospatial and hydrologic information to the public, and lessons learned

StreamStats is a U.S. Geological Survey (USGS) web application that provides streamflow statistics, such as the 1-percent annual exceedance probability peak flow, the mean flow, and the 7-day, 10-year low flow, to the public through a map-based user interface. These statistics are used in many ways, such as in the design of roads, bridges, and other structures; in delineation of floodplains for land-use zoning and setting of insurance rates; for regulatory purposes, such as the permitting of wastewater discharges; and for hydrologic and climate change studies. StreamStats was first developed for Massachusetts and released in 2001. The application provided users with the ability to obtain streamflow statistics computed from data collected at USGS streamgages and to obtain estimates of streamflow statistics for user-selected ungaged sites. Massachusetts StreamStats used geographic information system software and digital mapping to compute drainage-basin characteristics, which were then used in statistical models to estimate streamflow statistics for the user-selected sites. The statistical models were in the form of equations that were developed through a process known as regression analysis. StreamStats was the first known web application with the ability to do interactive geoprocessing. The utility of Massachusetts StreamStats was instantly apparent, leading the USGS to develop a version of StreamStats that could be implemented nationally. USGS State offices normally were required to develop custom regression equations and prepare local digital mapping data needed for implementing StreamStats for their States. Funding needed to complete this work usually was provided through cooperative agreements between the USGS and State agencies. In 2004, Idaho became the first to be released in the national version of StreamStats. By 2023, 44 States were fully implemented and six were undergoing implementation. StreamStats has undergone many modifications over the years to keep up with changes to the underlying software and to add functionality. Customized functionality and separate linked StreamStats applications were developed for several States. Meeting the high demand for additions and improvements to StreamStats while also adhering to budgetary constraints has, at times, been challenging. The StreamStats development team has identified numerous additional improvements that could be made to provide better performance and more functionality. The lessons learned from the experience of building and operating StreamStats for nearly 25 years could be relevant to others interested in pursuing efforts of a similar scale.

Circular