Search USGSSearch

SEARCH · Search USGS

Results for “Computational Statistics and Data Analysis”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Predicting paleoclimate from compositional data using multivariate Gaussian process inverse prediction

Multivariate compositional count data arise in many applications including ecology, microbiology, genetics and paleoclimate. A frequent question in the analysis of multivariate compositional count data is what underlying values of a covariate(s) give rise to the observed composition. Learning the relationship between covariates and the compositional count allows for inverse prediction of unobserved covariates given compositional count observations. Gaussian processes provide a flexible framework for modeling functional responses with respect to a covariate without assuming a functional form. Many scientific disciplines use Gaussian process approximations to improve prediction and make inference on latent processes and parameters. When prediction is desired on unobserved covariates given realizations of the response variable, this is called inverse prediction. Because inverse prediction is often mathematically and computationally challenging, predicting unobserved covariates often requires fitting models that are different from the hypothesized generative model. We present a novel computational framework that allows for efficient inverse prediction using a Gaussian process approximation to generative models. Our framework enables scientific learning about how the latent processes co-vary with respect to covariates while simultaneously providing predictions of missing covariates. The proposed framework is capable of efficiently exploring the high dimensional, multi-modal latent spaces that arise in the inverse problem. To demonstrate flexibility, we apply our method in a generalized linear model framework to predict latent climate states given multivariate count data. Based on cross-validation, our model has predictive skill competitive with current methods while simultaneously providing formal, statistical inference on the underlying community dynamics of the biological system previously not available.

Annals of Applied Statistics

Statistical models for estimating daily streamflow in Michigan

Statistical models for estimating daily streamflow were analyzed for 25 pairs of streamflow-gaging stations in Michigan. Stations were paired by randomly choosing a station operated in 1989 at which 10 or more years of continuous flow data had been collected and at which flow is virtually unregulated; a nearby station was chosen where flow characteristics are similar. Streamflow data from the 25 randomly selected stations were used as the response variables; streamflow data at the nearby stations were used to generate a set of explanatory variables. Ordinary-least squares regression (OLSR) equations, autoregressive integrated moving-average (ARIMA) equations, and transfer function-noise (TFN) equations were developed to estimate the log transform of flow for the 25 randomly selected stations. The precision of each type of equation was evaluated on the basis of the standard deviation of the estimation errors. OLSR equations produce one set of estimation errors; ARIMA and TFN models each produce l sets of estimation errors corresponding to the forecast lead. The lead- l forecast is the estimate of flow l days ahead of the most recent streamflow used as a response variable in the estimation. In this analysis, the standard deviation of lead l ARIMA and TFN forecast errors were generally lower than the standard deviation of OLSR errors for l < 2 days and l < 9 days, respectively. Composite estimates were computed as a weighted average of forecasts based on TFN equations and backcasts (forecasts of the reverse-ordered series) based on ARIMA equations. The standard deviation of composite errors varied throughout the length of the estimation interval and generally was at maximum near the center of the interval. For comparison with OLSR errors, the mean standard deviation of composite errors were computed for intervals of length 1 to 40 days. The mean standard deviation of length- l composite errors were generally less than the standard deviation of the OLSR errors for l < 32 days. In addition, the composite estimates ensure a gradual transition between periods of estimated and measured flows. Model performance among stations of differing model error magnitudes were compared by computing ratios of the mean standard deviation of the length l composite errors to the standard deviation of OLSR errors. The mean error ratio for the set of 25 selected stations was less than 1 for intervals l < 32 days. Considering the frequency characteristics of the length of intervals of estimated record in Michigan, the effective mean error ratio for intervals < 30 days was 0.52. Thus, for intervals of estimation of 1 month or less, the error of the composite estimate is substantially lower than error of the OLSR estimate.

Michigan

Effect of land-applied biosolids on surface-water nutrient yields and groundwater quality in Orange County, North Carolina

Land application of municipal wastewater biosolids is the most common method of biosolids management used in North Carolina and the United States. Biosolids have characteristics that may be beneficial to soil and plants. Land application can take advantage of these beneficial qualities, whereas disposal in landfills or incineration poses no beneficial use of the waste. Some independent studies and laboratory analysis, however, have shown that land-applied biosolids can pose a threat to human health and surface-water and groundwater quality. The effect of municipal biosolids applied to agriculture fields is largely unknown in relation to the delivery of nutrients, bacteria, metals, and contaminants of emerging concern to surface-water and groundwater resources. Therefore, the North Carolina Department of Environment and Natural Resources (NCDENR) collaborated with the U.S. Geological Survey (USGS) through the 319 Nonpoint Source Program to better understand the transport of nutrients and bacteria from biosolids application fields to groundwater and surface water and to provide a scientific basis for evaluating the effectiveness of the current regulations. The USGS conducted a paired agricultural watershed study in the Collins Creek and Cane Creek Reservoir watersheds in Orange County, North Carolina. Field activities were conducted from March 2011 through May 2013 at two field study sites, including biosolids field application sites owned by Orange County Water and Sewer Authority (OWASA) in the Collins Creek watershed and a background study site in the Cane Creek watershed that has no fields receiving biosolids applications. Samples of biosolids source material and soil were collected from the land-application fields for laboratory analyses. Soil samples were also collected from a background agricultural field in the Cane Creek watershed that has never received land-applied municipal biosolids. Shallow groundwater samples were collected quarterly from new monitoring wells installed by NCDENR along the edge of the biosolids land-application fields and a background agricultural field for laboratory analyses. Two surface-water monitoring sites were established on Collins Creek to compute continuous streamflow and collect discrete baseflow and stormwater runoff water-quality data upstream and downstream from the biosolids land-application fields. Surface water-quality samples were also collected for baseflow and stormwater runoff conditions at an existing USGS streamgage on Cane Creek to monitor water-quality conditions in the background study watershed. The study primarily focused on nutrients and bacteria; however, data for field properties and water-quality constituents, including metals, major ions, and contaminants of emerging concern (household-, industrial-, and agricultural-use compounds, pharmaceutical compounds, hormones, and antibiotics) also were collected and used in the analyses. There were no exceedances of the 10 elements with designated U.S. Environmental Protection Agency (EPA) ceiling concentrations for land-applied biosolids in any of the biosolids samples. Treatment processes and storage techniques used by OWASA are effective in eliminating Escherichia coli and fecal coliform bacteria from biosolids. Copper, molybdenum, total Kjeldahl nitrogen, and total phosphorus were elevated in the soil from biosolids land-application fields relative to the background field. The relative richness of these constituents in the biosolids land-application fields is consistent with biosolids being the source of the elevated concentrations given the relatively high concentrations of these constituents in the biosolids samples that were collected. Shallow groundwater in the transitional zone wells, which were located adjacent to and topographically downgradient from all the biosolids land-application fields, were found to be statistically different and had higher nitrate concentrations (medians greater than 12 milligrams per liter) than all the other wells sampled as part of the study. Surface-water nutrient concentrations and yields, primarily nitrate, were higher at the monitoring site on Collins Creek downstream from the biosolids land-application fields than the other study sites that drained watersheds without biosolids land application. The largest differences in concentrations between sites were measured at baseflow conditions, which indicate that the main cause of these differences, particularly between Cane Creek and the Collins Creek site downstream from the OWASA application fields, is related to nitrate contribution from the shallow groundwater. Contaminants of emerging concern were detected in approximately 40 percent of the laboratory analyses of the biosolids samples and more frequently in soil samples from the biosolids land-application fields (approximately 40 percent of laboratory analyses) relative to the soil samples from the background field (approximately 12 percent of laboratory analyses). However, contaminants of emerging concern detected in the laboratory analysis for this study do not appear to be good indicators of human-waste contaminants derived from land-applied biosolids in groundwater or surface-water because the number of detections and concentrations at the background wells and surface-water monitoring sites are similar to or higher than those at wells and monitoring sites adjacent to or downstream from the biosolids land-application fields. The data, analysis, and conclusions associated with this study can be used by regulatory agencies, resource managers, and wastewater-treatment operators to (1) better understand the quantity and characteristics of nutrients, bacteria, metals, and contaminants of emerging concern that are transported away from biosolids land-application fields to surface water and groundwater under current regulations for the purposes of establishing effective total maximum daily loads (TMDLs) and restoring impaired water resources, (2) assess how well existing regulations protect waters of the State and potentially recommend effective changes to regulations or land-application procedures, and (3) establish a framework for developing guidance on effective techniques for monitoring and regulatory enforcement of permitted biosolids land-application fields.

North Carolina

Age, growth, and production of the yellow perch, Perca flavescens (Mitchill), of Saginaw Bay

Ages were determined and individual growth histories computed from the examination and measurement of scales from 820 yellow perch collected in 1929 and 1930. Calculated lengths greater than 101 millimeters were computed on the assumption (supported by empirical data) that the ratio of body length to scale length is constant. Lengths below 101 millimeters were determined with the aid of an empirical curve of the body-scale relationship of small fish. Yellow perch of age-groups III and IV (in the fourth and fifth years of life) made up the bulk of the collection (78 per cent). Females grew slightly more rapidly than males, but members of both sexes attained the legal length of 8 1/2 inches during the fourth year of life, just as they were entering on the period of most rapid growth in weight. The greatest growth in weight of both sexes occurred in the sixth year of life. In the combined samples of the two years the females exceeded the males in abundance in the ratio, 296:100. The weight of the Saginaw Bay yellow perch was found to increase as the 3.117 power of the length. The relative length of the tail decreased with increase in the length of the fish. The Saginaw Bay yellow perch is now far less abundant than it was in the early years of the fishery. The average annual production of 548,000 pounds over the period, 1917-1938, was only 28 per cent of the earlier (1891-1916) "normal" annual production of 1,961,000 pounds. A detailed analysis of statistical data available for more recent years made possible a description of annual fluctuations in the abundance and production of yellow perch and in the intensity of the yellow perch fishery in Saginaw Bay over the period, 1929-1938.

Transactions of the American Fisheries Society

A spatially referenced regression model (SPARROW) for suspended sediment in streams of the Conterminous U.S.

Suspended sediment has long been recognized as an important contaminant affecting water resources. Besides its direct role in determining water clarity, bridge scour and reservoir storage, sediment serves as a vehicle for the transport of many binding contaminants, including nutrients, trace metals, semi-volatile organic compounds, a nd numerous pesticides (U.S. Environmental Protection Agency, 2000a). Recent efforts to addr ess water-quality concerns through the Total Maximum Daily Load (TMDL) process have iden tified sediment as the single most prevalent cause of impairment in the Nation’s streams a nd rivers (U.S. Environmental Protection Agency, 2000b). Moreover, sediment has been identified as a medium for the tran sport and sequestration of organic carbon, playing a potentia lly important role in understa nding sources and sinks in the global carbon budget (Stallard, 1998). A comprehensive understanding of sediment fate a nd transport is considered essential to the design and implementation of effective plans for sediment management (Osterkamp and others, 1998, U.S. General Accounting Office, 1990). An exte nsive literature addr essing the problem of quantifying sediment transport has produced a nu mber of methods for estimating its flux (see Cohn, 1995, and Robertson and Roerish, 1999, for us eful surveys). The accuracy of these methods is compromised by uncertainty in the concentration measurements and by the highly episodic nature of sediment movement, particul arly when the methods are applied to smaller basins. However, for annual or decadal flux es timates, the methods are generally reliable if calibrated with extended periods of data (Robertson and Roerish, 1999). A substantial literature also supports the Universal Soil Loss Equation (U SLE) (Soil Conservation Service, 1983), an engineering method for estimating sheet and rill erosion, although the empirical credentials of the USLE have recently been questioned (Tri mble and Crosson, 2000). Conversely, relatively little direct evidence is available concerning the fate of sediment. The common practice of quantifying sediment fate with a sediment deliv ery ratio, estimated from a simple empirical relation with upstream basin area, does not artic ulate the relative importance of individual storage sites within a basin (Wolman, 1977). Rates of sediment deposition in reservoirs and flood plains can be determined from empirical measurement s , but only a limited number of sites have been monitored, and net rates of deposition or loss from other potential sinks and sources is largely unknown (Stallard, 1998). In particular, little is known about how much sediment loss from fields ultimately makes its way to stream channels, and how much sediment is subsequently stored in or lost from th e streambed (Meade and Parker, 1985, Trimble and Crosson, 2000). This paper reports on recent progress made to a ddress empirically the question of sediment fate and transport on a national scale. The model pres ented here is based on the SPAtially Referenced Regression On Watershed attr ibutes (SPARROW) methodology, fi rst used to estimate the distribution of nutrients in str eams and rivers of the United Stat es, and subsequently shown to describe land and stream processes affecting the delivery of nutrients (Smith and others, 1997, Alexander and others, 2000, Preston and Brakeb ill, 1999). The model makes use of numerous spatial datasets, available at the national level, to explain long-term sediment water-quality conditions in major streams and rivers throughou t the United States. Sediment sources are identified using sediment erosion rates from the National Resources I nventory (NRI) (Natural Resources Conservation Service, 2000) and apportioned over the landscape according to 30- meter resolution land-use information from th e National Land Cover Data set (NLCD) (U.S. Geological Survey, 2000a). More than 76,000 reservoirs from the National Inventory of Dams (NID) (U.S. Army Corps of Engin eers, 1996) are identified as pot ential sediment sinks. Other, non-anthropogenic sources and sinks are identified using soil in formation from the State Soil Survey Geographic (STATSGO) data base (Schwarz and Alexander, 1995) and spatial coverages representing surficial rock t ype and vegetative cover. The SPA RROW model empirically relates these diverse spatial datasets to estimates of long-term, mean annual sediment flux computed from concentration and flow measurements co llected over the period 1985 -95 from more than 400 monitoring stations maintained by the Na tional Stream Quality Accounting Network (Alexander and others, 1998), the National Wa ter Quality Assessment Program, and U.S. Geological Survey District offices (Turcios and Gray, in press). Th e calibrated model is used to estimate sediment flux for over 60,000 stream segments included in the River Reach File 1 (RF1) stream network (Alexander and others, 1999). SPARROW uses statis tical methods to calibrate a simple, structural model of riverine water quality, one that imposes mass ba lance in accounting for changes in contaminant flux. As applied here, the mass-balance approach facilitates the interpretation of model results in terms of physical processes affecting sediment transport, and makes possible the estimation of various rates of sediment generation and loss associated with stream channels and features of the landscape. The statistical approach provides a basi s for assessing the error of these inferred rates and of the error in extrapolated estimates of sediment flux made for streams in the RF1 network. An important implication of the holistic modeling approach adopted in this analysis is that estimates of sediment production and loss ar e based on, and therefore consistent with, measurements of in-stream flux. Other ancillary information, such as direct measurements of long-term sediment storage and release from rese rvoirs (Steffen, 1996), is incorporated into the analysis by specifying additional equations expl aining these ancillary variables. The imposition of cross-equation constraints affords this info rmation a statistically consistent weight in explaining in-stream sediment flux. Thus, the me thodology described here represents a general framework for synthesizing a wide spectrum of available information relevant to the understanding of sediment fate and transport.

Conterminous United States

Analysis of sensitivity of simulated recharge to selected parameters for seven watersheds modeled using the precipitation-runoff modeling system

Recharge is a vital component of the ground-water budget and methods for estimating it range from extremely complex to relatively simple. The most commonly used techniques, however, are limited by the scale of application. One method that can be used to estimate ground-water recharge includes process-based models that compute distributed water budgets on a watershed scale. These models should be evaluated to determine which model parameters are the dominant controls in determining ground-water recharge. Seven existing watershed models from different humid regions of the United States were chosen to analyze the sensitivity of simulated recharge to model parameters. Parameter sensitivities were determined using a nonlinear regression computer program to generate a suite of diagnostic statistics. The statistics identify model parameters that have the greatest effect on simulated ground-water recharge and that compare and contrast the hydrologic system responses to those parameters. Simulated recharge in the Lost River and Big Creek watersheds in Washington State was sensitive to small changes in air temperature. The Hamden watershed model in west-central Minnesota was developed to investigate the relations that wetlands and other landscape features have with runoff processes. Excess soil moisture in the Hamden watershed simulation was preferentially routed to wetlands, instead of to the ground-water system, resulting in little sensitivity of any parameters to recharge. Simulated recharge in the North Fork Pheasant Branch watershed, Wisconsin, demonstrated the greatest sensitivity to parameters related to evapotranspiration. Three watersheds were simulated as part of the Model Parameter Estimation Experiment (MOPEX). Parameter sensitivities for the MOPEX watersheds, Amite River, Louisiana and Mississippi, English River, Iowa, and South Branch Potomac River, West Virginia, were similar and most sensitive to small changes in air temperature and a user-defined flow routing parameter. Although the primary objective of this study was to identify, by geographic region, the importance of the parameter value to the simulation of ground-water recharge, the secondary objectives proved valuable for future modeling efforts. The value of a rigorous sensitivity analysis can (1) make the calibration process more efficient, (2) guide additional data collection, (3) identify model limitations, and (4) explain simulated results.

Scientific Investigations Report

Continuous real-time water-quality monitoring and regression analysis to compute constituent concentrations and loads in the North Fork Ninnescah River upstream from Cheney Reservoir, south-central Kansas, 1999–2012

Cheney Reservoir, located in south-central Kansas, is the primary water supply for the city of Wichita. The U.S. Geological Survey has operated a continuous real-time water-quality monitoring station since 1998 on the North Fork Ninnescah River, the main source of inflow to Cheney Reservoir. Continuously measured water-quality physical properties include streamflow, specific conductance, pH, water temperature, dissolved oxygen, and turbidity. Discrete water-quality samples were collected during 1999 through 2009 and analyzed for sediment, nutrients, bacteria, and other water-quality constituents. Regression models were developed to establish relations between discretely sampled constituent concentrations and continuously measured physical properties to compute concentrations of those constituents of interest that are not easily measured in real time because of limitations in sensor technology and fiscal constraints. Regression models were published in 2006 that were based on data collected during 1997 through 2003. This report updates those models using discrete and continuous data collected during January 1999 through December 2009. Models also were developed for four new constituents, including additional nutrient species and indicator bacteria. In addition, a conversion factor of 0.68 was established to convert the Yellow Springs Instruments (YSI) model 6026 turbidity sensor measurements to the newer YSI model 6136 sensor at the North Ninnescah River upstream from Cheney Reservoir site. Newly developed models and 14 years of hourly continuously measured data were used to calculate selected constituent concentrations and loads during January 1999 through December 2012. The water-quality information in this report is important to the city of Wichita because it allows the concentrations of many potential pollutants of interest to Cheney Reservoir, including nutrients and sediment, to be estimated in real time and characterized over conditions and time scales that would not be possible otherwise. In general, model forms and the amount of variance explained by the models was similar between the original and updated models. The amount of variance explained by the updated models changed by 10 percent or less relative to the original models. Total nitrogen, nitrate, organic nitrogen, E. coli bacteria, and total organic carbon models were newly developed for this report. Additional data collection over a wider range of hydrological conditions facilitated the development of these models. The nitrate model is particularly important because it allows for comparison to Cheney Reservoir Task Force goals. Mean hourly computed total suspended solids concentration during 1999 through 2012 was 54 milligrams per liter (mg/L). The total suspended solids load during 1999 through 2012 was 174,031 tons. On an average annual basis, the Cheney Reservoir Task Force runoff (550 mg/L) and long-term (100 mg/L) total suspended solids goals were never exceeded, but the base flow goal was exceeded every year during 1999 through 2012. Mean hourly computed nitrate concentration was 1.08 mg/L during 1999 through 2012. The total nitrate load during 1999 through 2012 was 1,361 tons. On an annual average basis, the Cheney Reservoir Task Force runoff (6.60 mg/L) nitrate goal was never exceeded, the long-term goal (1.20 mg/L) was exceeded only in 2012, and the base flow goal of 0.25 mg/L was exceeded every year. Mean nitrate concentrations that were higher during base flow, rather than during runoff conditions, suggest that groundwater sources are the main contributors of nitrate to the North Fork Ninnescah River above Cheney Reservoir. Mean hourly computed phosphorus concentration was 0.14 mg/L during 1999 through 2012. The total phosphorus load during 1999 through 2012 was 328 tons. On an average annual basis, the Cheney Reservoir Task Force runoff goal of 0.40 mg/L for total phosphorus was exceeded in 2002, the year with the largest yearly mean turbidity, and the long-term goal (0.10 mg/L) was exceeded in every year except 2011 and 2012, the years with the smallest mean streamflows. The total phosphorus base flow goal of 0.05 mg/L was exceeded every year. Given that base flow goals for total suspended solids, nitrate, and total phosphorus were exceeded every year despite hydrologic conditions, the established base flow goals are either unattainable or substantially more best management practices will need to be implemented to attain them. On an annual average basis, no discernible patterns were evident in total suspended sediment, nitrate, and total phosphorus concentrations or loads over time, in large part because of hydrologic variability. However, more rigorous statistical analyses are required to evaluate temporal trends. A more rigorous analysis of temporal trends will allow evaluation of watershed investments in best management practices.

Kansas

Estimating the Magnitude and Frequency of Floods in Small Urban Streams in South Carolina, 2001

The magnitude and frequency of floods at 20 streamflowgaging stations on small, unregulated urban streams in or near South Carolina were estimated by fitting the measured wateryear peak flows to a log-Pearson Type-III distribution. The period of record (through September 30, 2001) for the measured water-year peak flows ranged from 11 to 25 years with a mean and median length of 16 years. The drainage areas of the streamflow-gaging stations ranged from 0.18 to 41 square miles. Based on the flood-frequency estimates from the 20 streamflow-gaging stations (13 in South Carolina; 4 in North Carolina; and 3 in Georgia), generalized least-squares regression was used to develop regional regression equations. These equations can be used to estimate the 2-, 5-, 10-, 25-, 50-, 100-, 200-, and 500-year recurrence-interval flows for small urban streams in the Piedmont, upper Coastal Plain, and lower Coastal Plain physiographic provinces of South Carolina. The most significant explanatory variables from this analysis were mainchannel length, percent impervious area, and basin development factor. Mean standard errors of prediction for the regression equations ranged from -25 to 33 percent for the 10-year recurrence-interval flows and from -35 to 54 percent for the 100-year recurrence-interval flows. The U.S. Geological Survey has developed a Geographic Information System application called StreamStats that makes the process of computing streamflow statistics at ungaged sites faster and more consistent than manual methods. This application was developed in the Massachusetts District and ongoing work is being done in other districts to develop a similar application using streamflow statistics relative to those respective States. Considering the future possibility of implementing StreamStats in South Carolina, an alternative set of regional regression equations was developed using only main channel length and impervious area. This was done because no digital coverages are currently available for basin development factor and, therefore, it could not be included in the StreamStats application. The average mean standard error of prediction for the alternative equations was 2 to 5 percent larger than the standard errors for the equations that contained basin development factor. For the urban streamflow-gaging stations in South Carolina, measured water-year peak flows were compared with those from an earlier urban flood-frequency investigation. The peak flows from the earlier investigation were computed using a rainfall-runoff model. At many of the sites, graphical comparisons indicated that the variance of the measured data was much less than the variance of the simulated data. Several statistical tests were applied to compare the variances and the means of the measured and simulated data for each site. The results indicated that the variances were significantly different for 11 of the 13 South Carolina streamflow-gaging stations. For one streamflow-gaging station, the test for normality, which is one of the assumptions of the data when comparing variances, indicated that neither the measured data nor the simulated data were distributed normally; therefore, the test for differences in the variances was not used for that streamflow-gaging station. Another statistical test was used to test for statistically significant differences in the means of the measured and simulated data. The results indicated that for 5 of the 13 urban streamflowgaging stations in South Carolina there was a statistically significant difference in the means of the two data sets. For comparison purposes and to test the hypothesis that there may have been climatic differences between the period in which the measured peak-flow data were measured and the period for which historic rainfall data were used to compute the simulated peak flows, 16 rural streamflow-gaging stations with long-term records were reviewed using similar techniques as those used for the measured an

South Carolina

Time of travel of water in the Ohio River, Pittsburgh to Cincinnati

This report presents a procedure for estimating the time of travel of water in the Ohio River from Pittsburgh, Pa., to Cincinnati, Ohio, under various river stage conditions. This information is primarily for use by civil defense officials and by others concerned with problems involving travel time of river water. Tables and charts are presented to show, for a particular stage or discharge at Cincinnati, the average time it would take for water to travel through the entire reach from Pittsburgh, or through successive intermediate segments of the reach. For example, when the discharge at Cincinnati is 200,000 cfs, travel time from Pittsburgh to Cincinnati, a distance of 470 miles, averages about 7 days; and for discharges of more than 200,000 cfs, the travel time decreases very slowly with increasing discharge. When the discharge is 30,000 cfs, travel time is about 28 days; and for discharges of less than 30,000 cfs, the travel time increases very rapidly with decreasing discharge. Estimates of travel time at low discharge are subject to large errors. Statistical analysis of the possible variations of upstream discharge for a given discharge at Cincinnati indicates that the shortest probable travel time from Pittsburgh to Cincinnati ranges from 56 percent of that under average conditions when the discharge at Cincinnati is 15,000 cfs to 93 percent of that under average conditions when the discharge at Cincinnati is 894,000 cfs. A chart showing the time distribution of flow at Cincinnati is presented so that the probable travel time of Ohio River water can be determined for any time of the year. This chart provides information which, when applied to the time-of-travel chart, shows that the most probable travel time of water from Pittsburgh to Cincinnati ranges from 160 hours in February to 1,250 hours in September. Also presented is a flow-duration curve that can be used to predict future discharges and, subsequently, times of travel, for use in long-range planning. The procedure used to compute time of travel is described in sufficient detail to make it usable as a guide for similar studies of other rivers that have deans and pools in the reach being studied. The computations for the time-of-travel charts were made as follows: (a) by dividing the reach between Pittsburgh and Cincinnati into four subreaches with a full-range streamgaging station at or near the ends of each; (b) by computing for each subreach mean velocities corresponding to various discharges at Cincinnati, using data obtained from river survey maps and data available from gaging station operations; (c) by assuming that any mass of contaminated water would travel at a rate equal to that of the mean velocity of the river water.

Ohio River

A Bayesian approach for temporally scaling climate for modeling ecological systems

With climate change becoming more of concern, many ecologists are including climate variables in their system and statistical models. The Standardized Precipitation Evapotranspiration Index (SPEI) is a drought index that has potential advantages in modeling ecological response variables, including a flexible computation of the index over different timescales. However, little development has been made in terms of the choice of timescale for SPEI. We developed a Bayesian modeling approach for estimating the timescale for SPEI and demonstrated its use in modeling wetland hydrologic dynamics in two different eras (i.e., historical [pre-1970] and contemporary [post-2003]). Our goal was to determine whether differences in climate between the two eras could explain changes in the amount of water in wetlands. Our results showed that wetland water surface areas tended to be larger in wetter conditions, but also changed less in response to climate fluctuations in the contemporary era. We also found that the average timescale parameter was greater in the historical period, compared with the contemporary period. We were not able to determine whether this shift in timescale was due to a change in the timing of wet&ndash;dry periods or whether it was due to changes in the way wetlands responded to climate. Our results suggest that perhaps some interaction between climate and hydrologic response may be at work, and further analysis is needed to determine which has a stronger influence. Despite this, we suggest that our modeling approach enabled us to estimate the relevant timescale for SPEI and make inferences from those estimates. Likewise, our approach provides a mechanism for using prior information with future data to assess whether these patterns may continue over time. We suggest that ecologists consider using temporally scalable climate indices in conjunction with Bayesian analysis for assessing the role of climate in ecological systems.

Ecology and Evolution

Virginia Bridge Scour Pilot Study—Hydrological Tools

Hydrologic and geophysical components interact to produce streambed scour. This study investigates methods for improving the utility of estimates of hydrologic flow in streams and rivers used when evaluating potential pier scour over the design-life of highway bridges in Virginia. Recent studies of streambed composition identify potential bridge design cost savings when attributes of cohesive soil and weathered rock unique to certain streambeds are considered within the bridge planning design. To achieve potential cost savings, however, attributes and effects of scour forces caused by water movement across the streambed surface must be accurately described and estimated. This study explores the potential for improving estimates of the hydrologic component, namely hydrologic flow, afforded by empirically based deterministic, probabilistic, and statistical modeling of flows using streamgage data from 10 selected sites in Virginia. Methods are described and tools are provided that may assist with estimating hydrological components of flow duration and potential cumulative stream power for bridge designs in specific settings, and calculation of comprehensive projections of anticipated individual bridge pier scour rates. Examples of hydrologic properties needed to determine the rates of streambed scour are described for sites spanning a range of basin sizes and locations in Virginia. Deterministic, probabilistic, and statistical modeling methods are demonstrated for estimating hydrological components of streambed scour over a bridge design lifespan. Eight tools provide examples of streamflow analysis using daily and instantaneous streamflow data collected at 10 study sites in Virginia. Tool 1 provides a generalized system dynamics model of streamflow and sediment motion that may be used to estimate hydrologic flow over time. Tool 2 illustrates at-a-station hydraulic geometry using methods pioneered by Leopold and others. Tool 3 provides a system dynamics model developed to test the use of Monte-Carlo sampling of instantaneous streamflow measurements to augment and increase precision of site-specific period-of-record daily-flow values useful for driving stream-power and streambed scour estimates. Tool 4 integrates deterministic modeling, maximum likelihood logistic regression, and Monte-Carlo sampling to identify probable hydrologic flows. Tool 5 provides instantaneous flow hydrologic envelope profiles, using measured instantaneous flow data integrated with measured daily-flow value data. Tool 6 provides precise estimates of hydrologic flow over entire data time-series suitable for driving scour simulation models. Tool 7 provides a threshold of flow and probability of time-under-load interactive calculator that allows selection of a desired bridge design lifespan, ranging from 1 to 250 years, and identification of a flow interval of interest. Tool 8 provides a flow-random sampling interactive tool, developed to facilitate easy access to large datasets of randomly sampled flow data measurements from unique locations for purposes of computing and testing future models of bridge pier scour.

Virginia

Generalized estimates from streamflow data of annual and seasonal ground-water-recharge rates for drainage basins in New Hampshire

This report presents regression equations to estimate generalized annual and seasonal ground-water-recharge rates in drainage basins in New Hampshire. The ultimate source of water for a ground-water withdrawal is aquifer recharge from a combination of precipitation on the aquifer, ground-water flow from upland basin areas, and infiltration from streambeds to the aquifer. An assessment of ground-water availability in a basin requires that recharge rates be estimated under `normal' conditions and under assumed drought conditions. Recharge equations were developed by analyzing streamflow, basin characteristics, and precipitation at 55 unregulated continuous record stream-gaging stations in New Hampshire and in adjacent states. In the initial step, streamflow records were analyzed to estimate a series of annual and seasonal ground-water-recharge components of streamflow in each drainage basin evaluated in this study. Regression equations were then developed relating the series of annual and seasonal ground-water-recharge values to the corresponding series of annual and seasonal precipitation values as determined at the centroid of each drainage basin. This resulted in one equation for each of the 55 basins for each of the four seasonal periods and the annual period, or a total of 275 regression equations. Average annual and seasonal precipitation data for 1961-90 were then used to compute a set of normalized ground-water-recharge values that reflected the long-term average annual and seasonal variations (normalized) and mean recharge characteristics of each drainage basin. Ordinary-least-squares regression was applied in the process of selecting 10 out of 93 possible basin and climatic characteristics for further testing in the development of the equations for computing the generalized estimate of annual and seasonal ground-water recharge based on the set of normalized recharge values. Generalized-least-squares regression was used for the final parameter estimation and error evaluation. The following basin and climatic characteristics were found to be statistically significant predictors for at least one of the dependent variables: average annual, summer, and spring precipitation as determined at U.S. Geological Survey stream-gaging stations; average annual basin-centroid precipitation; average mean annual basin temperature; average minimum winter basin temperature; percent coniferous forest in a basin; percent mixed coniferous and deciduous forest in a basin; average fall basin-centroid precipitation; and average annual snowcover. These 10 basin and climatic characteristics were selected because they were statistically significant based on several statistical parameters that evaluated which combination of characteristics contributed the most to the predictive accuracy of the regression-equation models. A geographic information system is required to measure the values of the predictor variables for the equations developed in the study. The average annual normalized ground-water recharge was 21.0 in. This value was determined by generalized-least-squares (GLS) regression analysis for all of the basins used in the normalized ground-water recharge analysis for precipitation from 1961-90. The average winter (January 1-March 15) ground-water recharge was 4.3 in., average spring (March 16-May 31) ground-water recharge was 9.0 in., average summer (June 1-October 31) ground-water recharge was 4.0 in., and average fall (November 1-December 31) ground-water recharge was 3.6 in. Normalized ground-water recharge ranged annually from 12.3 to 31.8 in., for winter from 2.30 to 7.82 in., for spring from 5.16 to 13.7 in., for summer from 1.45 to 10.2 in., and for fall from 2.21 to 6.06 in.

New Hampshire

Effect of uncertainty of discharge data on uncertainty of discharge simulation for the Lake Michigan Diversion, northeastern Illinois and northwestern Indiana

Simulation models of watershed hydrology (also referred to as “rainfall-runoff models”) are calibrated to the best available streamflow data, which are typically published discharge time series at the outlet of the watershed. Even after calibration, the model generally cannot replicate the published discharges because of simplifications of the physical system embedded in the model structure and uncertainties of the input data and of the estimated model parameters, which, although optimized for the given calibration data, remain uncertain. The input data errors are caused by uncertainties in the forcing data, such as precipitation and other climatological data, and in the published discharges used for calibration. In the numerical algorithms used for calibration, the published discharges are often assumed to be without error, but they are themselves uncertain, typically having been computed using ratings, which are models fitted to uncertain discharge measurements. In this study, uncertainty of published daily discharge data and how the discharge uncertainty is transmitted to the parameter values of the Hydrological Simulation Program–FORTRAN (HSPF) rainfall-runoff model and to the simulated discharge at both calibration and prediction locations were investigated for the Lake Michigan diversion in northeastern Illinois and northwestern Indiana. The HSPF model used in this study is used by the U.S. Army Corps of Engineers as part of quantifying the diversion of water from Lake Michigan by the State of Illinois. In this study, the model is calibrated jointly at two watersheds in the study area; the resulting model is considered the base model in this study. Seven other gaged watersheds in the study area are used for testing predictive simulations. A Bayesian rating curve estimation (BaRatin) approach, the BaRatin stage-period-discharge (SPD) method, was used to estimate the uncertainty of the published discharge from the calibration watersheds. To characterize the effect of the discharge uncertainty on parameter values, the HSPF model parameters were recalibrated to 17 nonrandomly selected pairs of discharge series from the BaRatin SPD analysis. To provide an indicator of the effect of parameter uncertainty to compare to the effect of discharge uncertainty, 1,000 parameter sets also were randomly generated from the estimated parameter covariance matrix of the base model. The recalibrated and random parameter sets were then used in HSPF simulations of discharge at the two calibration watersheds and at the seven prediction watersheds. Selected discharge summary statistics—the period-of-study (POS, water years 1997 to 2015) mean discharge, selected flow-duration curve (FDC) quantiles, and water year mean discharges—are used to characterize the variability between simulated and published discharge. A normalized variability index ( V N ) is used as a measure of the uncertainty of flow statistics arising from the uncertainty of the sources considered in this study. When this index is at least 1, the variability of the simulations is large enough to explain the median error between simulated and published values, although offsetting errors from other sources are also likely. When the index is appreciably less than 1, the variability of the simulations is clearly insufficient to explain the median error between simulated and published values. At the two calibration watersheds and for results of the two simulation sets considered together, the V N values ranged from 0.2 to 0.8 for POS mean discharge, from 0.3 to 0.6 in the median for a set of FDC quantiles, and from 0.1 to 0.2 in the median for water year mean discharges. These values indicate that substantial uncertainty remains unexplained. Even though two watersheds were used in calibration, that calibration was highly constrained because it was applied to the watersheds simultaneously and was subject to parameter regularization that constrained the adjustment of the parameters from their initial values. These constraints were applied to avoid overfitting to the calibration watersheds and thus to increase the likelihood that the resulting parameters would give accurate results at watersheds not used in the calibration, but they created a parameter transfer error in the calibration watershed results shown by the balancing of errors between the two watersheds. Additional remaining error sources include model structural error and meteorological forcing error to the degree that the calibration was unable to adjust the parameters to account for these errors. At the prediction watersheds, the corresponding V N values were almost always substantially lower than those values at the calibration watersheds. This result is expected because the prediction watersheds have additional uncertainty, including parameter transfer error. The work described in this report provides preliminary estimates of a limited range of sources of error in predicted discharge uncertainty. Future work would be beneficial to obtain a better statistical characterization of the effect of the uncertainty of calibration discharge series and to address additional sources of uncertainty, such as from precipitation input data used in calibration and prediction and from structural (model) errors.

Illinois, Indiana

User's manual for the National Water-Quality Assessment Program Invertebrate Data Analysis System (IDAS) software, version 5

The Invertebrate Data Analysis System (IDAS) software was developed to provide an accurate, consistent, and efficient mechanism for analyzing invertebrate data collected as part of the U.S. Geological Survey National Water-Quality Assessment (NAWQA) Program. The IDAS software is a stand-alone program for personal computers that run Microsoft Windows(Registered). It allows users to read data downloaded from the NAWQA Program Biological Transactional Database (Bio-TDB) or to import data from other sources either as Microsoft Excel(Registered) or Microsoft Access(Registered) files. The program consists of five modules: Edit Data, Data Preparation, Calculate Community Metrics, Calculate Diversities and Similarities, and Data Export. The Edit Data module allows the user to subset data on the basis of taxonomy or sample type, extract a random subsample of data, combine or delete data, summarize distributions, resolve ambiguous taxa (see glossary) and conditional/provisional taxa, import non-NAWQA data, and maintain and create files of invertebrate attributes that are used in the calculation of invertebrate metrics. The Data Preparation module allows the user to select the type(s) of sample(s) to process, calculate densities, delete taxa on the basis of laboratory processing notes, delete pupae or terrestrial adults, combine lifestages or keep them separate, select a lowest taxonomic level for analysis, delete rare taxa on the basis of the number of sites where a taxon occurs and (or) the abundance of a taxon in a sample, and resolve taxonomic ambiguities by one of four methods. The Calculate Community Metrics module allows the user to calculate 184 community metrics, including metrics based on organism tolerances, functional feeding groups, and behavior. The Calculate Diversities and Similarities module allows the user to calculate nine diversity and eight similarity indices. The Data Export module allows the user to export data to other software packages (CANOCO, Primer, PC-ORD, MVSP) and produce tables of community data that can be imported into spreadsheet, database, graphics, statistics, and word-processing programs. The IDAS program facilitates the documentation of analyses by keeping a log of the data that are processed, the files that are generated, and the program settings used to process the data. Though the IDAS program was developed to process NAWQA Program invertebrate data downloaded from Bio-TDB, the Edit Data module includes tools that can be used to convert non-NAWQA data into Bio-TDB format. Consequently, the data manipulation, analysis, and export procedures provided by the IDAS program can be used to process data generated outside of the NAWQA Program.

Techniques and Methods

Magnitude and frequency of floods on Kauaʻi, Oʻahu, Molokaʻi, Maui, and Hawaiʻi, State of Hawaiʻi, based on data through water year 2020

Accurate estimates of flood magnitude and frequency are needed to (1) optimize the design and location of infrastructure, including dams, culverts, bridges, industrial buildings, and highways, and (2) inform flood-zoning and flood-insurance studies. The U.S. Geological Survey (USGS), in cooperation with the State of Hawaiʻi Department of Transportation, estimated flood magnitudes for the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities (AEP) for unregulated streamgages in Kauaʻi, Oʻahu, Molokaʻi, Maui, and Hawaiʻi, State of Hawaiʻi, using data through water year 2020. Regression equations were developed to estimate flood magnitude and associated frequency at ungaged streams. This study improves upon a previous USGS flood-frequency report (Oki and others, 2010) by including more peak-flow data, implementing new statistical methods in flood-frequency analysis, and using updated techniques to estimate the regional-skewness coefficient (regional skew). Flood magnitude and frequency at 238 streamgages were estimated—following national guidelines established in Bulletin 17C (England and others, 2019)—by fitting annual peak-flow data to the Log-Pearson Type III distribution using the expected moments algorithm and the PeakFQ flood-frequency software. Potentially influential low outliers in the data were identified and removed using the Multiple Grubbs-Beck Test. An updated regional skew for Hawaiʻi was estimated using the Bayesian weighted least squares/Bayesian generalized least squares method. The updated regional skew employs a constant model for the five islands in the study area and has a value of −0.157 (mean square error of 0.212). Multiple linear regression techniques were used to develop regression equations that relate basin and climatic characteristics to peak flows at streamgages. The regression equations can be applied to estimate flood magnitude and frequency at ungaged sites. The study area was split into 10 regions—2 regions per island, generally following a leeward/windward division—containing from 9 to 49 streamgages each. The final regression equations for each region were determined with generalized least-squares analysis using the USGS weighted-multiple-linear regression (WREG) program. The standard error of prediction at the 1-percent AEP for the regression equations ranged from 18 to 164 percent; the pseudo coefficient of determination (pseudo-R2) at the 1-percent AEP ranged from 46 to 100 percent. The regression equations performed well for all regions except leeward Molokaʻi and southern Island of Hawaiʻi; for all other regions, the pseudo-R2 values ranged from about 75 to 100 percent. Compared to the regression equations developed by Oki and others (2010), the regression equations in this study generally showed modest improvements, although the magnitude of differences varied for each region. Peak-flow estimates at the 238 streamgages included in this study are improved by weighting the at-site statistics computed with PeakFQ and the predicted flows based on the regression equations. Results of this study—including the final peak-flow estimates at streamgages and the regional regression equations—are implemented in the USGS StreamStats web application (U.S. Geological Survey, 2023, StreamStats: https://streamstats.usgs.gov/ss/ ). StreamStats provides a consistent approach for obtaining peak-flow estimates at streamgages and for applying the regional regression equations for estimating peak flows at ungaged locations.

Hawaii

Methods for estimating flow-duration and annual mean-flow statistics for ungaged streams in Oklahoma

Flow statistics can be used to provide decision makers with surface-water information needed for activities such as water-supply permitting, flow regulation, and other water rights issues. Flow statistics could be needed at any location along a stream. Most often, streamflow statistics are needed at ungaged sites, where no flow data are available to compute the statistics. Methods are presented in this report for estimating flow-duration and annual mean-flow statistics for ungaged streams in Oklahoma. Flow statistics included the (1) annual (period of record), (2) seasonal (summer-autumn and winter-spring), and (3) 12 monthly duration statistics, including the 20th, 50th, 80th, 90th, and 95th percentile flow exceedances, and the annual mean-flow (mean of daily flows for the period of record). Flow statistics were calculated from daily streamflow information collected from 235 streamflow-gaging stations throughout Oklahoma and areas in adjacent states. A drainage-area ratio method is the preferred method for estimating flow statistics at an ungaged location that is on a stream near a gage. The method generally is reliable only if the drainage-area ratio of the two sites is between 0.5 and 1.5. Regression equations that relate flow statistics to drainage-basin characteristics were developed for the purpose of estimating selected flow-duration and annual mean-flow statistics for ungaged streams that are not near gaging stations on the same stream. Regression equations were developed from flow statistics and drainage-basin characteristics for 113 unregulated gaging stations. Separate regression equations were developed by using U.S. Geological Survey streamflow-gaging stations in regions with similar drainage-basin characteristics. These equations can increase the accuracy of regression equations used for estimating flow-duration and annual mean-flow statistics at ungaged stream locations in Oklahoma. Streamflow-gaging stations were grouped by selected drainage-basin characteristics by using a k-means cluster analysis. Three regions were identified for Oklahoma on the basis of the clustering of gaging stations and a manual delineation of distinguishable hydrologic and geologic boundaries: Region 1 (western Oklahoma excluding the Oklahoma and Texas Panhandles), Region 2 (north- and south-central Oklahoma), and Region 3 (eastern and central Oklahoma). A total of 228 regression equations (225 flow-duration regressions and three annual mean-flow regressions) were developed using ordinary least-squares and left-censored (Tobit) multiple-regression techniques. These equations can be used to estimate 75 flow-duration statistics and annual mean-flow for ungaged streams in the three regions. Drainage-basin characteristics that were statistically significant independent variables in the regression analyses were (1) contributing drainage area; (2) station elevation; (3) mean drainage-basin elevation; (4) channel slope; (5) percentage of forested canopy; (6) mean drainage-basin hillslope; (7) soil permeability; and (8) mean annual, seasonal, and monthly precipitation. The accuracy of flow-duration regression equations generally decreased from high-flow exceedance (low-exceedance probability) to low-flow exceedance (high-exceedance probability) . This decrease may have happened because a greater uncertainty exists for low-flow estimates and low-flow is largely affected by localized geology that was not quantified by the drainage-basin characteristics selected. The standard errors of estimate of regression equations for Region 1 (western Oklahoma) were substantially larger than those standard errors for other regions, especially for low-flow exceedances. These errors may be a result of greater variability in low flow because of increased irrigation activities in this region. Regression equations may not be reliable for sites where the drainage-basin characteristics are outside the range of values of independent vari

Scientific Investigations Report

Relationship of geological and geothermal field properties: Midcontinent area, USA, an example

Quantitative approaches to data analysis in the last decade have become important in basin modeling and mineral-resource estimation. The interrelation of geological, geophysical, geochemical, and geohydrological variables is important in adjusting a model to a real-world situation. Revealing the interdependences of variables can contribute in understanding the processes interacting in sedimentary basins. It is reasonably simple to compare spatial data of the same type but more difficult if different properties are involved. Statistical techniques, such as cluster analysis or principal components analysis, or some algebraic approaches can be used to ascertain the relations of standardized spatial data. In this example, structural configuration on five different stratigraphic horizons, one total sediment thickness map, and four maps of geothermal data were copared. As expected, the structural maps are highly related because all had undergone about the same deformation with differing degrees of intensity. The temperature gradients derived (1) from shallow borehole logging measurements under equilibrium conditions with the surrounding rock, and (2) from non-equilibrium bottom-hole temperatures (BHT) from deeper depths are mainly independent of each other. This was expected and confirmed also for the two temperature maps at 1000 ft which were constructed using both types of gradient values. Thus, it is evident that the use of a 2-point (BHT and surface temperature) straightline calculation of a mean temperature gradient gives different information about the geothermal regime than using gradients from temperatures logged under equilibrium conditions. Nevertheless, it is useful to determine to what a degree the larger dataset of nonequilibrium temperatures could reflect quantitative relationships to geologic conditions. Comparing all maps of geothermal information vs. the structural and the sediment thickness maps, it was determined that all correlations are moderately negative or slightly positive. These results are clearly shown by the cluster analysis and the principal components. Considering a close relationship between temperature and thermal conductivity of the sediments as observed for most of the Midcontinent area and relatively homogeneous heat-flow density conditions for the study area these results support the following assumptions: (1) undifferentiated geothermal gradients, computed from temperatures of different depth intervals and differing sediment properties, cannot contribute to an improved understanding of the temperature structure and its controls within the sedimentary cover, and (2) the quantitative approach of revealing such relations needs refined datasets of temperature information valid for the different depth levels or stratigraphic units. ?? 1993 International Association for Mathematical Geology.

Mathematical Geology

Small-mammal density estimation: A field comparison of grid-based vs. web-based density estimators

Statistical models for estimating absolute densities of field populations of animals have been widely used over the last century in both scientific studies and wildlife management programs. To date, two general classes of density estimation models have been developed: models that use data sets from capture–recapture or removal sampling techniques (often derived from trapping grids) from which separate estimates of population size ( NÌ‚ ) and effective sampling area ( AÌ‚ ) are used to calculate density ( DÌ‚ = NÌ‚ / AÌ‚ ); and models applicable to sampling regimes using distance-sampling theory (typically transect lines or trapping webs) to estimate detection functions and densities directly from the distance data. However, few studies have evaluated these respective models for accuracy, precision, and bias on known field populations, and no studies have been conducted that compare the two approaches under controlled field conditions. In this study, we evaluated both classes of density estimators on known densities of enclosed rodent populations. Test data sets ( n = 11) were developed using nine rodent species from capture–recapture live-trapping on both trapping grids and trapping webs in four replicate 4.2-ha enclosures on the Sevilleta National Wildlife Refuge in central New Mexico, USA. Additional “saturation” trapping efforts resulted in an enumeration of the rodent populations in each enclosure, allowing the computation of true densities. Density estimates ( DÌ‚ ) were calculated using program CAPTURE for the grid data sets and program DISTANCE for the web data sets, and these results were compared to the known true densities ( D ) to evaluate each model's relative mean square error, accuracy, precision, and bias. In addition, we evaluated a variety of approaches to each data set's analysis by having a group of independent expert analysts calculate their best density estimates without a priori knowledge of the true densities; this “blind” test allowed us to evaluate the influence of expertise and experience in calculating density estimates in comparison to simply using default values in programs CAPTURE and DISTANCE. While the rodent sample sizes were considerably smaller than the recommended minimum for good model results, we found that several models performed well empirically, including the web-based uniform and half-normal models in program DISTANCE, and the grid-based models M b and M bh in program CAPTURE (with AÌ‚ adjusted by species-specific full mean maximum distance moved (MMDM) values). These models produced accurate DÌ‚ values (with 95% confidence intervals that included the true D values) and exhibited acceptable bias but poor precision. However, in linear regression analyses comparing each model's DÌ‚ values to the true D values over the range of observed test densities, only the web-based uniform model exhibited a regression slope near 1.0; all other models showed substantial slope deviations, indicating biased estimates at higher or lower density values. In addition, the grid-based DÌ‚ analyses using full MMDM values for WÌ‚ area adjustments required a number of theoretical assumptions of uncertain validity, and we therefore viewed their empirical successes with caution. Finally, density estimates from the independent analysts were highly variable, but estimates from web-based approaches had smaller mean square errors and better achieved confidence-interval coverage of D than did grid-based approaches. Our results support the contention that web-based approaches for density estimation of small-mammal populations are both theoretically and empirically superior to grid-based approaches, even when sample size is far less than often recommended. In view of the increasing need for standardized environmental measures for comparisons among ecosystems and through time, analytical models based on distance sampling appear to offer accurate density estimation approaches for research studies involving small-mammal abundances.

Ecological Monographs