Search USGSSearch

SEARCH · Search USGS

Results for “Computational Statistics and Data Analysis”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

An alternative basin characteristic for use in estimating impervious area in urban Missouri basins

A previous regression analysis of flood peaks on urban basins in St. Louis County, Missouri, indicated that the basin characteristics of percentage of impervious area and drainage area were statistically significant for estimating the 2-, 5-, 10-, 25-, 50-. and 100-yr peak discharges at ungaged urban basins. In this statewide regression analysis of the urban basins for Missouri, an alternative basin characteristic called the percentage of developed area was evaluated. A regression analysis of the percentage of developed area (independent variable), resulted in a simple equation for computing percentage of impervious area. The percentage of developed area also was evaluated using flood-frequency data for 23 streamflow gaging stations, and the use of this variable was determined to be valid. Using nationwide data, an urban basin characteristic known as the basin development factor was determined to be valid for inclusion in urban regression equations for estimating flood flows. The basin development factor and the percentage of developed area were compared for use in regression equations to estimate peak flows of streams in Missouri. The equations with the basin development factor produced peak flow estimates with slightly smaller average standard errors of estimate than the equation with the percentage of developed area; however, this study indicates that there was not enough statistical or numerical difference to warrant using the basin development factor instead of the percentage of developed area in Missouri. The selection of a basin characteristic to describe the physical conditions of a drainage basin will depend not only on its contribution to accuracy of regression equations, but also on the ease of determining the characteristics; the percentage of developed area has this advantage. A correlation analysis was made by correlating drainage area to percentage of impervious area, the percentage of developed area, and the basin development factor. The results of the analysis indicate that the three basin characteristics are independent of drainage area and appropriate to use in multiple-regression analysis. (Author 's abstract)

Water-Resources Investigations Report

Fort Collins Science Center: Species and Habitats of Federal Interest

Ecosystem changes directly affect a wide variety of plant and animal species, floral and faunal communities, and groups of species such as amphibians and grassland birds. Appropriate management of public lands plays a crucial role in the conservation and recovery of endangered species and can be a key element in preventing a species from being listed under the Endangered Species Act. The Species and Habitats of Federal Interest Branch of the Fort Collins Science Center (FORT) conducts research on the ecology, habitat requirements, distribution and abundance, population dynamics, and genetics and systematics of many species facing threatened or endangered status or of special concern to resource management agencies. FORT scientists develop reintroduction and restoration techniques, technologies for monitoring populations, and novel methods to analyze data on population trends and habitat requirements. FORT expertise encompasses both traditional and specialized natural resource disciplines within wildlife biology, including population dynamics, animal behavior, plant and community ecology, inventory and monitoring, statistics and computer applications, conservation genetics, stable isotope analysis, and curatorial expertise.

Fact Sheet

Notes on numerical reliability of several statistical analysis programs

This report presents a benchmark analysis of several statistical analysis programs currently in use in the USGS. The benchmark consists of a comparison between the values provided by a statistical analysis program for variables in the reference data set ANASTY and their known or calculated theoretical values. The ANASTY data set is an amendment of the Wilkinson NASTY data set that has been used in the statistical literature to assess the reliability (computational correctness) of calculated analytical results.

Open-File Report

Local sensitivity analysis for inverse problems solved by singular value decomposition

Local sensitivity analysis provides computationally frugal ways to evaluate models commonly used for resource management, risk assessment, and so on. This includes diagnosing inverse model convergence problems caused by parameter insensitivity and(or) parameter interdependence (correlation), understanding what aspects of the model and data contribute to measures of uncertainty, and identifying new data likely to reduce model uncertainty. Here, we consider sensitivity statistics relevant to models in which the process model parameters are transformed using singular value decomposition (SVD) to create SVD parameters for model calibration. The statistics considered include the PEST identifiability statistic, and combined use of the process-model parameter statistics composite scaled sensitivities and parameter correlation coefficients (CSS and PCC). The statistics are complimentary in that the identifiability statistic integrates the effects of parameter sensitivity and interdependence, while CSS and PCC provide individual measures of sensitivity and interdependence. PCC quantifies correlations between pairs or larger sets of parameters; when a set of parameters is intercorrelated, the absolute value of PCC is close to 1.00 for all pairs in the set. The number of singular vectors to include in the calculation of the identifiability statistic is somewhat subjective and influences the statistic. To demonstrate the statistics, we use the USDA’s Root Zone Water Quality Model to simulate nitrogen fate and transport in the unsaturated zone of the Merced River Basin, CA. There are 16 log-transformed process-model parameters, including water content at field capacity (WFC) and bulk density (BD) for each of five soil layers. Calibration data consisted of 1,670 observations comprising soil moisture, soil water tension, aqueous nitrate and bromide concentrations, soil nitrate concentration, and organic matter content. All 16 of the SVD parameters could be estimated by regression based on the range of singular values. Identifiability statistic results varied based on the number of SVD parameters included. Identifiability statistics calculated for four SVD parameters indicate the same three most important process-model parameters as CSS/PCC (WFC1, WFC2, and BD2), but the order differed. Additionally, the identifiability statistic showed that BD1 was almost as dominant as WFC1. The CSS/PCC analysis showed that this results from its high correlation with WCF1 (-0.94), and not its individual sensitivity. Such distinctions, combined with analysis of how high correlations and(or) sensitivities result from the constructed model, can produce important insights into, for example, the use of sensitivity analysis to design monitoring networks. In conclusion, the statistics considered identified similar important parameters. They differ because (1) with CSS/PCC can be more awkward because sensitivity and interdependence are considered separately and (2) identifiability requires consideration of how many SVD parameters to include. A continuing challenge is to understand how these computationally efficient methods compare with computationally demanding global methods like Markov-Chain Monte Carlo given common nonlinear processes and the often even more nonlinear models.

Book

Distribution and density of bird species hazardous to aircraft

Only in the past 5 years has it become feasible to map the relative abundance of North American birds. Two programs presently under way and a third that is in the experimental phase are making possible the up-to-date mapping of abundance as well as distribution. A fourth program that has been used successfully in Europe and on a small scale in parts of North America yields detailed information on breeding distribution. The Breeding Bird Survey, sponsored by the U.S. Bureau of Sport Fisheries and Wildlife and the Canadian Wildlife Service, involves 2,000 randomly distributed roadside counts that are conducted during the height of the breeding season in all U.S. States and Canadian Provinces. Observations of approximately 1.4 million birds per year are entered on magnetic tape and subsequently used both for statistical analysis of population trends and for computer mapping of distribution and abundance. The National Audubon Society's Christmas Bird Count is conducted in about 1,000 circles, each 15 miles (24 km) in diameter, in the latter half of December. Raw data for past years have been published in voluminous reports, but not in a form for ready analysis. Under a contract between the U.S. Air Force and the U. S. Bureau of Sport Fisheries and Wildlife (in cooperation with the National Audubon Society), preliminary maps showing distribution and abundance of selected species that are potential hazards to aircraft are presently being mapped and prepared for publication. The Winter Bird Survey, which is in its fifth season of experimental study in a limited area in Central Maryland, may ultimately replace the Christmas Bird Count source. This Survey consists of a standardized 8-kilometer (5-mile) route covered uniformly once a year during midwinter. Bird Atlas programs, which map distribution but not abundance, are well established in Europe and are gaining interest in America

Book chapter

Methods for estimating selected low-flow frequency and mean annual flow statistics at gaged and ungaged locations on streams in Georgia, North Carolina, and South Carolina

The U.S. Geological Survey, in cooperation with the Georgia Department of Natural Resources (Environmental Protection Division), North Carolina Department of Environmental Quality (Division of Water Resources), North Carolina Department of Public Safety (Office of Recovery and Resiliency), and South Carolina Department of Environmental Services, updated low-flow frequency, mean annual flow, and flow-duration statistics at 843 streamgages in and near Georgia, North Carolina, and South Carolina. The low-flow frequency statistics are annual minimum 1-day average flow with a 10-year recurrence interval (1Q10), annual minimum 7-day average flow for 2- and 10-year recurrence intervals (7Q2 and 7Q10, respectively), and annual minimum 30-day average flow with 2- and 3-year recurrence intervals (30Q2 and 30Q3, respectively). Monthly 1Q10 and 7Q10, and W7Q10 flow statistics for the winter period (November–March) also are presented. By using data from 604 of the streamgages on streams with streamflows that are not substantially affected by regulation or diversion and are not tidally influenced, regional regression equations were developed to predict flow statistics with prediction intervals at ungaged locations on streams with those same criteria. The regional regression analysis included data from 132 streamgages from adjacent States Alabama, Florida, Tennessee, and Virginia. The final regional regression equations include variables such as drainage area, streamflow variability, precipitation, percentage of impervious area, and percentage of the basin in various ecoregions. The low-flow statistics for the streamgages analyzed and the regional regression equations will be integrated into the U.S. Geological Survey StreamStats application ( https://www.usgs.gov/streamstats ) for Georgia, North Carolina, and South Carolina. StreamStats generates basin characteristics needed to compute low-flow frequency statistics for ungaged locations. A trend analysis of annual minimum 7-day average flows was done for 78 streamgages with at least 30 years of continuous record. Trends were evaluated for 30-, 50‑, 70-, and 90-year periods, ending in climate year 2021, and independence and short- and long-term persistence assumptions were considered. For all trend analysis assumptions, most streamgages did not exhibit significant trends in annual minimum 7-day average flows. Trends in annual precipitation and air temperature were similarly evaluated for the period 1895–2021 to assess the variability of climate for Georgia, North Carolina, and South Carolina.

Georgia, North Carolina, South Carolina

Testing small-aperture array analysis on well-located earthquakes, and application to the location of deep tremor

We have here analyzed local and regional earthquakes using array techniques with the double aim of quantifying the errors associated with the estimation of propagation parameters of seismic signals and testing the suitability of a probabilistic location method for the analysis of nonimpulsive signals. We have applied the zero-lag cross-correlation method to earthquakes recorded by three dense arrays in Puget Sound and Vancouver Island to estimate the slowness and back azimuth of direct P waves and S waves. The results are compared with the slowness and back azimuth computed from the source location obtained by the analysis of data recorded by the Pacific Northwest seismic network (PNSN). This comparison has allowed a quantification of the errors associated with the estimation of slowness and back azimuth obtained through the analysis of array data. The statistical analysis gives ??BP = 10?? and ??BS = 8?? as standard deviations for the back azimuth and ??SP = 0.021 sec/km and ??SS = 0.033 sec /km for the slowness results of the P and S phases, respectively. These values are consistent with the theoretical relationship between slowness and back azimuth and their uncertainties. We have tested a probabilistic source location method on the local earthquakes based on the use of the slowness estimated for two or three arrays without taking into account travel-time information. Then we applied the probabilistic method to the deep, nonvolcanic tremor recorded by the arrays during July 2004. The results of the tremor location using the probabilistic method are in good agreement with those obtained by other techniques. The wide depth range, of between 10 and 70 km, and the source migration with time are evident in our results. The method is useful for locating the source of signals characterized by the absence of pickable seismic phases.

Bulletin of the Seismological Society of America

Bayesian approaches to proxy uncertainty quantification in paleoecology: A mathematical justification and practical integration

Paleoenvironmental data are essential for reconstructing environmental conditions in the distant past, and these reconstructions strongly depend on proxies and age–depth models. Proxies are indirect measurements that substitute for variables that cannot be directly measured, such as past precipitation. Conversely, an age–depth model is a tool that correlates the observed proxy with a specific moment in time. Bayesian age–depth modelling has proved to be a powerful method for estimating sediment ages and their associated uncertainties. However, there remains considerable potential for further integration into proxy analysis. In this paper, we explore a mathematical justification and a computational approach that integrates uncertainty at the age–depth level and propagates it to the proxy scale in the form of a posterior predictive distribution. This method mitigates potential biases and errors by removing the need to assign a single age to a given proxy measurement. It allows for quantifying the likelihood that proxy data values correspond to modelled ages, thus enabling the quantification of uncertainty in both the temporal and proxy value domains. The use of Bayesian statistics in proxy analysis represents a relatively recent advancement. We aim to mathematically justify incorporating the Markov chain Monte Carlo output from age–depth models into proxy analysis and to present a novel methodology for constructing environmental reconstructions using this approach.

Journal of Agricultural, Biological and Environmen

Statistical methods in water resources

This text began as a collection of class notes for a course on applied statistical methods for hydrologists taught at the U.S. Geological Survey (USGS) National Training Center. Course material was formalized and organized into a textbook, first published in 1992 by Elsevier as part of their Studies in Environmental Science series. In 2002, the work was made available online as a USGS report. The text has now been updated as a USGS Techniques and Methods Report. It is intended to be a text in applied statistics for hydrology, environmental science, environmental engineering, geology, or biology that addresses distinctive features of environmental data. For example, water resources data tend to have many variables with a lower bound of zero, tend to be more skewed than data from many other disciplines, commonly contain censored data (less than values), and assumptions that the data are normally distributed are not appropriate. Computer-intensive methods (bootstrapping and permutation tests) now improve upon and replace the dependence on t-intervals, t-tests, and analysis of variance. A new chapter on sampling design addresses questions such as “How many observations do I need?” The chapter also presents distribution-free methods to help plan sampling efforts. The trends chapter has been updated to include the WRTDS (Weighted Regressions on Time, Discharge, and Season) method for analysis of water-quality data. This new version contains updated graphics and updated guidance on the use of statistical techniques. The text utilizes R, a programming language and open-source software environment, for all exercises and most graphics, and the R code used to generate figures and examples is provided for download.

Techniques and Methods

Numerical analysis of regional water levels to define aquifer hydrology

Two fundamental methods for studying aquifer hydrology are now in use. The first, applied many years ago, consists of detailed observation of aquifer inflow, outflow, and storage changes, and their variations in time. By analysis of these observations, estimates of the perennial recharge to the aquifer and other pertinent hydrologic data are obtained, all as gross characteristics of the aquifer. The need for greater detail gave rise to a second fundamental method: special field tests, such as pumping tests, by which the hydrologic coefficients could be measured in a comparatively short time. In order to evaluate properly the ability of an aquifer to serve as a source of perennial water supply, the geology and hydrology of the aquifer must be known in some detail over its entire area. The first method cannot supply the necessary detail in most cases, and the second method cannot ordinarily provide the needed areal coverage because of the lack of appropriate testing facilities. Thus an auxiliary third approach was sought which would combine the features of a simple data‐collection program with a final analysis yielding both adequate detail and areal coverage. A method designed to satisfy these requirements is described. Water‐level altitudes, usually observed in the course of more general ground‐water studies, are analyzed by numerical methods, using finite‐difference approximations of the basic differential equations which describe ground‐water flow. Analytical methods are given for nonsteady flow through homogeneous and nonhomogeneous aquifers. Both direct and statistical solutions are shown. The hydrologic factors are computed as functions of transmissibility, and for the nonhomogeneous aquifer the variations of transmissibility in space are computed also from the water‐level data. Knowledge of the absolute value of any one of the hydrologic factors at some location in the aquifer permits conversion of the computed functions to absolute terms for all the aquifer flow field studied.

Eos, Transactions, American Geophysical Union

Combining particle-tracking and geochemical data to assess public supply well vulnerability to arsenic and uranium

Flow-model particle-tracking results and geochemical data from seven study areas across the United States were analyzed using three statistical methods to test the hypothesis that these variables can successfully be used to assess public supply well vulnerability to arsenic and uranium. Principal components analysis indicated that arsenic and uranium concentrations were associated with particle-tracking variables that simulate time of travel and water fluxes through aquifer systems and also through specific redox and pH zones within aquifers. Time-of-travel variables are important because many geochemical reactions are kinetically limited, and geochemical zonation can account for different modes of mobilization and fate. Spearman correlation analysis established statistical significance for correlations of arsenic and uranium concentrations with variables derived using the particle-tracking routines. Correlations between uranium concentrations and particle-tracking variables were generally strongest for variables computed for distinct redox zones. Classification tree analysis on arsenic concentrations yielded a quantitative categorical model using time-of-travel variables and solid-phase-arsenic concentrations. The classification tree model accuracy on the learning data subset was 70%, and on the testing data subset, 79%, demonstrating one application in which particle-tracking variables can be used predictively in a quantitative screening-level assessment of public supply well vulnerability. Ground-water management actions that are based on avoidance of young ground water, reflecting the premise that young ground water is more vulnerable to anthropogenic contaminants than is old ground water, may inadvertently lead to increased vulnerability to natural contaminants due to the tendency for concentrations of many natural contaminants to increase with increasing ground-water residence time.

Journal of Hydrology

Methods for estimating annual exceedance-probability discharges for streams in Iowa, based on data through water year 2010

A statewide study was performed to develop regional regression equations for estimating selected annual exceedance-probability statistics for ungaged stream sites in Iowa. The study area comprises streamgages located within Iowa and 50 miles beyond the State’s borders. Annual exceedance-probability estimates were computed for 518 streamgages by using the expected moments algorithm to fit a Pearson Type III distribution to the logarithms of annual peak discharges for each streamgage using annual peak-discharge data through 2010. The estimation of the selected statistics included a Bayesian weighted least-squares/generalized least-squares regression analysis to update regional skew coefficients for the 518 streamgages. Low-outlier and historic information were incorporated into the annual exceedance-probability analyses, and a generalized Grubbs-Beck test was used to detect multiple potentially influential low flows. Also, geographic information system software was used to measure 59 selected basin characteristics for each streamgage. Regional regression analysis, using generalized least-squares regression, was used to develop a set of equations for each flood region in Iowa for estimating discharges for ungaged stream sites with 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent annual exceedance probabilities, which are equivalent to annual flood-frequency recurrence intervals of 2, 5, 10, 25, 50, 100, 200, and 500 years, respectively. A total of 394 streamgages were included in the development of regional regression equations for three flood regions (regions 1, 2, and 3) that were defined for Iowa based on landform regions and soil regions. Average standard errors of prediction range from 31.8 to 45.2 percent for flood region 1, 19.4 to 46.8 percent for flood region 2, and 26.5 to 43.1 percent for flood region 3. The pseudo coefficients of determination for the generalized least-squares equations range from 90.8 to 96.2 percent for flood region 1, 91.5 to 97.9 percent for flood region 2, and 92.4 to 96.0 percent for flood region 3. The regression equations are applicable only to stream sites in Iowa with flows not significantly affected by regulation, diversion, channelization, backwater, or urbanization and with basin characteristics within the range of those used to develop the equations. These regression equations will be implemented within the U.S. Geological Survey StreamStats Web-based geographic information system tool. StreamStats allows users to click on any ungaged site on a river and compute estimates of the eight selected statistics; in addition, 90-percent prediction intervals and the measured basin characteristics for the ungaged sites also are provided by the Web-based tool. StreamStats also allows users to click on any streamgage in Iowa and estimates computed for these eight selected statistics are provided for the streamgage.

Iowa

Computed statistics at streamgages, and methods for estimating low-flow frequency statistics and development of regional regression equations for estimating low-flow frequency statistics at ungaged locations in Missouri

The weather and precipitation patterns in Missouri vary considerably from year to year. In 2008, the statewide average rainfall was 57.34 inches and in 2012, the statewide average rainfall was 30.64 inches. This variability in precipitation and resulting streamflow in Missouri underlies the necessity for water managers and users to have reliable streamflow statistics and a means to compute select statistics at ungaged locations for a better understanding of water availability. Knowledge of surface-water availability is dependent on the streamflow data that have been collected and analyzed by the U.S. Geological Survey for more than 100 years at approximately 350 streamgages throughout Missouri. The U.S. Geological Survey, in cooperation with the Missouri Department of Natural Resources, computed streamflow statistics at streamgages through the 2010 water year, defined periods of drought and defined methods to estimate streamflow statistics at ungaged locations, and developed regional regression equations to compute selected streamflow statistics at ungaged locations. Streamflow statistics and flow durations were computed for 532 streamgages in Missouri and in neighboring States of Missouri. For streamgages with more than 10 years of record, Kendall’s tau was computed to evaluate for trends in streamflow data. If trends were detected, the variable length method was used to define the period of no trend. Water years were removed from the dataset from the beginning of the record for a streamgage until no trend was detected. Low-flow frequency statistics were then computed for the entire period of record and for the period of no trend if 10 or more years of record were available for each analysis. Three methods are presented for computing selected streamflow statistics at ungaged locations. The first method uses power curve equations developed for 28 selected streams in Missouri and neighboring States that have multiple streamgages on the same streams. Statistical estimates on one of these streams can be calculated at an ungaged location that has a drainage area that is between 40 percent of the drainage area of the farthest upstream streamgage and within 150 percent of the drainage area of the farthest downstream streamgage along the stream of interest. The second method may be used on any stream with a streamgage that has operated for 10 years or longer and for which anthropogenic effects have not changed the low-flow characteristics at the ungaged location since collection of the streamflow data. A ratio of drainage area of the stream at the ungaged location to the drainage area of the stream at the streamgage was computed to estimate the statistic at the ungaged location. The range of applicability is between 40- and 150-percent of the drainage area of the streamgage, and the ungaged location must be located on the same stream as the streamgage. The third method uses regional regression equations to estimate selected low-flow frequency statistics for unregulated streams in Missouri. This report presents regression equations to estimate frequency statistics for the 10-year recurrence interval and for the N-day durations of 1, 2, 3, 7, 10, 30, and 60 days. Basin and climatic characteristics were computed using geographic information system software and digital geospatial data. A total of 35 characteristics were computed for use in preliminary statewide and regional regression analyses based on existing digital geospatial data and previous studies. Spatial analyses for geographical bias in the predictive accuracy of the regional regression equations defined three low-flow regions with the State representing the three major physiographic provinces in Missouri. Region 1 includes the Central Lowlands, Region 2 includes the Ozark Plateaus, and Region 3 includes the Mississippi Alluvial Plain. A total of 207 streamgages were used in the regression analyses for the regional equations. Of the 207 U.S. Geological Survey streamgages, 77 were located in Region 1, 120 were located in Region 2, and 10 were located in Region 3. Streamgages located outside of Missouri were selected to extend the range of data used for the independent variables in the regression analyses. Streamgages included in the regression analyses had 10 or more years of record and were considered to be affected minimally by anthropogenic activities or trends. Regional regression analyses identified three characteristics as statistically significant for the development of regional equations. For Region 1, drainage area, longest flow path, and streamflow-variability index were statistically significant. The range in the standard error of estimate for Region 1 is 79.6 to 94.2 percent. For Region 2, drainage area and streamflow variability index were statistically significant, and the range in the standard error of estimate is 48.2 to 72.1 percent. For Region 3, drainage area and streamflow-variability index also were statistically significant with a range in the standard error of estimate of 48.1 to 96.2 percent. Limitations on the use of estimating low-flow frequency statistics at ungaged locations are dependent on the method used. The first method outlined for use in Missouri, power curve equations, were developed to estimate the selected statistics for ungaged locations on 28 selected streams with multiple streamgages located on the same stream. A second method uses a drainage-area ratio to compute statistics at an ungaged location using data from a single streamgage on the same stream with 10 or more years of record. Ungaged locations on these streams may use the ratio of the drainage area at an ungaged location to the drainage area at a streamgage location to scale the selected statistic value from the streamgage location to the ungaged location. This method can be used if the drainage area of the ungaged location is within 40 to 150 percent of the streamgage drainage area. The third method is the use of the regional regression equations. The limits for the use of these equations are based on the ranges of the characteristics used as independent variables and that streams must be affected minimally by anthropogenic activities.

Missouri

Radioactive springs geochemical data related to uranium exploration: basic data and use of multivariate factor scores

Radioactive springs and wells at 33 localities in the States of Colorado, Utah, Arizona, and New Mexico have been studied and sampled to obtain geochemical data to determine whether such data are useful in a uranium exploration program. Most samples were collected from mineral-rich springs probably related to hydrothermal systems of various ages. Two sets of data were obtained, the first based on the chemical composition and physical and chemical properties of spring and ground water, and the second based on the chemical composition of mineral precipitates deposited by radioactive springs. Multivariate statistical analysis of the water data suggests four major geochemical factors affecting the 23 parameters measured. These factors were labeled as total dissolved solids, alkalinity, temperature, and Fe-U concentration. Multivariate statistical analysis of the precipitate data suggests five factors affecting the 32 element values measured. These factors were labeled as mineral contamination, Mn precipitation, Fe-As-Be precipitation, heavy metals precipitation, and Ba-Ra precipitation. Relative intensities of the geochemical processes represented by the factors were computed using factor scores. Sample localities were ranked on the basis of relative intensities, and the five localities with the highest intensities were selected as being the most favorable for more intensive exploration for uranium. Immediate use of such selection would be experimental because of the lack of industry experience at this time in the exploration of active hydrothermal systems for uranium.

Open-File Report

Methods for estimating selected low-flow frequency statistics and harmonic mean flows for streams in Iowa

A statewide study was conducted to develop regression equations for estimating six selected low-flow frequency statistics and harmonic mean flows for ungaged stream sites in Iowa. The estimation equations developed for the six low-flow frequency statistics include: the annual 1-, 7-, and 30-day mean low flows for a recurrence interval of 10 years, the annual 30-day mean low flow for a recurrence interval of 5 years, and the seasonal (October 1 through December 31) 1- and 7-day mean low flows for a recurrence interval of 10 years. Estimation equations also were developed for the harmonic-mean-flow statistic. Estimates of these seven selected statistics are provided for 208 U.S. Geological Survey continuous-record streamgages using data through September 30, 2006. The study area comprises streamgages located within Iowa and 50 miles beyond the State's borders. Because trend analyses indicated statistically significant positive trends when considering the entire period of record for the majority of the streamgages, the longest, most recent period of record without a significant trend was determined for each streamgage for use in the study. The median number of years of record used to compute each of these seven selected statistics was 35. Geographic information system software was used to measure 54 selected basin characteristics for each streamgage. Following the removal of two streamgages from the initial data set, data collected for 206 streamgages were compiled to investigate three approaches for regionalization of the seven selected statistics. Regionalization, a process using statistical regression analysis, provides a relation for efficiently transferring information from a group of streamgages in a region to ungaged sites in the region. The three regionalization approaches tested included statewide, regional, and region-of-influence regressions. For the regional regression, the study area was divided into three low-flow regions on the basis of hydrologic characteristics, landform regions, and soil regions. A comparison of root mean square errors and average standard errors of prediction for the statewide, regional, and region-of-influence regressions determined that the regional regression provided the best estimates of the seven selected statistics at ungaged sites in Iowa. Because a significant number of streams in Iowa reach zero flow as their minimum flow during low-flow years, four different types of regression analyses were used: left-censored, logistic, generalized-least-squares, and weighted-least-squares regression. A total of 192 streamgages were included in the development of 27 regression equations for the three low-flow regions. For the northeast and northwest regions, a censoring threshold was used to develop 12 left-censored regression equations to estimate the 6 low-flow frequency statistics for each region. For the southern region a total of 12 regression equations were developed; 6 logistic regression equations were developed to estimate the probability of zero flow for the 6 low-flow frequency statistics and 6 generalized least-squares regression equations were developed to estimate the 6 low-flow frequency statistics, if nonzero flow is estimated first by use of the logistic equations. A weighted-least-squares regression equation was developed for each region to estimate the harmonic-mean-flow statistic. Average standard errors of estimate for the left-censored equations for the northeast region range from 64.7 to 88.1 percent and for the northwest region range from 85.8 to 111.8 percent. Misclassification percentages for the logistic equations for the southern region range from 5.6 to 14.0 percent. Average standard errors of prediction for generalized least-squares equations for the southern region range from 71.7 to 98.9 percent and pseudo coefficients of determination for the generalized-least-squares equations range from 87.7 to 91.8 percent. Average standard errors of prediction for weighted-least-squares equations developed for estimating the harmonic-mean-flow statistic for each of the three regions range from 66.4 to 80.4 percent. The regression equations are applicable only to stream sites in Iowa with low flows not significantly affected by regulation, diversion, or urbanization and with basin characteristics within the range of those used to develop the equations. If the equations are used at ungaged sites on regulated streams, or on streams affected by water-supply and agricultural withdrawals, then the estimates will need to be adjusted by the amount of regulation or withdrawal to estimate the actual flow conditions if that is of interest. Caution is advised when applying the equations for basins with characteristics near the applicable limits of the equations and for basins located in karst topography. A test of two drainage-area ratio methods using 31 pairs of streamgages, for the annual 7-day mean low-flow statistic for a recurrence interval of 10 years, indicates a weighted drainage-area ratio method provides better estimates than regional regression equations for an ungaged site on a gaged stream in Iowa when the drainage-area ratio is between 0.5 and 1.4. These regression equations will be implemented within the U.S. Geological Survey StreamStats web-based geographic-information-system tool. StreamStats allows users to click on any ungaged site on a river and compute estimates of the seven selected statistics; in addition, 90-percent prediction intervals and the measured basin characteristics for the ungaged sites also are provided. StreamStats also allows users to click on any streamgage in Iowa and estimates computed for these seven selected statistics are provided for the streamgage.

Iowa

Models for inference in dynamic metacommunity systems

A variety of processes are thought to be involved in the formation and dynamics of species assemblages. For example, various metacommunity theories are based on differences in the relative contributions of dispersal of species among local communities and interactions of species within local communities. Interestingly, metacommunity theories continue to be advanced without much empirical validation. Part of the problem is that statistical models used to analyze typical survey data either fail to specify ecological processes with sufficient complexity or they fail to account for errors in detection of species during sampling. In this paper, we describe a statistical modeling framework for the analysis of metacommunity dynamics that is based on the idea of adopting a unified approach, multispecies occupancy modeling, for computing inferences about individual species, local communities of species, or the entire metacommunity of species. This approach accounts for errors in detection of species during sampling and also allows different metacommunity paradigms to be specified in terms of species‐ and location‐specific probabilities of occurrence, extinction, and colonization: all of which are estimable. In addition, this approach can be used to address inference problems that arise in conservation ecology, such as predicting temporal and spatial changes in biodiversity for use in making conservation decisions. To illustrate, we estimate changes in species composition associated with the species‐specific phenologies of flight patterns of butterflies in Switzerland for the purpose of estimating regional differences in biodiversity.

Ecology

Exploratory analysis of environmental interactions in central California

As part of its global change research program, the United States Geological Survey (USGS) has produced raster data that describe the land cover of the United States using a consistent format. The data consist of elevations, satellite measurements, computed vegetation indices, land cover classes, and ancillary political, topographic and hydrographic information. This open-file report uses some of these data to explore the environment of a (256-km)? region of central California. We present various visualizations of the data, multiscale correlations between topography and vegetation, a path analysis of more complex statistical interactions, and a map that portrays the influence of agriculture on the region's vegetation. An appendix contains C and Mathematica code used to generate the graphics and some of the analysis.

Open-File Report

Oregon ground-water quality and its relation to hydrogeologic factors — A statistical approach

An appraisal of Oregon ground-water quality was made using existing data accessible through the U.S. Geological Survey computer system. The data available for about 1,000 sites were separated by aquifer units and hydrologic units. Selected statistical moments were described for 19 constituents including major ions. About 96 percent of all sites in the data base were sampled only once. The sample data were classified by aquifer unit and hydrologic unit and analysis of variance was run to determine if significant differences exist between the units within each of these two classifications for the same 19 constituents on which statistical moments were determined. Results of the analysis of variance indicated both classification variables performed about the same, but aquifer unit did provide more separation for some constituents. Samples from the Rogue River basin were classified by location within the flow system and type of flow system. The samples were then analyzed using analysis of variance on 14 constituents to determine if there were significant differences between subsets classified by flow path. Results of this analysis were not definitive, but classification as to the type of flow system did indicate potential for segregating water-quality data into distinct subsets.

Oregon