Search USGSSearch

SEARCH · Search USGS

Results for “Computational Statistics and Data Analysis”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Cost-Benefit Analysis of Computer Resources for Machine Learning

Machine learning describes pattern-recognition algorithms - in this case, probabilistic neural networks (PNNs). These can be computationally intensive, in part because of the nonlinear optimizer, a numerical process that calibrates the PNN by minimizing a sum of squared errors. This report suggests efficiencies that are expressed as cost and benefit. The cost is computer time needed to calibrate the PNN, and the benefit is goodness-of-fit, how well the PNN learns the pattern in the data. There may be a point of diminishing returns where a further expenditure of computer resources does not produce additional benefits. Sampling is suggested as a cost-reduction strategy. One consideration is how many points to select for calibration and another is the geometric distribution of the points. The data points may be nonuniformly distributed across space, so that sampling at some locations provides additional benefit while sampling at other locations does not. A stratified sampling strategy can be designed to select more points in regions where they reduce the calibration error and fewer points in regions where they do not. Goodness-of-fit tests ensure that the sampling does not introduce bias. This approach is illustrated by statistical experiments for computing correlations between measures of roadless area and population density for the San Francisco Bay Area. The alternative to training efficiencies is to rely on high-performance computer systems. These may require specialized programming and algorithms that are optimized for parallel performance.

Open-File Report

Mapping cropland extent of Southeast and Northeast Asia using multi-year time-series Landsat 30-m data using Random Forest classifier on Google Earth Engine

Cropland extent maps are useful components for assessing food security. Ideally, such products are a useful addition to countrywide agricultural statistics since they are not politically biased and can be used to calculate cropland area for any spatial unit from an individual farm to various administrative unites (e.g., state, county, district) within and across nations, which in turn can be used to estimate agricultural productivity as well as degree of disturbance on food security from natural disasters and political conflict. However, existing cropland extent maps over large areas (e.g., Country, region, continent, world) are derived from coarse resolution imagery (250 m to 1 km pixels) and have many limitations such as missing fragmented and\or small farms with mixed signatures from different crop types and\or farming practices that can be, confused with other land cover. As a result, the coarse resolution maps have limited useflness in areas where fields are small (<1 ha), such as in Southeast Asia. Furthermore, coarse resolution cropland maps have known uncertainties in both geo-precision of cropland location as well as accuracies of the product. To overcome these limitations, this research was conducted using multi-date, multi-year 30-m Landsat time-series data for 3 years chosen from 2013 to 2016 for all Southeast and Northeast Asian Countries (SNACs), which included 7 refined agro-ecological zones (RAEZ) and 12 countries (Indonesia, Thailand, Myanmar, Vietnam, Malaysia, Philippines, Cambodia, Japan, North Korea, Laos, South Korea, and Brunei). The 30-m (1 pixel = 0.09 ha) data from Landsat 8 Operational Land Imager (OLI) and Landsat 7 Enhanced Thematic Mapper (ETM+) were used in the study. Ten Landsat bands were used in the analysis (blue, green, red, NIR, SWIR1, SWIR2, Thermal, NDVI, NDWI, LSWI) along with additional layers of standard deviation of these 10 bands across 1 year, and global digital elevation model (GDEM)-derived slope and elevation bands. To reduce the impact of clouds, the Landsat imagery was time-composited over four time-periods (Period 1: January- April, Period 2: May-August, and Period 3: September-December) over 3-years. Period 4 was the standard deviation of all 10 bands taken over all images acquired during the 2015 calendar year. These four period composites, totaling 42 band data-cube, were generated for each of the 7 RAEZs. The reference training data (N = 7849) generated for the 7 RAEZ using sub-meter to 5-m very high spatial resolution imagery (VHRI) helped generate the knowledge-base to separate croplands from non-croplands. This knowledge-base was used to code and run a pixel-based random forest (RF) supervised machine learning algorithm on the Google Earth Engine (GEE) cloud computing environment to separate croplands from non-croplands. The resulting cropland extent products were evaluated using an independent reference validation dataset (N = 1750) in each of the 7 RAEZs as well as for the entire SNAC area. For the entire SNAC area, the overall accuracy was 88.1% with a producer’s accuracy of 81.6% (errors of omissions = 18.4%) and user’s accuracy of 76.7% (errors of commissions = 23.3%). For each of the 7 RAEZs overall accuracies varied from 83.2 to 96.4%. Cropland areas calculated for the 12 countries were compared with country areas reported by the United Nations Food and Agriculture Organization and other national cropland statistics resulting in an R 2 value of 0.93. The cropland areas of provinces were compared with the province statistics that showed an R 2 = 0.95 for South Korea and R 2 = 0.94 for Thailand. The cropland products are made available on an interactive viewer at www.croplands.org and for download at National Aeronautics and Space Administration’s (NASA) Land Processes Distributed Active Archive Center (LP DAAC): https://lpdaac.usgs.gov/node/1281 .

International Journal of Applied Earth Observation

Concentrations of Escherichia coli in streams in the Ohio River Watershed in Indiana, May—August 2000

Water samples collected from 40 stream sites in the Ohio River Watershed in Indiana from May through August 2000 were analyzed for concentrations of Escherichia coli (E. coli) bacteria. Each site was sampled five times in a 30-day period. Concentrations of E. coli in 72 of the 200 samples exceeded the State of Indiana single-sample standard of 235 colonies per 100 milliliters for waters used for recreation. A five-sample geometric mean of concentrations was computed for each site. Concentrations in samples from 24 of the 40 sites exceeded the State of Indiana standard for a five-sample geometric mean of 125 colonies per 100 milliliters for waters used for recreation. Samples collected from 34 sites had E. coli concentrations that exceeded either or both the single-sample standard and the five-sample geometric mean standard. Five of the 40 sampling sites were at or near U.S. Geological Survey streamflowgaging stations. On the basis of records from these stations, 16 percent of the samples from these sites were collected at streamflows above the median daily mean discharge for each station. E. coli concentrations and turbidity measurements collected during 2000 were analyzed in concert with E. coli concentration and turbidity data collected in 1999 at streams within the Kankakee and Lower Wabash River Watersheds in Indiana and in 1998 at streams within the Upper Wabash River Watershed in Indiana. These data were grouped together to investigate the relation between concentrations of bacteria and turbidity. The resulting analysis indicated a statistically significant correlation between concentrations of E. coli and turbidity.

Indiana

Analysis of multinomial models with unknown index using data augmentation

Multinomial models with unknown index ('sample size') arise in many practical settings. In practice, Bayesian analysis of such models has proved difficult because the dimension of the parameter space is not fixed, being in some cases a function of the unknown index. We describe a data augmentation approach to the analysis of this class of models that provides for a generic and efficient Bayesian implementation. Under this approach, the data are augmented with all-zero detection histories. The resulting augmented dataset is modeled as a zero-inflated version of the complete-data model where an estimable zero-inflation parameter takes the place of the unknown multinomial index. Interestingly, data augmentation can be justified as being equivalent to imposing a discrete uniform prior on the multinomial index. We provide three examples involving estimating the size of an animal population, estimating the number of diabetes cases in a population using the Rasch model, and the motivating example of estimating the number of species in an animal community with latent probabilities of species occurrence and detection.

Journal of Computational and Graphical Statistics

Estimating species richness and accumulation by modeling species occurrence and detectability

A statistical model is developed for estimating species richness and accumulation by formulating these community-level attributes as functions of model-based estimators of species occurrence while accounting for imperfect detection of individual species. The model requires a sampling protocol wherein repeated observations are made at a collection of sample locations selected to be representative of the community. This temporal replication provides the data needed to resolve the ambiguity between species absence and nondetection when species are unobserved at sample locations. Estimates of species richness and accumulation are computed for two communities, an avian community and a butterfly community. Our model-based estimates suggest that detection failures in many bird species were attributed to low rates of occurrence, as opposed to simply low rates of detection. We estimate that the avian community contains a substantial number of uncommon species and that species richness greatly exceeds the number of species actually observed in the sample. In fact, predictions of species accumulation suggest that even doubling the number of sample locations would not have revealed all of the species in the community. In contrast, our analysis of the butterfly community suggests that many species are relatively common and that the estimated richness of species in the community is nearly equal to the number of species actually detected in the sample. Our predictions of species accumulation suggest that the number of sample locations actually used in the butterfly survey could have been cut in half and the asymptotic richness of species still would have been attained. Our approach of developing occurrence-based summaries of communities while allowing for imperfect detection of species is broadly applicable and should prove useful in the design and analysis of surveys of biodiversity.

Ecology

Low-flow frequency and flow duration of selected South Carolina streams in the Pee Dee River basin through March 2007

Part of the mission of the South Carolina Department of Health and Environmental Control and the South Carolina Department of Natural Resources is to protect and preserve South Carolina's water resources. Doing so requires an ongoing understanding of streamflow characteristics of the rivers and streams in South Carolina. A particular need is information concerning the low-flow characteristics of streams; this information is especially important for effectively managing the State's water resources during critical flow periods such as the severe drought that occurred between 1998 and 2002 and the most recent drought that occurred between 2006 and 2009. In 2008, the U.S. Geological Survey, in cooperation with the South Carolina Department of Health and Environmental Control, initiated a study to update low-flow statistics at continuous-record streamgaging stations operated by the U.S. Geological Survey in South Carolina. Under this agreement, the low-flow characteristics at continuous-record streamgaging stations will be updated in a systematic manner during the monitoring and assessment of the eight major basins in South Carolina as defined and grouped according to the South Carolina Department of Health and Environmental Control's Watershed Water Quality Management Strategy. Depending on the length of record available at the continuous-record streamgaging stations, low-flow frequency characteristics are estimated for annual minimum 1-, 3-, 7-, 14-, 30-, 60-, and 90-day average flows with recurrence intervals of 2, 5, 10, 20, 30, and 50 years. Low-flow statistics are presented for 18 streamgaging stations in the Pee Dee River basin. In addition, daily flow durations for the 5-, 10-, 25-, 50-, 75-, 90-, and 95-percent probability of exceedance also are presented for the stations. The low-flow characteristics were computed from records available through March 31, 2007. The last systematic update of low-flow characteristics in South Carolina occurred more than 20 years ago and included data through March 1987. Of the 17 streamgaging stations included in this study, 15 had low-flow characteristics that were published in previous U.S. Geological Survey reports. A comparison of the low-flow characteristic for the minimum average flow for a 7-consecutive-day period with a 10-year recurrence interval from this study with the most recently published values indicated that 10 of the 15 streamgaging stations had values that were within &plusmn;25 percent of each other. Nine of the 15 streamgaging stations had negative percentage differences indicating the low-flow statistic had decreased since the previous study, 4 streamgaging stations had positive percent differences indicating that the low-flow statistic had increased since the previous study, and 2 streamgaging stations had a zero percent difference indicating no change since the previous study. The low-flow characteristics are influenced by length of record, hydrologic regime under which the record was collected, techniques used to do the analysis, and other changes that may have occurred in the watershed.

North Carolina, South Carolina

Estimating the Magnitude and Frequency of Peak Streamflows for Ungaged Sites on Streams in Alaska and Conterminous Basins in Canada

Estimates of the magnitude and frequency of peak streamflow are needed across Alaska for floodplain management, cost-effective design of floodway structures such as bridges and culverts, and other water-resource management issues. Peak-streamflow magnitudes for the 2-, 5-, 10-, 25-, 50-, 100-, 200-, and 500-year recurrence-interval flows were computed for 301 streamflow-gaging and partial-record stations in Alaska and 60 stations in conterminous basins of Canada. Flows were analyzed from data through the 1999 water year using a log-Pearson Type III analysis. The State was divided into seven hydrologically distinct streamflow analysis regions for this analysis, in conjunction with a concurrent study of low and high flows. New generalized skew coefficients were developed for each region using station skew coefficients for stations with at least 25 years of systematic peak-streamflow data. Equations for estimating peak streamflows at ungaged locations were developed for Alaska and conterminous basins in Canada using a generalized least-squares regression model. A set of predictive equations for estimating the 2-, 5-, 10-, 25-, 50-, 100-, 200-, and 500-year peak streamflows was developed for each streamflow analysis region from peak-streamflow magnitudes and physical and climatic basin characteristics. These equations may be used for unregulated streams without flow diversions, dams, periodically releasing glacial impoundments, or other streamflow conditions not correlated to basin characteristics. Basin characteristics should be obtained using methods similar to those used in this report to preserve the statistical integrity of the equations.

Water-Resources Investigations Report

A practical primer on geostatistics

Introduction The Challenge —Most geological phenomena are extraordinarily complex in their interrelationships and vast in their geographical extension. Ordinarily, engineers and geoscientists are faced with corporate or scientific requirements to properly prepare geological models with measurements involving a small fraction of the entire area or volume of interest. Exact description of a system such as an oil reservoir is neither feasible nor economically possible. The results are necessarily uncertain. Note that the uncertainty is not an intrinsic property of the systems; it is the result of incomplete knowledge by the observer. The Aim of Geostatistics —The main objective of geostatistics is the characterization of spatial systems that are incompletely known, systems that are common in geology. A key difference from classical statistics is that geostatistics uses the sampling location of every measurement. Unless the measurements show spatial correlation, the application of geostatistics is pointless. Ordinarily the need for additional knowledge goes beyond a few points, which explains the display of results graphically as fishnet plots, block diagrams, and maps. Geostatistical Methods —Geostatistics is a collection of numerical techniques for the characterization of spatial attributes using primarily two tools: probabilistic models, which are used for spatial data in a manner similar to the way in which time-series analysis characterizes temporal data, or pattern recognition techniques. The probabilistic models are used as a way to handle uncertainty in results away from sampling locations, making a radical departure from alternative approaches like inverse distance estimation methods. Differences with Time Series —On dealing with time-series analysis, users frequently concentrate their attention on extrapolations for making forecasts. Although users of geostatistics may be interested in extrapolation, the methods work at their best interpolating. This simple difference has significant methodological implications. Historical Remarks —As a discipline, geostatistics was firmly established in the 1960s by the French engineer Georges Matheron, who was interested in the appraisal of ore reserves in mining. Geostatistics did not develop overnight. Like other disciplines, it has built on previous results, many of which were formulated with different objectives in various fields. Pioneers —Seminal ideas conceptually related to what today we call geostatistics or spatial statistics are found in the work of several pioneers, including: 1940s: A.N. Kolmogorov in turbulent flow and N. Wiener in stochastic processing; 1950s: D. Krige in mining; 1960s: B. Mathern in forestry and L.S. Gandin in meteorology Calculations —Serious applications of geostatistics require the use of digital computers. Although for most geostatistical techniques rudimentary implementation from scratch is fairly straightforward, coding programs from scratch is recommended only as part of a practice that may help users to gain a better grasp of the formulations. Software —For professional work, the reader should employ software packages that have been thoroughly tested to handle any sampling scheme, that run as efficiently as possible, and that offer graphic capabilities for the analysis and display of results. This primer employs primarily the package Stanford Geomodeling Software (SGeMS) - recently developed at the Energy Resources Engineering Department at Stanford University - as a way to show how to obtain results practically. This applied side of the primer should not be interpreted as the notes being a manual for the use of SGeMS. The main objective of the primer is to help the reader gain an understanding of the fundamental concepts and tools in geostatistics. Organization of the Primer —The chapters of greatest importance are those covering kriging and simulation. All other materials are peripheral and are included for better comprehension of these main geostatistical modeling tools. The choice of kriging versus simulation is often a big puzzle to the uninitiated, let alone the different variants of both of them. Chapters 14, 18, and 19 are intended to shed light on those subjects. The critical aspect of assessing and modeling spatial correlation is covered in chapter 7. Chapters 2 and 3 review relevant concepts in classical statistics. Course Objectives —This course offers stochastic solutions to common problems in the characterization of complex geological systems. At the end of the course, participants should have: an understanding of the theoretical foundations of geostatistics; a good grasp of its possibilities and limitations; and reasonable familiarity with the SGeMS software, thus opening the possibility of practically applying geostatistics.

Open-File Report

Development of regression equations for the estimation of the magnitude and frequency of floods at rural, unregulated gaged and ungaged streams in Puerto Rico through water year 2017

The methods of computation and estimates of the magnitude of flood flows were updated for the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent chance exceedance levels for 91 streamgages on the main island of Puerto Rico by using annual peak-flow data through 2017. Since the previous flood frequency study in 1994, the U.S. Geological Survey has collected additional peak flows at additional streamgages, and Puerto Rico has experienced numerous flood events. This updated study was performed using longer annual peak-flow datasets from more stations to provide more representative equations to predict flood flows. Screening criteria for these streamgages included 10 or more years of annual peak-flow data, unregulated flow, and less than 10 percent impervious drainage area. The magnitude and frequency of floods at selected streamgages in Puerto Rico were estimated using updated methods outlined in Bulletin 17C. The new procedures include a regional skew analysis that incorporates Bayesian regression techniques, the Expected Moments Algorithm to better represent missing record and estimate parameters of the log-Pearson Type III distribution, and the Multiple Grubbs-Beck test for low outlier detection. Regional regression equations were developed to estimate peak-flow statistics at ungaged locations by using selected basin and climatic characteristics as explanatory variables. These variables were determined from digital spatial datasets and geographic information systems by using the most recent data available. Ordinary least-squares regression techniques were used to filter the basin characteristics and determine two separate regions, region 1 (west) and region 2 (east), based on residuals. A generalized least-squares procedure was used to account for cross-correlation of sites and develop the final set of equations that have drainage area as the only explanatory variable. The average standard errors of prediction ranged from 18.7 to 46.7 percent in region 1 and 33.4 to 57.6 percent in region 2 for all annual exceedance probabilities (AEPs) examined. The updated statistics showed a greater accuracy of prediction when compared to those from the previous study using drainage area as the only explanatory variable for all AEPs examined in region 1 and the 0.01 and 0.002 AEP flows for region 2. When compared to equations developed in the previous study that have drainage area, mean annual rainfall, and (or) depth-to-rock as explanatory variables, the updated statistics show a greater accuracy of prediction in region 1 at AEP flows of 0.02 and lower (that is, higher flows). Those developed for region 2 do not show a greater accuracy of prediction for any AEP flows when compared to the equations having multiple explanatory variables in the previous study. The calculated regression equations, basin characteristics, and at-site statistics will be incorporated into the U.S. Geological Survey web application, StreamStats ( https://streamstats.usgs.gov/ss/ ). This application allows users to select a location on a stream, whether gaged or ungaged, to obtain estimates of basin characteristics and flow statistics.

Puerto Rico

Statistical approaches for modeling correlated grade and tonnage distributions and applications for mineral resource assessments

Correlations between grade and tonnage exist in mineral resource data compiled from published reports, but they are not always addressed during quantitative assessment of undiscovered mineral resources. Failure to account for correlated grade and tonnage distributions can result in geologically unrealistic assessment results. Current software tools simulate univariate ore tonnage and multivariate resource grades of undiscovered deposits independently. As a result, analysts are forced to rely on ad-hoc solutions to minimize the correlation issues by: 1) creating subsets of data with restricted criteria; 2) truncating grade and tonnage distributions; and 3) testing model robustness using exploratory data analysis. While these methods represent pragmatic solutions, the statistical solutions presented here provide additional options to address real correlations in grade and tonnage data used for mineral resource assessments. We present a modified version of the MapMark4 package in R that introduces two alternatives for modeling grade and tonnage distributions, consisting of a multivariate solution that accounts for correlations between ore tonnage and metal grades and an empirical solution that utilizes simple random sampling with replacement to reproduce coupled grades and tonnages from the input data. We present simulations for contained ore and metal for three case studies representing tungsten skarn, komatiite-hosted nickel, and sediment-hosted carbonate amagmatic zinc-lead (Mississippi Valley-type) deposits. Employing the methods presented here yields quantitative mineral resource assessment results that more closely reflect the empirical distributions of grades and tonnages observed in nature and expands the applicability of these tools for ongoing critical mineral resource assessments.

Applied Computing and Geosciences

Monitoring and modeling to predict Escherichia coli at Presque Isle Beach 2, City of Erie, Erie County, Pennsylvania

The Lake Erie shoreline in Pennsylvania spans nearly 40 miles and is a valuable recreational resource for Erie County. Nearly 7 miles of the Lake Erie shoreline lies within Presque Isle State Park in Erie, Pa. Concentrations of Escherichia coli (E. coli) bacteria at permitted Presque Isle beaches occasionally exceed the single-sample bathing-water standard, resulting in unsafe swimming conditions and closure of the beaches. E. coli concentrations and other water-quality and environmental data collected at Presque Isle Beach 2 during the 2004 and 2005 recreational seasons were used to develop models using tobit regression analyses to predict E. coli concentrations. All variables statistically related to E. coli concentrations were included in the initial regression analyses, and after several iterations, only those explanatory variables that made the models significantly better at predicting E. coli concentrations were included in the final models. Regression models were developed using data from 2004, 2005, and the combined 2-year dataset. Variables in the 2004 model and the combined 2004-2005 model were log10 turbidity, rain weight, wave height (calculated), and wind direction. Variables in the 2005 model were log10 turbidity and wind direction. Explanatory variables not included in the final models were water temperature, streamflow, wind speed, and current speed; model results indicated these variables did not meet significance criteria at the 95-percent confidence level (probabilities were greater than 0.05). The predicted E. coli concentrations produced by the models were used to develop probabilities that concentrations would exceed the single-sample bathing-water standard for E. coli of 235 colonies per 100 milliliters. Analysis of the exceedence probabilities helped determine a threshold probability for each model, chosen such that the correct number of exceedences and nonexceedences was maximized and the number of false positives and false negatives was minimized. Future samples with computed exceedence probabilities higher than the selected threshold probability, as determined by the model, will likely exceed the E. coli standard and a beach advisory or closing may need to be issued; computed exceedence probabilities lower than the threshold probability will likely indicate the standard will not be exceeded. Additional data collected each year can be used to test and possibly improve the model. This study will aid beach managers in more rapidly determining when waters are not safe for recreational use and, subsequently, when to issue beach advisories or closings.

Pennsylvania

National shoreline change—Summary statistics of shoreline change from the 1800s to the 2010s for the coast of California

Rates of shoreline change have been updated for the open-ocean sandy coastline of California as part of studies conducted by the U.S. Geological Survey. Shorelines from the original assessment (1800s through 1998 or 2002), as well as additional shoreline position data from 2009 to 2011, 2015, and 2016 extracted from light detection and ranging (lidar) data, were used to compute long-term rates (approximately 150 years) that incorporate the proxy-datum bias on a transect-by-transect basis. The proxy-datum bias accounts for the unidirectional onshore bias of proxy-based high water line shorelines relative to datum-based mean high water shorelines. In areas where the methods for delineating shorelines did not make it possible to compute a bias correction, the rates are reported without that correction. In this study, the coasts of northern and central California exhibited the highest average rates of erosion, whereas southern California exhibited the highest average rate of accretion. The maximum erosion rate was in San Mateo County in central California. The maximum rate of accretion was in Humboldt County in northern California. Rates were calculated at 19,063 transect locations. Shoreline positions from the mid-1800s through 2016 were used to update shoreline change rates in California using the Digital Shoreline Analysis System (DSAS) software.

California

Regression equations for estimation of annual peak-streamflow frequency for undeveloped watersheds in Texas using an L-moment-based, PRESS-minimized, residual-adjusted approach

Annual peak-streamflow frequency estimates are needed for flood-plain management; for objective assessment of flood risk; for cost-effective design of dams, levees, and other flood-control structures; and for design of roads, bridges, and culverts. Annual peak-streamflow frequency represents the peak streamflow for nine recurrence intervals of 2, 5, 10, 25, 50, 100, 200, 250, and 500 years. Common methods for estimation of peak-streamflow frequency for ungaged or unmonitored watersheds are regression equations for each recurrence interval developed for one or more regions; such regional equations are the subject of this report. The method is based on analysis of annual peak-streamflow data from U.S. Geological Survey streamflow-gaging stations (stations). Beginning in 2007, the U.S. Geological Survey, in cooperation with the Texas Department of Transportation and in partnership with Texas Tech University, began a 3-year investigation concerning the development of regional equations to estimate annual peak-streamflow frequency for undeveloped watersheds in Texas. The investigation focuses primarily on 638 stations with 8 or more years of data from undeveloped watersheds and other criteria. The general approach is explicitly limited to the use of L-moment statistics, which are used in conjunction with a technique of multi-linear regression referred to as PRESS minimization. The approach used to develop the regional equations, which was refined during the investigation, is referred to as the 'L-moment-based, PRESS-minimized, residual-adjusted approach'. For the approach, seven unique distributions are fit to the sample L-moments of the data for each of 638 stations and trimmed means of the seven results of the distributions for each recurrence interval are used to define the station specific, peak-streamflow frequency. As a first iteration of regression, nine weighted-least-squares, PRESS-minimized, multi-linear regression equations are computed using the watershed characteristics of drainage area, dimensionless main-channel slope, and mean annual precipitation. The residuals of the nine equations are spatially mapped, and residuals for the 10-year recurrence interval are selected for generalization to 1-degree latitude and longitude quadrangles. The generalized residual is referred to as the OmegaEM parameter and represents a generalized terrain and climate index that expresses peak-streamflow potential not otherwise represented in the three watershed characteristics. The OmegaEM parameter was assigned to each station, and using OmegaEM, nine additional regression equations are computed. Because of favorable diagnostics, the OmegaEM equations are expected to be generally reliable estimators of peak-streamflow frequency for undeveloped and ungaged stream locations in Texas. The mean residual standard error, adjusted R-squared, and percentage reduction of PRESS by use of OmegaEM are 0.30log 10 , 0.86, and -21 percent, respectively. Inclusion of the OmegaEM parameter provides a substantial reduction in the PRESS statistic of the regression equations and removes considerable spatial dependency in regression residuals. Although the OmegaEM parameter requires interpretation on the part of analysts and the potential exists that different analysts could estimate different values for a given watershed, the authors suggest that typical uncertainty in the OmegaEM estimate might be about +or-0.10 10 . Finally, given the two ensembles of equations reported herein and those in previous reports, hydrologic design engineers and other analysts have several different methods, which represent different analytical tracks, to make comparisons of peak-streamflow frequency estimates for ungaged watersheds in the study area.

Texas

Real-time validation of the Dst Predictor model

The Dst Predictor model, which has been running real-time in the Space Weather Analysis and Forecast System (SWAFS), provides 1-hour and 4-hour forecasts of the Dst index. This is useful for awareness of impending geomagnetic activity, as well as driving other real-time models that use Dst as an input. In this report, we examine the performance of this forecast model in detail. When validating indices it should be noted that performance is only with respect to a reference index as they are derived quantities assumed to reflect a state of the magnetosphere that cannot be directly measured. In this case U.S. Geological Survey (USGS) Definitive Dst is the reference index (Section 3). Whether or not the model better reflects the actual activity level is nearly impossible to discern and is outside the scope of this report. We evaluate the performance of the model by computing continuous predictant skill scores against USGS Definitive Dst values as “observations” (Section 4.2). The two sets of data are not well-correlated for both 1-hour and 4-hour forecasts. The Dst Predictor Prediction Efficiency for both the 1- and 4-hour forecasts suggests poor performance versus the climatological mean. However, the skill score against a nowcast persistence model is positive, suggesting value added by the Dst Predictor model. We further examine statistics for storm times (Section 4.3) with similar results: nowcast persistence performs worse than Dst Predictor. Dst Predictor is superior to the nowcast persistence model for the metric used in this study. We recommend continued use of the DstPredictor model for 1-and4-hour Dst predictions along with active study of other Dst forecast models that do not rely on nowcast inputs (Section 6). The lack of certified requirements makes further recommendations difficult. A study of how the error in Dst translates to error in models and a better understanding of operational needs for magnetic storm warning are needed to determine such requirements. Nowcast persistence is often hard to beat for short term forecasts and specification and Dst Predictor clearly performs well against that standard (with 1-hour and 4-hour skill-scores of 0.233 and 0.485 respectively), although poor in absolute terms (with1-hourand4-hour prediction efficiencies of-64.6and-43.1, respectively).

Air Force Research Laboratory Technical Report

Synthesis of monthly and annual streamflow records (water years 1950-2003) for Big Sandy, Clear, Peoples, and Beaver Creeks in the Milk River basin, Montana

To address concerns expressed by the State of Montana about the apportionment of water in the St. Mary and Milk River basins between Canada and the United States, the International Joint Commission requested information from the United States government about water that originates in the United States but does not cross the border into Canada. In response to this request, the U.S. Geological Survey synthesized monthly and annual streamflow records for Big Sandy, Clear, Peoples, and Beaver Creeks, all of which are in the Milk River basin in Montana, for water years 1950-2003. This report presents the synthesized values of monthly and annual streamflow for Big Sandy, Clear, Peoples, and Beaver Creeks in Montana. Synthesized values were derived from recorded and estimated streamflows. Statistics, including long-term medians and averages and flows for various exceedance probabilities, were computed from the synthesized data. Beaver Creek had the largest median annual discharge (19,490 acre-feet), and Clear Creek had the smallest median annual discharge (6,680 acre-feet). Big Sandy Creek, the stream with the largest drainage area, had the second smallest median annual discharge (9,640 acre-feet), whereas Peoples Creek, the stream with the second smallest drainage area, had the second largest median annual discharge (11,700 acre-feet). The combined median annual discharge for the four streams was 45,400 acre-feet. The largest combined median monthly discharge for the four creeks was 6,930 acre-feet in March, and the smallest combined median monthly discharge was 48 acre-feet in January. The combined median monthly values were substantially smaller than the average monthly values. Overall, synthesized flow records for the four creeks are considered to be reasonable given the prevailing climatic conditions in the region during the 1950-2003 base period. Individual estimates of monthly streamflow may have large errors, however. Linear regression was used to relate logarithms of combined annual streamflow to water years 1950-2003. The results of the regression analysis indicated a significant downward trend (regression line slope was -0.00977) for combined annual streamflow. A regression analysis using data from 1956-2003 indicated a slight, but not significant, downward trend for combined annual streamflow.

Scientific Investigations Report

Greater than the sum of its parts: Computationally flexible Bayesian hierarchical modeling

We propose a multistage method for making inference at all levels of a Bayesian hierarchical model (BHM) using natural data partitions to increase efficiency by allowing computations to take place in parallel form using software that is most appropriate for each data partition. The full hierarchical model is then approximated by the product of independent normal distributions for the data component of the model. In the second stage, the Bayesian maximum a posteriori (MAP) estimator is found by maximizing the approximated posterior density with respect to the parameters. If the parameters of the model can be represented as normally distributed random effects, then the second-stage optimization is equivalent to fitting a multivariate normal linear mixed model. We consider a third stage that updates the estimates of distinct parameters for each data partition based on the results of the second stage. The method is demonstrated with two ecological data sets and models, a generalized linear mixed effects model (GLMM) and an integrated population model (IPM). The multistage results were compared to estimates from models fit in single stages to the entire data set. In both cases, multistage results were very similar to a full MCMC analysis. Supplementary materials accompanying this paper appear online.

Journal of Agricultural, Biological and Environmen

Evaluation of the Ott Hydromet Qliner for measuring discharge in laboratory and field conditions

The U.S. Geological Survey, in collaboration with the University of Iowa IIHR &ndash; Hydroscience and Engineering, evaluated the use of the Ott Hydromet Qliner using laboratory flume tests along with field validation tests. Analysis of the flume testing indicates the velocities measured by the Qliner at a 40-second exposure time results in higher dispersion of velocities from the mean velocity of data collected with a 5-minute exposure time. The percent data spread from the mean of a 100-minute mean of Qliner velocities for a 40-second exposure time averaged 16.6 percent for the entire vertical, and a 5-minute mean produced a 6.2 percent data spread from the 100-minute mean. This 16.6 percent variation in measured velocity would result in a 3.32 percent variation in computed discharge assuming 25 verticals while averaging 4 bins in each vertical. The flume testing also provided results that indicate the blanking distance of 0.20 meters is acceptable when using beams 1 and 2, however beam 3 is negatively biased near the transducer and the 0.20-meter blanking distance is not sufficient. Field testing included comparing the measured discharge by the Qliner to the discharge measured by a Price AA mechanical current meter and a Teledyne RDI Rio Grande 1200 kilohertz acoustic Doppler current profiler. The field tests indicated a difference between the discharges measured with the Qliner and the field reference discharge between -14.0 and 8.0 percent; however the average percent difference for all 22 field comparisons was 0.22, which was not statistically significant.

Iowa

Rainfall and runoff quantity and quality characteristics of four urban land-use catchments in Fresno, California, October 1981 to April 1983

Rainfall and runoff quantity and quality were monitored for industrial, single-dwelling residential, multiple-dwelling residential, and commercial land-use catchments during the 1981-82 and 1982-83 rain seasons. Storm-composite rainfall and discrete run6ff samples were analyzed for numerous inorganic, biological, physical, and organic constituents. Atmospheric dry-deposition and street-surface particulate samples also were collected and analyzed. With the exception of the industrial catchment, the highest runoff concentrations for most constituents occurred during the initial storm runoff and then decreased throughout the remainder of the storm, independent of hydraulic conditions. Metal concentrations were high during initial runoff, but also increased as flow increased. Constituent concentrations for the industrial catchment fluctuated greatly during storms. Statistical tests showed higher ammonia plus organic nitrogen, ammonia, pH, and phenol concentrations in rainfall at the industrial site than at the single-dwelling residential and laboratory sites. Statistical testing of runoff quality data showed higher concentrations for the industrial catchment than for the two residential and commercial catchments for most constituents. Total recoverable lead was one of the few constituents that had lower concentrations for the industrial catchment than for the other three catchments. The two residential catchments showed no significant difference in runoff concentrations for 50 of the 57 constituents used in the statistical analysis. The commercial catchment runoff concentrations for most constituents generally were similar to the residential catchments. Although constituent concentrations generally were higher for the industrial catchment than for the commercial catchment, constituent storm loads from the commercial catchment were similar to the industrial catchment because of the greater runoff volume from the highly impervious commercial catchment. Between 10 and 50 percent of the constituent runoff loads for the two residential catchments were attributed to the rainfall load, with the percentages generally considerably less for the industrial catchment. Event mean concentrations (EMC) for most constituents for all but the industrial catchment were highest for the first two or three storms of the rain season after which they became almost constant. Constituent event mean concentrations for the industrial catchment generally did not show any pattern throughout a rain season. Multiple-regression predictor equations for event mean concentrations were developed for several constituents for all sites. Average annual constituent unit loads were computed for 18 constituents for each catchment. The organophosphorus compounds, diazinon, malathion, and parathion were the most prevalent pesticides detected in rainfall. Diazinon was detected in all 54 rainfall samples. Parathion and malathion were detected in 49 and 50 samples, respectively. Other pesticides detected in rainfall included chlordane, lindane, methoxychlor, endosulfan, and 2,4-D. Of these, only methoxychlor and endosulfan were not consistently detected in runoff.

Water Supply Paper