Search USGSSearch

SEARCH · Search USGS

Results for “Computational Statistics and Data Analysis”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Estimation of stream conditions in tributaries of the Klamath River, northern California

Because of their critical ecological role, stream temperature and discharge are requisite inputs for models of salmonid population dynamics. Coho Salmon inhabiting the Klamath Basin spend much of their freshwater life cycle inhabiting tributaries, but environmental data are often absent or only seasonally available at these locations. To address this information gap, we constructed daily averaged water temperature models that used simulated meteorological data to estimate daily tributary temperatures, and we used flow differentials recorded on the mainstem Klamath River to estimate daily tributary discharge. Observed temperature data were available for fourteen of the major salmon bearing tributaries, which enabled estimation of tributary-specific model parameters at those locations. Water temperature data from six mid-Klamath Basin tributaries were used to estimate a global set of parameters for predicting water temperatures in the remaining tributaries. The resulting parameter sets were used to simulate water temperatures for each of 75 tributaries from 1980-2015. Goodness-of-fit statistics computed from a cross-validation analysis demonstrated a high precision of the tributary-specific models in predicting temperature in unobserved years and of the global model in predicting temperatures in unobserved streams. Klamath River discharge has been monitored by four gages that broadly intersperse the 292 kilometers from the Iron Gate Dam to the Klamath River mouth. These gages defined the upstream and downstream margins of three reaches. Daily discharge of tributaries within a reach was estimated from 1980-2015 based on drainage-area proportionate allocations of the discharge differential between the upstream and downstream margin. Comparisons with measured discharge on Indian Creek, a moderate-sized tributary with naturally regulated flows, revealed that the estimates effectively approximated both the variability and magnitude of discharge.

California

Guidelines for determining flood flow frequency: Bulletin #17B of the Hydrology Subcommittee

In December 1967, Bulletin No. 15, "A Uniform Technique for Determining Flood Flow Frequencies," was issued by the Hydrology Committee of the Water Resources Council. The report recommended use of the Pearson Type III distribution with log transformation of the data (log-Pearson Type III distribution) as a base method for flood flow frequency studies. As pointed out in that report, further studies were needed covering various aspects of flow frequency determinations. In March 1976, Bulletin 17, "Guidelines for Determining Flood Flow Frequency" was issued by the Water Resources Council. The guide was an extension and update of Bulletin No. 15. It provided a more complete guide for flood flow frequency analysis incorporating currently accepted technical methods with sufficient detail to promote uniform application. It was limited to defining flood potentials in terms of peak discharge and exceedance probability at locations where a systematic record of peak flood flows is available. The recommended set of procedures was selected from those used or described in the literature prior to 1976, based on studies conducted for this purpose at the Center for Research in Water Resources of the University of Texas at Austin (summarized in Appendix 14) and on studies by the Work Group on Flood Flow Frequency. The "Guidelines" were revised and reissued in June 1977 as Bulletin 17A. Bulletin 17B is the latest effort to improve and expand upon the earlier publications. Bulletin 17B provides revised procedures for weighting a station skew value with the results from a generalized skew study, detecting and treating outliers, making two station comparisons, and computing confidence limits about a frequency curve. The Work Group that prepared this revision did not address the suitability of the original distribution or the generalized skew map. Major problems are encountered when developing guides for flood flow frequency determinations. There is no procedure or set of procedures that can be adopted which, when rigidly applied to the available data, will accurately define the flood potential of any given watershed. Statistical analysis alone will not resolve all flood frequency problems. As discussed in subsequent sections of this guide, elements of risk and uncertainty are inherent in any flood frequency analysis. User decisions must be based on properly applied procedures and proper interpretation of results considering risk and uncertainty. Therefore, the judgment of a professional experienced in hydrologic analysis will enhance the usefulness of a flood frequency analysis and promote appropriate application. It is possible to standarize many elements of flood frequency analysis. This guide describes each major element of the process of defining the flood potential at a specific location in terms of peak discharge and exceedance probability. Use is confined to stations where available records are adequate to warrant statistical analysis of the data. Special situations may require other approaches. In those cases where the procedures of this guide are not followed, deviations must be supported by appropriate study and accompanied by a comparison of results using the recommended procedures. As a further means of achieving consistency and improving results, the Work Group recommends that studies be coordinated when more than one analyst is working currently on data for the same location. This recommendation holds particularly when defining exceedance probabilities for rare events, where this guide allows more latitude. Flood records are limited. As more years of record become available at each location, the determination of flood potential may change. Thus, an estimate may be outdated a few years after it is made. Additional flood data alone may be sufficient reason for a fresh assessment of the flood potential. When making a new assessment, the analyst should incorporate in his study a review of earlier estimates. Where differences appear, they should be acknowledged and explained.

Bulletin

Magnitude of flood flows for selected annual exceedance probabilities for streams in Massachusetts

The U.S. Geological Survey, in cooperation with the Massachusetts Department of Transportation, determined the magnitude of flood flows at selected annual exceedance prob­abilities (AEPs) at streamgages in Massachusetts and from these data developed equations for estimating flood flows at ungaged locations in the State. Flood magnitudes were deter­mined for the 50-, 20-, 10-, 4-, 2-, 1-, 0.5-, and 0.2-percent AEPs at 220 streamgages, 125 of which are in Massachusetts and 95 are in the adjacent States of Connecticut, New Hamp­shire, New York, Rhode Island, and Vermont. AEP flood flows were computed for streamgages using the expected moments algorithm weighted with a recently computed regional skew­ness coefficient for New England. Regional regression equations were developed to estimate the magnitude of floods for selected AEP flows at ungaged sites from 199 selected streamgages and for 60 potential explanatory basin characteristics. AEP flows for 21 of the 125 streamgages in Massachusetts were not used in the final regional regression analysis, primarily because of regulation or redundancy. The final regression equations used general­ized least squares methods to account for streamgage record length and correlation. Drainage area, mean basin elevation, and basin storage explained 86 to 93 percent of the variance in flood magnitude from the 50- to 0.2-percent AEPs, respec­tively. The estimates of AEP flows at streamgages can be improved by using a weighted estimate that is based on the magnitude of the flood and associated uncertainty from the at-site analysis and the regional regression equations. Weighting procedures for estimating AEP flows at an ungaged site on a gaged stream also are provided that improve estimates of flood flows at the ungaged site when hydrologic characteristics do not abruptly change. Urbanization expressed as the percentage of imperviousness provided some explanatory power in the regional regression; however, it was not statistically significant at the 95-percent confidence level for any of the AEPs examined. The effect of urbanization on flood flows indicates a complex interaction with other basin characteristics. Another complicating factor is the assumption of stationarity, that is, the assumption that annual peak flows exhibit no significant trend over time. The results of the analysis show that stationarity does not prevail at all of the streamgages. About 27 percent of streamgages in Massachusetts and about 42 percent of streamgages in adjacent States with 20 or more years of systematic record used in the study show a significant positive trend at the 95-percent confidence level. The remaining streamgages had both positive and negative trends, but the trends were not statistically significant. Trends were shown to vary over time. In particular, during the past decade (2004–2013), peak flows were persistently above normal, which may give the impression of positive trends. Only continued monitoring will provide the information needed to determine whether recent increases in annual peak flows are a normal oscillation or a true trend. The analysis used 37 years of additional data obtained since the last comprehensive study of flood flows in Massa­chusetts. In addition, new methods for computing flood flows at streamgages and regionalization improved estimates of flood magnitudes at gaged and ungaged locations and better defined the uncertainty of the estimates of AEP floods.

Massachusetts

A proposed streamflow data program for Oklahoma

An evaluation of the streamflow data available in Oklahoma has been made to provide guidelines for planning future data-collection programs. The basic steps in the evaluation procedure were (1) definition of the long-terms goals of the streamflow-data program in quantitative form, (2) examination and analysis of streamflow data to determine which goals have been met, and (3) consideration of alternate programs and techniques to meet the remaining goals. The study defines the individual relation between certain statistical streamflow characteristics and selected basin parameters. This relation is a multiple regression equation that could be used on a statewide basis to compute a selected natural-flow characteristic at any site on a stream. The study shows that several streamflow characteristics can be estimated within an accuracy equivalent to 10 years of record by use of a regression related to at least three climatic or basin parameters for any basin of 50 square miles or more. The study indicates that significant changes in the scope and character of the data-collection program would enhance the possibility of attaining the remaining goals. A streamflow-data program based on the guidelines developed in this study is proposed for the future.

Open-File Report

Augmenting two-dimensional hydrodynamic simulations with measured velocity data to identify flow paths as a function of depth on Upper St. Clair River in the Great Lakes basin

Upper St. Clair River, which receives outflow from Lake Huron, is characterized by flow velocities that exceed 7 feet per second and significant channel curvature that creates complex flow patterns downstream from the Blue Water Bridge in the Port Huron, Michigan, and Sarnia, Ontario, area. Discrepancies were detected between depth-averaged velocities previously simulated by a two-dimensional (2D) hydrodynamic model and surface velocities determined from drifting buoy deployments. A detailed ADCP (acoustic Doppler current profiler) survey was done on Upper St. Clair River during July 1–3, 2003, to help resolve these discrepancies. As part of this study, a refined finite-element mesh of the hydrodynamic model used to identify source areas to public water intakes was developed for Upper St. Clair River. In addition, a numerical procedure was used to account for radial accelerations, which cause secondary flow patterns near channel bends. The refined model was recalibrated to better reproduce local velocities measured in the ADCP survey. ADCP data also were used to help resolve the remaining discrepancies between simulated and measured velocities and to describe variations in velocity with depth. Velocity data from ADCP surveys have significant local variability, and statistical processing is needed to compute reliable point estimates. In this study, velocity innovations were computed for seven depth layers posited within the river as the differences between measured and simulated velocities. For each layer, the spatial correlation of velocity innovations was characterized by use of variogram analysis. Results were used with kriging to compute expected innovations within each layer at applicable model nodes. Expected innovations were added to simulated velocities to form integrated velocities, which were used with reverse particle tracking to identify the expected flow path near a sewage outfall as a function of flow depth. Expected particle paths generated by use of the integrated velocities showed that surface velocities in the upper layers tended to originate nearer the Canadian shoreline than velocities near the channel bottom in the lower layers. Therefore, flow paths to U.S. public water intakes located on the river bottom are more likely to be in the United States than withdrawals near the water surface. Integrated velocities in the upper layers are generally consistent with the surface velocities indicated by drifting-buoy deployments. Information in the 2D hydrodynamic model and the ADCP measurements was insufficient to describe the vertical flow component. This limitation resulted in the inability to account for vertical movements on expected flow paths through Upper St. Clair River. A three dimensional hydrodynamic model would be needed to account for these effects.

Scientific Investigations Report

Estimates of flow duration, mean flow, and peak-discharge frequency values for Kansas stream locations

Streamflow statistics of flow duration and peak-discharge frequency were estimated for 4,771 individual locations on streams listed on the 1999 Kansas Surface Water Register. These statistics included the flow-duration values of 90, 75, 50, 25, and 10 percent, as well as the mean flow value. Peak-discharge frequency values were estimated for the 2-, 5-, 10-, 25-, 50-, and 100-year floods. Least-squares multiple regression techniques were used, along with Tobit analyses, to develop equations for estimating flow-duration values of 90, 75, 50, 25, and 10 percent and the mean flow for uncontrolled flow stream locations. The contributing-drainage areas of 149 U.S. Geological Survey streamflow-gaging stations in Kansas and parts of surrounding States that had flow uncontrolled by Federal reservoirs and used in the regression analyses ranged from 2.06 to 12,004 square miles. Logarithmic transformations of climatic and basin data were performed to yield the best linear relation for developing equations to compute flow durations and mean flow. In the regression analyses, the significant climatic and basin characteristics, in order of importance, were contributing-drainage area, mean annual precipitation, mean basin permeability, and mean basin slope. The analyses yielded a model standard error of prediction range of 0.43 logarithmic units for the 90-percent duration analysis to 0.15 logarithmic units for the 10-percent duration analysis. The model standard error of prediction was 0.14 logarithmic units for the mean flow. Regression equations used to estimate peak-discharge frequency values were obtained from a previous report, and estimates for the 2-, 5-, 10-, 25-, 50-, and 100-year floods were determined for this report. The regression equations and an interpolation procedure were used to compute flow durations, mean flow, and estimates of peak-discharge frequency for locations along uncontrolled flow streams on the 1999 Kansas Surface Water Register. Flow durations, mean flow, and peak-discharge frequency values determined at available gaging stations were used to interpolate the regression-estimated flows for the stream locations where available. Streamflow statistics for locations that had uncontrolled flow were interpolated using data from gaging stations weighted according to the drainage area and the bias between the regression-estimated and gaged flow information. On controlled reaches of Kansas streams, the streamflow statistics were interpolated between gaging stations using only gaged data weighted by drainage area.

Kansas

Trends in groundwater levels in and near the Rosebud Indian Reservation, South Dakota, water years 1956–2017

The U.S. Geological Survey (USGS), in cooperation with the Rosebud Sioux Tribe, completed a study to characterize water-level fluctuations in observation wells to examine driving factors that affect water levels in and near the Rosebud Indian Reservation, which comprises all of Todd County. The study investigates concerns regarding potential effects of groundwater withdrawals and climate conditions on groundwater levels within an area that includes Todd County and a surrounding area that extends 10 miles north, east, and west of the county border. Characterization of water-level fluctuations in observation wells and relative driving factors was accomplished by statistical trend analysis. Two statistical methods were used for analysis of temporal trends for climatic and hydrologic data. To determine which trend analysis to use, applicable datasets were tested for statistically significant short-term persistence (STP). In the absence of significant STP, existence of statistical trends was determined using the standard Mann-Kendall test for probability values less than or equal to 0.10 (90-percent confidence level); however, a modified Mann-Kendall test was used for datasets where statistically significant STP was detected. Trend magnitudes were computed using the Sen’s slope estimator. Monthly data from the Parameter-elevation Regressions on Independent Slopes Model (PRISM) were aggregated to obtain annual and seasonal datasets for total precipitation, minimum air temperature ( T min ), and maximum air temperature ( T max ) for the study area and a surrounding buffer area. Trend tests for total precipitation, T min , and T max were completed for annual and seasonal time series for water years 1956–2017, which is about 2 years before the earliest available water-level measurements. A 2-year offset was arbitrarily selected because scrutiny of water-level and precipitation data indicated that responses of groundwater levels for many of the observation wells lagged major changes in precipitation patterns by about 2 years. Statistically significant upward trends were detected for annual precipitation and annual T min for almost all of the study area and the surrounding buffer area. Statistically significant downward trends in T max were detected for a very small part of the study area; however, the sparse spatial coverage reduces confidence that these are true trends. Spatial distributions of statistically significant trends in seasonal climate data were generally similar to the annual trends, but with substantial differences in the spatial density of the trends. Groundwater trends for 58 observation wells were analyzed for three separate water-level parameters (minimum, median, and maximum) because wells are measured sporadically and data are biased towards more frequent measurements during periods of heaviest irrigation demand. Trends in the time series of annual precipitation (from PRISM) starting 2 years earlier than for the associated water-level trend also were analyzed for the location of each individual observation well. Sen’s slope and Mann-Kendall probability values (p-values) were computed for the three water-level parameters and for the annual precipitation time series. Graphs showing results of trend analyses for each observation well also showed changes over time in the sum of licensed groundwater withdrawals within six specified radii (0.5, 1, 2, 3, 4, and 5 miles) of each well as a qualitative indicator of proximal groundwater demand. Of all 58 observation wells considered, 28 wells had significant upward trends for at least one of the three water-level parameters, 11 wells had significant downward trends for at least one water-level parameter, and 19 wells did not have any significant trends. Significant upward trends in annual precipitation were detected for 48 of the 58 wells. Results of trend analyses likely show the effects of groundwater withdrawals on water levels in the Ogallala aquifer in areas of substantial demand. Precipitation trends are significantly upward for 43 of the 48 wells completed in the Ogallala aquifer that were analyzed. Of the 48 Ogallala aquifer wells, 24 had significant upward trends for at least one water-level parameter (17 with all 3); however, 10 wells had statistically significant downward trends for at least one water-level parameter (8 with all 3 parameters). All but one of the wells with significant downward trends are located in the south-central part of the study area where licensed irrigation withdrawals are concentrated.

South Dakota

Sensitivity analysis, calibration, and testing of a distributed hydrological model using error‐based weighting and one objective function

We evaluate the utility of three interrelated means of using data to calibrate the fully distributed rainfall‐runoff model TOPKAPI as applied to the Maggia Valley drainage area in Switzerland. The use of error‐based weighting of observation and prior information data, local sensitivity analysis, and single‐objective function nonlinear regression provides quantitative evaluation of sensitivity of the 35 model parameters to the data, identification of data types most important to the calibration, and identification of correlations among parameters that contribute to nonuniqueness. Sensitivity analysis required only 71 model runs, and regression required about 50 model runs. The approach presented appears to be ideal for evaluation of models with long run times or as a preliminary step to more computationally demanding methods. The statistics used include composite scaled sensitivities, parameter correlation coefficients, leverage, Cook's D, and DFBETAS. Tests suggest predictive ability of the calibrated model typical of hydrologic models.

Water Resources Research

Adjusted geomagnetic data—Theoretical basis and validation

Adjusted geomagnetic data are magnetometer measurements with provisional correction factors applied such that vector quantities are oriented in a local Cartesian frame in which the X axis points north, the Y axis points east, and the Z axis points down. These correction factors are determined from so-called absolute measurements, which are “ground truth” observations made in the field using specialized magnetometers and survey equipment that are (nearly) colocated with the automated and continuously running magnetic measurement instrumentation. Correction factors can be substantial, up to hundreds of nanoTeslas, depending on the geologic and geomagnetic characteristics of the observatory site. They also tend to evolve over time because of instrument response instability and changing site characteristics. Historically, correction factors were determined offline, up to 1 year or more post-measurement, and applied to raw measurements to produce “Definitive” data for scientific analysis. Growing demand for corrected real-time geomagnetic data to better support space weather operations motivated development of an “Adjusted” geomagnetic data product. Modern computational tools, and some notable practical concerns, dictated a transition to affine transformations in lieu of more traditional baseline corrections, as well as a calibration parameter estimation algorithm that is more robust and statistically optimal, and therefore better suited for automated and unsupervised execution. A theoretical basis for this algorithm is presented, along with a demonstration and validation based on a comparison of results obtained with traditional techniques. Discrepancies between Definitive corrected data and near real-time Adjusted data obtained using affine transformations are minimal, generally much less than 5 nanoTeslas per vector component, and less than 1 nanoTesla for the total field magnitude, which satisfies International Real-Time Magnetic Observatory Network (INTERMAGNET) standards.

Open-File Report

Spectral analysis of aeromagnetic profiles for depth estimation principles, software, and practical application

Fourier spectral analysis in recent years has become a widely utilized tool for the processing and interpretation of potential field data. It is particularly well suited to analysis of aeromagnetic maps and profiles, where coverage commonly is of broad scope and statistical treatment is appropriate. The techniques developed by earlier workers for map data are readily adapted for depth estimates using aeromagnetic profiles. Three subroutines are presented: "FRQAN", which employs the complete Fourier transform to convert field intensity to the frequency domain and then computes the logarithmic energy spectrum; "CSIZE", which refines the spectrum to correct for the finite horizontal dimensions of magnetic sources, and "ENSMTH", which smooths the spectrum to clarify its decay characteristics. The average depths to sources of ensembles are obtained by manually fitting a straight line to each linear interval of the logarithmic energy-decay curve. If proper account is taken of the constraints of the method, it is capable of providing depth estimates to within an accuracy of about 10 percent under suitable circumstances. The estimates are unaffected by source magnetization and are relatively insensitive to assumptions as to source shape or distribution. The validity of the method is demonstrated by analyses of synthetic profiles and profiles recorded over Harrat Rahat, Saudi Arabia, and Diyur, Egypt, where source depths have been proved by drilling.

Open-File Report

Streambed stability and scour potential at selected bridge sites in Michigan

Contraction scour in the main stream channel at a bridge and local scour near piers and abutments can result in bridge failure. Estimates of contraction-scour and local-scour potentials associated with the 100-year flood were computed for 13 bridge sites in Michigan by use of semi-theoretical equations and procedures recommended by the Federal Highway Administration. These potentials were compared with measures of Streambed stability obtained by use of data from 773 historical streamflow measurements, documenting 20,741 individual Streambed soundings between 1959 and 1995. Analysis of these data indicate small, but statistically significant, monotonic trends in Streambed elevation at 10 sites. No consistent patterns in relations between changes in Streambed elevations and streamflow, flow velocity, or flow depth were evident. Also, estimates of contraction-scour potential were not correlated with measures of Streambed stability, and no differences were detected between measures of Streambed stability in the main channel and stability adjacent to piers. Despite the inconsistencies between measures of Streambed stability and scour potential, data from a single, large flood (greater than a 100-year event) provided field evidence that the relation between scour and streamflow is highly nonlinear. This nonlinearity and the limited availability of measurements of extreme flood events may have reduced the utility of the empirical measures for confirming the nonlinear scour-potential equations and procedures. Results of field surveys using ground-penetrating radar and tuned transducers showed limited ability to aid interpretation of historical scour conditions at four bridge sites. Additional research is needed to confirm the applicability of scour-potential equations for hydrogeologic conditions in Michigan.

Michigan

Statistical analysis of water-quality data containing multiple detection limits II: S-language software for nonparametric distribution modeling and hypothesis testing

Analysis of low concentrations of trace contaminants in environmental media often results in left-censored data that are below some limit of analytical precision. Interpretation of values becomes complicated when there are multiple detection limits in the data-perhaps as a result of changing analytical precision over time. Parametric and semi-parametric methods, such as maximum likelihood estimation and robust regression on order statistics, can be employed to model distributions of multiply censored data and provide estimates of summary statistics. However, these methods are based on assumptions about the underlying distribution of data. Nonparametric methods provide an alternative that does not require such assumptions. A standard nonparametric method for estimating summary statistics of multiply-censored data is the Kaplan-Meier (K-M) method. This method has seen widespread usage in the medical sciences within a general framework termed "survival analysis" where it is employed with right-censored time-to-failure data. However, K-M methods are equally valid for the left-censored data common in the geosciences. Our S-language software provides an analytical framework based on K-M methods that is tailored to the needs of the earth and environmental sciences community. This includes routines for the generation of empirical cumulative distribution functions, prediction or exceedance probabilities, and related confidence limits computation. Additionally, our software contains K-M-based routines for nonparametric hypothesis testing among an unlimited number of grouping variables. A primary characteristic of K-M methods is that they do not perform extrapolation and interpolation. Thus, these routines cannot be used to model statistics beyond the observed data range or when linear interpolation is desired. For such applications, the aforementioned parametric and semi-parametric methods must be used.

Computers & Geosciences

Review of the Water Resources Information System of Argentina

A representative of the U.S. Geological Survey traveled to Buenos Aires, Argentina, in November 1986, to discuss water information systems and data bank implementation in the Argentine Government Center for Water Resources Information. Software has been written by Center personnel for a minicomputer to be used to manage inventory (index) data and water quality data. Additional hardware and software have been ordered to upgrade the existing computer. Four microcomputers, statistical and data base management software, and network hardware and software for linking the computers have also been ordered. The Center plans to develop a nationwide distributed data base for Argentina that will include the major regional offices as nodes. Needs for continued development of the water resources information system for Argentina were reviewed. Identified needs include: (1) conducting a requirements analysis to define the content of the data base and insure that all user requirements are met, (2) preparing a plan for the development, implementation, and operation of the data base, and (3) developing a conceptual design to inform all development personnel and users of the basic functionality planned for the system. A quality assurance and configuration management program to provide oversight to the development process was also discussed. (USGS)

Open-File Report

Estimator selection for closed-population capture: recapture

For valid statistical inference, it is important to select an appropriate statistical model. In the analysis of capture-recapture data under the closed-population models of Otis et al. (1978), information theoretic and hypothesis testing approaches to model selection are not practical, because some of the models have likelihoods with nonidenti- fiable parameters. A further problem is that, for some of the Otis et al. models, multiple estimators exist but there is no objective basis for deciding which estimator to use for a particular dataset. In CAPTURE, a computer program for estimating parameters un- der the closed models of Otis et al., a linear discriminant classifier is used to select an appropriate model. This classifier frequently selects the incorrect generating model in simulation studies, and it provides no guidance on which estimator to use once a model has been selected. In this study, we develop new classifiers for selecting the best esti- mator (as opposed to the generating model) and evaluate their performance. In addition, we investigate an estimator averaging approach to estimation that is a modification of the model averaging approach described by Buckland et al. (1997). We found that, in general, the overall performance of the new classifiers was unimpressive. In contrast, the estimator averaging approach we investigated performed well.

Journal of Agricultural, Biological, and Environme

Comparison of two U.S. power-plant carbon dioxide emissions data sets

Estimates of fossil-fuel CO2 emissions are needed to address a variety of climate-change mitigation concerns over a broad range of spatial and temporal scales. We compared two data sets that report power-plant CO 2 emissions in the conterminous U.S. for 2004, the most recent year reported in both data sets. The data sets were obtained from the Department of Energy's Energy Information Administration (EIA) and the Environmental Protection Agency's eGRID database. Conterminous U.S. total emissions computed from the data sets differed by 3.5% for total plant emissions (electricity plus useful thermal output) and 2.3% for electricity generation only. These differences are well within previous estimates of uncertainty in annual U.S. fossil-fuel emissions. However, the corresponding average absolute differences between estimates of emissions from individual power plants were much larger, 16.9% and 25.3%, respectively. By statistical analysis, we identified several potential sources of differences between EIA and eGRID estimates for individual plants. Estimates that are based partly or entirely on monitoring of stack gases (reported by eGRID only) differed significantly from estimates based on fuel consumption (as reported by EIA). Differences in accounting methods appear to explain differences in estimates for emissions from electricity generation from combined heat and power plants, and for total and electricity generation emissions from plants that burn nonconventional fuels (e.g., biomass). Our analysis suggests the need for care in utilizing emissions data from individual power plants, and the need for transparency in documenting the accounting and monitoring methods used to estimate emissions.

Environmental Science & Technology

Geochemical modeling of water-rock interactions in mining environments

Geochemical modeling is a powerful tool for evaluating geochemical processes in mining environments. Properly constrained and judiciously applied, modeling can provide valuable insights into processes controlling the release, transport, and fate of contaminants in mine drainage. This chapter contains 1) an overview of geochemical modeling, 2) discussion of the types of models and computer programs used, 3) description of a procedure for screening water analyses for modeling input, and 4) examples of the application of modeling for interpreting geochemical processes in mining environments. Three general strategies in current use to interpret water-rock interactions are statistical analysis, “inverse” modeling, and “forward” modeling. Multivariate correlation analysis, factor analysis, cluster analysis, and other statistical techniques can group water-chemistry data into sets that may relate to hydrogeochemical processes (Drever, 1988; Puckett and Bricker, 1992). In the field of geochemical exploration, statistical analysis is used widely to treat large data sets of rock and sediment chemistry (e.g., Garrett, 1989). No physical or chemical principles are involved directly in these statistical treatments, hence they are not considered further in this chapter. Nevertheless, statistical analysis can be a useful tool in organizing complex geochemical data for interpretation.

Book chapter

Guidelines and Procedures for Computing Time-Series Suspended-Sediment Concentrations and Loads from In-Stream Turbidity-Sensor and Streamflow Data

In-stream continuous turbidity and streamflow data, calibrated with measured suspended-sediment concentration data, can be used to compute a time series of suspended-sediment concentration and load at a stream site. Development of a simple linear (ordinary least squares) regression model for computing suspended-sediment concentrations from instantaneous turbidity data is the first step in the computation process. If the model standard percentage error (MSPE) of the simple linear regression model meets a minimum criterion, this model should be used to compute a time series of suspended-sediment concentrations. Otherwise, a multiple linear regression model using paired instantaneous turbidity and streamflow data is developed and compared to the simple regression model. If the inclusion of the streamflow variable proves to be statistically significant and the uncertainty associated with the multiple regression model results in an improvement over that for the simple linear model, the turbidity-streamflow multiple linear regression model should be used to compute a suspended-sediment concentration time series. The computed concentration time series is subsequently used with its paired streamflow time series to compute suspended-sediment loads by standard U.S. Geological Survey techniques. Once an acceptable regression model is developed, it can be used to compute suspended-sediment concentration beyond the period of record used in model development with proper ongoing collection and analysis of calibration samples. Regression models to compute suspended-sediment concentrations are generally site specific and should never be considered static, but they represent a set period in a continually dynamic system in which additional data will help verify any change in sediment load, type, and source.

Techniques and Methods

Reconstruction of spatio-temporal temperature from sparse historical records using robust probabilistic principal component regression

Scientific records of temperature and precipitation have been kept for several hundred years, but for many areas, only a shorter record exists. To understand climate change, there is a need for rigorous statistical reconstructions of the paleoclimate using proxy data. Paleoclimate proxy data are often sparse, noisy, indirect measurements of the climate process of interest, making each proxy uniquely challenging to model statistically. We reconstruct spatially explicit temperature surfaces from sparse and noisy measurements recorded at historical United States military forts and other observer stations from 1820 to 1894. One common method for reconstructing the paleoclimate from proxy data is principal component regression (PCR). With PCR, one learns a statistical relationship between the paleoclimate proxy data and a set of climate observations that are used as patterns for potential reconstruction scenarios. We explore PCR in a Bayesian hierarchical framework, extending classical PCR in a variety of ways. First, we model the latent principal components probabilistically, accounting for measurement error in the observational data. Next, we extend our method to better accommodate outliers that occur in the proxy data. Finally, we explore alternatives to the truncation of lower-order principal components using different regularization techniques. One fundamental challenge in paleoclimate reconstruction efforts is the lack of out-of-sample data for predictive validation. Cross-validation is of potential value, but is computationally expensive and potentially sensitive to outliers in sparse data scenarios. To overcome the limitations that a lack of out-of-sample records presents, we test our methods using a simulation study, applying proper scoring rules including a computationally efficient approximation to leave-one-out cross-validation using the log score to validate model performance. The result of our analysis is a spatially explicit reconstruction of spatio-temporal temperature from a very sparse historical record.

Advances in Statistical Climatology, Meteorology a