Search USGSSearch

SEARCH · Search USGS

Results for “Statistical Methods & Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

A novel approach to assessing natural resource injury with Bayesian networks

Quantifying the effects of environmental stressors on natural resources is problematic because of complex interactions among environmental factors that influence endpoints of interest. This complexity, coupled with data limitations, propagates uncertainty that can make it difficult to causally associate specific environmental stressors with injury endpoints. The Natural Resource Damage Assessment and Restoration (NRDAR) regulations under the Comprehensive Environmental Response, Compensation, and Liability Act and Oil Pollution Act aim to restore natural resources injured by oil spills and hazardous substances released into the environment; exploration of alternative statistical methods to evaluate effects could help address NRDAR legal claims. Bayesian networks (BNs) are statistical tools that can be used to estimate the influence and interrelatedness of abiotic and biotic environmental variables on environmental endpoints of interest. We investigated the application of a BN for injury assessment using a hypothetical case study by simulating data of acid mine drainage (AMD) affecting a fictional stream-dwelling bird species. We compared the BN-generated probability estimates for injury with a more traditional approach using toxicity thresholds for water and sediment chemistry. Bayesian networks offered several distinct advantages over traditional approaches, including formalizing the use of expert knowledge, probabilistic estimates of injury using intermediate direct and indirect effects, and the incorporation of a more nuanced and ecologically relevant representation of effects. Given the potential that BNs have for natural resource injury assessment, more research and field-based application are needed to determine their efficacy in NRDAR. We expect the resulting methods will be of interest to many US federal, state, and tribal programs devoted to the evaluation, mitigation, remediation, and/or restoration of natural resources injured by releases or spills of contaminants

Integrated Environmental Assessment and Management

Probabilistic tsunami hazard analysis: Multiple sources and global applications

Applying probabilistic methods to infrequent but devastating natural events is intrinsically challenging. For tsunami analyses, a suite of geophysical assessments should be in principle evaluated because of the different causes generating tsunamis (earthquakes, landslides, volcanic activity, meteorological events, and asteroid impacts) with varying mean recurrence rates. Probabilistic Tsunami Hazard Analyses (PTHAs) are conducted in different areas of the world at global, regional, and local scales with the aim of understanding tsunami hazard to inform tsunami risk reduction activities. PTHAs enhance knowledge of the potential tsunamigenic threat by estimating the probability of exceeding specific levels of tsunami intensity metrics (e.g., run-up or maximum inundation heights) within a certain period of time (exposure time) at given locations (target sites); these estimates can be summarized in hazard maps or hazard curves. This discussion presents a broad overview of PTHA, including (i) sources and mechanisms of tsunami generation, emphasizing the variety and complexity of the tsunami sources and their generation mechanisms, (ii) developments in modeling the propagation and impact of tsunami waves, and (iii) statistical procedures for tsunami hazard estimates that include the associated epistemic and aleatoric uncertainties. Key elements in understanding the potential tsunami hazard are discussed, in light of the rapid development of PTHA methods during the last decade and the globally distributed applications, including the importance of considering multiple sources, their relative intensities, probabilities of occurrence, and uncertainties in an integrated and consistent probabilistic framework.

Reviews of Geophysics

Evaluating the reliability of environmental concentration data to characterize exposure in environmental risk assessments

Environmental risk assessments often rely on measured concentrations in environmental matrices to characterize exposure of the population of interest—typically, humans, aquatic biota, or other wildlife. Yet, there is limited guidance available on how to report and evaluate exposure datasets for reliability and relevance, despite their importance to regulatory decision-making. This paper is the second of a four-paper series detailing the outcomes of a Society of Environmental Toxicology and Chemistry Technical Workshop that has developed Criteria for Reporting and Evaluating Exposure Datasets (CREED). It presents specific criteria to systematically evaluate the reliability of environmental exposure datasets. These criteria can help risk assessors understand and characterize uncertainties when existing data are used in various types of assessments and can serve as guidance on best practice for the reporting of data for data generators (to maximize utility of their datasets). Although most reliability criteria are universal, some practices may need to be evaluated considering the purpose of the assessment. Reliability refers to the inherent quality of the dataset and evaluation criteria address the identification of analytes, study sites, environmental matrices, sampling dates, sample collection methods, analytical method performance, data handling or aggregation, treatment of censored data, and generation of summary statistics. Each criterion is evaluated as “fully met,” “partly met,” “not met or inappropriate,” “not reported,” or “not applicable” for the dataset being reviewed. The evaluation concludes with a scheme for scoring the dataset as reliable with or without restrictions, not reliable, or not assignable, and is demonstrated with three case studies representing both organic and inorganic constituents, and different study designs and assessment purposes. Reliability evaluation can be used in conjunction with relevance evaluation (assessed separately) to determine the extent to which environmental monitoring datasets are “fit for purpose,” that is, suitable for use in various types of assessments. Integr Environ Assess Manag 2024;00:1–23. © 2024 Society of Environmental Toxicology & Chemistry (SETAC). This article has been contributed to by U.S. Government employees and their work is in the public domain in the USA.

Integrated Environmental Assessment and Management

Revisiting the declustering of spatial data with preferential sampling

Preferential sampling is a form of data collection that may significantly distort the histogram and the semivariogram of spatially correlated data . Typical situations are a higher sampling density at high-valued areas favorable for mining, and highly contaminated areas in need of environmental remediation. Multiple statistical procedures are devoted to obtaining representative statistics, whose magnitudes should be close to the respective population values. This paper proposes a resampling method that can compensate for preferential sampling of spatially correlated data without using declustering weights. The application of the method herein generates a dataset of median estimates of quantiles of multiple stratified resamples that is free of preferential sampling. The methodology is illustrated with two examples. The first one involves values actually measured in the field and has the advantage of representing a real scenario of spatial fluctuations and preferential sampling. A second dataset is synthetic and has the main benefit of a priori knowledge of the underlying spatial distribution, thus allowing a satisfactory evaluation of the results against the known baseline. Access to computer code is offered for practical application of the method.

Computers & Geosciences

Regional analysis of the dependence of peak-flow quantiles on climate with application to adjustment to climate trends

Standard flood-frequency analysis methods rely on an assumption of stationarity, but because of growing understanding of climatic persistence and concern regarding the effects of climate change, the need for methods to detect and model nonstationary flood frequency has become widely recognized. In this study, a regional statistical method for estimating the effects of climate variations on annual maximum (peak) flows that allows for the effect to vary by quantile is presented and applied. The method uses a panel–quantile regression framework based on a location-scale model with two fixed effects per basin. The model was fitted to 330 selected gauged basins in the midwestern United States, filtered to remove basins affected by reservoir regulation and urbanization. Precipitation and discharge simulated using a water-balance model at daily and annual time scales were tested as climate variables. Annual maximum daily discharge was found to be the best predictor of peak flows, and the quantile regression coefficients were found to depend monotonically on annual exceedance probability. Application of the models to gauged basins is demonstrated by estimating the peak-flow distributions at the end of the study period (2018) and, using the panel model, to the study basins as-if-ungauged by using leave-one-out cross validation, estimating the fixed effects using static basin characteristics, and parameterizing the water-balance model discharge using median parameters. The errors of the quantiles predicted as-if-ungauged approximately doubled compared to the errors of the fitted panel model.

Hydrology

Generalized linear and generalized additive models in studies of species distributions: Setting the scene

An important statistical development of the last 30 years has been the advance in regression analysis provided by generalized linear models (GLMs) and generalized additive models (GAMs). Here we introduce a series of papers prepared within the framework of an international workshop entitled: Advances in GLMs/GAMs modeling: from species distribution to environmental management, held in Riederalp, Switzerland, 6-11 August 2001. We first discuss some general uses of statistical models in ecology, as well as provide a short review of several key examples of the use of GLMs and GAMs in ecological modeling efforts. We next present an overview of GLMs and GAMs, and discuss some of their related statistics used for predictor selection, model diagnostics, and evaluation. Included is a discussion of several new approaches applicable to GLMs and GAMs, such as ridge regression, an alternative to stepwise selection of predictors, and methods for the identification of interactions by a combined use of regression trees and several other approaches. We close with an overview of the papers and how we feel they advance our understanding of their application to ecological modeling. ?? 2002 Elsevier Science B.V. All rights reserved.

Ecological Modelling

Methods of Fitting a Straight Line to Data: Examples in Water Resources

Three methods of fitting straight lines to data are described and their purposes are discussed and contrasted in terms of their applicability in various water resources contexts. The three methods are ordinary least squares (OLS), least normal squares (LNS), and the line of organic correlation (OC). In all three methods the parameters are based on moment statistics of the data. When estimation of an individual value is the objective, OLS is the most appropriate. When estimation of many values is the objective and one wants the set of estimates to have the appropriate variance, then OC is most appropriate. When one wishes to describe the relationship between two variables and measurement error is unimportant, then OC is most appropriate. Where the error is important in descriptive problems or in calibration problems, then structural analysis techniques may be most appropriate. Finally, if the problem is one of describing some geographic trajectory, then LNS is most appropriate.

Water Resources Bulletin

Frequency-duration analysis of dissolved-oxygen concentrations in two southwestern Wisconsin streams

Historically, dissolved-oxygen (DO) data have been collected in the same manner as other water-quality constituents, typically at infrequent intervals as a grab sample or an instantaneous meter reading. Recent years have seen an increase in continuous water-quality monitoring with electronic dataloggers. This new technique requires new approaches in the statistical analysis of the continuous record. This paper presents an application of frequency-duration analysis to the continuous DO records of a cold and a warm water stream in rural southwestern Wisconsin. This method offers a quick, concise way to summarize large time-series data bases in an easily interpretable manner. Even though the two streams had similar mean DO concentrations, frequency-duration analyses showed distinct differences in their DO-concentration regime. This type of analysis also may be useful in relating DO concentrations to biological effects and in predicting low DO occurrences.

Wisconsin

Heavy mineral prospecting in Pakistan

Heavy-mineral prospecting is a method of tracing potential ore minerals to their source by systematically examining the heavy minerals of stream sands. The heavy minerals are partially concentrated in the field by panning. In the laboratory the heavy minerals are further concentrated by floating off the quartz and feldspar in a heavy liquid, such as bromoform, having a specific gravity of 2.89. Magnetite is separated magnetically and scheelite is identified under short-wave ultraviolet light. To obtain quantitative results on the minerals present, the sample is reduced by coning and quartering, mounted on a slide in oil, and a count of a minimum of 250 grains made. The chi square test applied to the splitting and counting techniques shows them to be statistically valid. The properties of some heavy minerals most useful for their rapid identification are listed. Methods here described are those utilizing a minimum of field and laboratory equipment and may be appropriate for application in many areas of the world.

Open-File Report

A controlled experiment in ground water flow model calibration

Nonlinear regression was introduced to ground water modeling in the 1970s, but has been used very little to calibrate numerical models of complicated ground water systems. Apparently, nonlinear regression is thought by many to be incapable of addressing such complex problems. With what we believe to be the most complicated synthetic test case used for such a study, this work investigates using nonlinear regression in ground water model calibration. Results of the study fall into two categories. First, the study demonstrates how systematic use of a well designed nonlinear regression method can indicate the importance of different types of data and can lead to successive improvement of models and their parameterizations. Our method differs from previous methods presented in the ground water literature in that (1) weighting is more closely related to expected data errors than is usually the case; (2) defined diagnostic statistics allow for more effective evaluation of the available data, the model, and their interaction; and (3) prior information is used more cautiously. Second, our results challenge some commonly held beliefs about model calibration. For the test case considered, we show that (1) field measured values of hydraulic conductivity are not as directly applicable to models as their use in some geostatistical methods imply; (2) a unique model does not necessarily need to be identified to obtain accurate predictions; and (3) in the absence of obvious model bias, model error was normally distributed. The complexity of the test case involved implies that the methods used and conclusions drawn are likely to be powerful in practice.

Groundwater

Estimating post-fire debris-flow hazards prior to wildfire using a statistical analysis of historical distributions of fire severity from remote sensing data

Following wildfire, mountainous areas of the western United States are susceptible to debris flow during intense rainfall. Convective storms that can generate debris flows in recently burned areas may occur during or immediately after the wildfire, leaving insufficient time for development and implementation of risk mitigation strategies. We present a method for estimating post-fire debris-flow hazards prior to wildfire using historical data to define the range of potential fire severities for a given location based on the statistical distribution of severity metrics obtained from remote sensing. Estimates of debris-flow likelihood, magnitude, and triggering rainfall threshold based upon the statistically simulated fire severity data provide hazard predictions consistent with those calculated from fire severity data collected after wildfire. Simulated fire severity data also produce hazard estimates that replicate observed debris-flow occurrence, rainfall conditions, and magnitude at a monitored site in the San Gabriel Mountains of southern California. Future applications of this method should rely upon a range of potential fire severity scenarios for improved pre-fire estimates of debris-flow hazard. The method presented here is also applicable to modeling other post-fire hazards, such as flooding and erosion risk, and for quantifying trends in observed fire severity in a changing climate.

California, Colorado, Washington

Of bugs and birds: Markov Chain Monte Carlo for hierarchical modeling in wildlife research

Markov chain Monte Carlo (MCMC) is a statistical innovation that allows researchers to fit far more complex models to data than is feasible using conventional methods. Despite its widespread use in a variety of scientific fields, MCMC appears to be underutilized in wildlife applications. This may be due to a misconception that MCMC requires the adoption of a subjective Bayesian analysis, or perhaps simply to its lack of familiarity among wildlife researchers. We introduce the basic ideas of MCMC and software BUGS (Bayesian inference using Gibbs sampling), stressing that a simple and satisfactory intuition for MCMC does not require extraordinary mathematical sophistication. We illustrate the use of MCMC with an analysis of the association between latent factors governing individual heterogeneity in breeding and survival rates of kittiwakes ( Rissa tridactyla ). We conclude with a discussion of the importance of individual heterogeneity for understanding population dynamics and designing management plans.

Journal of Wildlife Management

Peak-flow and low-flow magnitude estimates at defined frequencies and durations for nontidal streams in Delaware

Reliable estimates of the magnitude of peak flows in streams are required for the economical and safe design of transportation and water conveyance structures. In addition, reliable estimates of the magnitude of low flows at defined frequencies and durations are needed for meeting regulatory requirements, quantifying base flows in streams and rivers, and evaluating time of travel and dilution of toxic spills. This report, in cooperation with the Delaware Department of Transportation and the Delaware Geological Survey, presents methods for estimating the magnitude of peak flows and low flows at defined frequencies and durations on nontidal streams in Delaware, at locations both monitored by streamflow-gage sites and ungaged. Methods are presented for estimating (1) the magnitude of peak flows for return periods ranging from 2 to 500 years (50-percent to 0.2-percent annual-exceedance probability), and (2) the magnitude of low flows as applied to 7-, 14-, and 30-consecutive day low-flow periods with recurrence intervals of 2, 10, and 20 years (50-, 10-, and 5-percent annual non-exceedance probabilities). These methods are applicable to watersheds that exhibit a full range of development conditions in Delaware. The report also describes StreamStats, a web application that allows users to easily obtain peak-flow and low-flow magnitude estimates for user-selected locations in Delaware. Peak-flow and low-flow magnitude estimates for ungaged sites are obtained using statistical regression analysis through a process known as regionalization, where information from a group of streamflow-gage sites within a region forms the basis for estimates for ungaged sites within the same region. Ninety-four streamflow-gage sites in and near Delaware with at least 10 years of nonregulated annual peak-flow data were used for the peak-flow regression analysis, a subset of the 121 sites for which peak-flow estimates were computed. These sites included both continuous-record streamflow-gage sites as well as partial record sites. Forty-five streamflow-gage sites with at least 10 years of nonregulated low-flow data available were used for the low-flow regression analyses, a subset of the 68 sites for which low-flow estimates were computed. Estimates for gaged sites are obtained by combining (1) the station peak-flow statistics (mean, standard deviation, and skew) and peak-flow estimates using the recent Bulletin 17C guidelines that incorporate the Expected Moments Algorithm with (2) regional estimates of peak-flow magnitude derived from regional regression equations and regional skew derived from sites with records greater than or equal to 35 years. Example peak-flow estimate calculations using the methods presented in the report are given for (1) ungaged sites, (2) gaged sites, (3) sites upstream or downstream from a gaged location, and (4) sites between gaged locations. Estimates for low-flow gaged sites are obtained by combining (1) the station low-flow statistics (mean, standard deviation, and skew) and low-flow estimates with (2) regional estimates of low-flow magnitude derived from regional regression equations. Example low-flow estimate calculations using the methods presented in the report are given for (1) ungaged sites, (2) gaged sites, (3) sites upstream or downstream from a gaged location, and (4) sites between gaged locations. A total of 54 sites in the Coastal Plain region were used to develop peak-flow regressions for the region and 40 sites were used for the Piedmont region. Similarly, 24 sites were used for low-flow regression equation development in the Coastal Plain, with 21 in the Piedmont. Peak and low-flow site inclusion in the Coastal Plain tended to be more restricted with tidal influence and ranges of basin characteristics, including drainage area, limiting regression equation development and application. Regional regression equations for peak flows and low flows, as applicable to ungaged sites in the Piedmont and Coastal Plain Physiographic Provinces in Delaware, are presented. Peak-flow regression equations used variables that quantified drainage area, basin slope, percent area with well-drained soils, percent area with poorly drained soils, impervious area, and percent area of surface water storage in estimating peak-flow estimates, whereas low-flow regression equations used only drainage area and percent poorly drained soils in the estimation of low flows. Average standard errors for peak-flow regressions tended to be lower than those for low- flow regressions, with lower errors in the Piedmont region for both peak- and low-flow regressions. For peak-flow estimates, a sensitivity analysis of Piedmont regression equation estimates to changes in impervious area is also presented. Additional topics associated with the analyses performed during the study are discussed, including (1) the availability and description of 32 basin and climatic characteristics considered during the development of the regional regression equations; (2) the treatment of increasing trends in the annual peak-flow series identified at 18 gaged sites and inclusion in or exclusion from the regional analysis; (3) regional skew analysis and determination of regression regions; (4) sample adjustments and removal of sites owing to regulation and redundancy; and (5) a brief comparison of peak- and low-flow estimates at gages used in previous studies.

Delaware

Parameter estimation for the 4-parameter Asymmetric Exponential Power distribution by the method of L-moments using R

The implementation characteristics of two method of L-moments (MLM) algorithms for parameter estimation of the 4-parameter Asymmetric Exponential Power (AEP4) distribution are studied using the R environment for statistical computing. The objective is to validate the algorithms for general application of the AEP4 using R. An algorithm was introduced in the original study of the L-moments for the AEP4. A second or alternative algorithm is shown to have a larger L-moment-parameter domain than the original. The alternative algorithm is shown to provide reliable parameter production and recovery of L-moments from fitted parameters. A proposal is made for AEP4 implementation in conjunction with the 4-parameter Kappa distribution to create a mixed-distribution framework encompassing the joint L-skew and L-kurtosis domains. The example application provides a demonstration of pertinent algorithms with L-moment statistics and two 4-parameter distributions (AEP4 and the Generalized Lambda) for MLM fitting to a modestly asymmetric and heavy-tailed dataset using R.

Computational Statistics and Data Analysis

Estimating migratory game-bird productivity by integrating age ratio and banding data

Context: Reproduction is a critical component of fitness, and understanding factors that influence temporal and spatial dynamics in reproductive output is important for effective management and conservation. Although several indices of reproductive output for wide-ranging species, such as migratory birds, exist, there has been no theoretical justification for their estimators or associated measures of variance. Aims: The aims of our research were to develop statistical justification for an estimator of reproduction and associated variances on the basis of an existing national wing-collection survey and banding data, and to demonstrate the applicability of this estimator to a migratory game bird. Methods: We used a Bayesian hierarchical modelling approach to integrate wing-collection data, which provides information on population age ratios, and band-recovery data, which provides information on recovery probabilities of various age classes, for American woodcock ( Scolopax minor ) to estimate productivity and associated measures of variance. We present two models of relative vulnerability between age classes: one model assumed that adult recovery probabilities were higher, but that annual fluctuations were synchronous between the two age classes (i.e. an additive effect of age and year). The second model assumed that adults, on average, had higher recovery probabilities than did juveniles and that annual fluctuations were asynchronous through time (i.e. an interaction between age and year). Key results: Fitting our models within a hierarchical Bayesian framework efficiently incorporates the two data types into a single estimator and derives appropriate variances for the productivity estimator. Further, use of Bayesian methods enabled us to derive credible intervals that avoid the reliance on asymptotic assumptions. When applied to American woodcock data, the additive model resulted in biologically realistic and more precise age-ratio estimates each year and is adequate when the relative vulnerability to sampling only slightly varies or does not vary among components of a population (e.g. age, sex class) among years. Therefore, we recommend using woodcock indices from our analysis based on this model. Conclusions: We provide a flexible modelling framework for estimating productivity and associated variances that can incorporate ecological covariates to explore various factors that could drive annual dynamics in productivity. Applying our model to the American woodcock data indicated that assumptions about the variability in relative recovery probabilities could greatly influence the precision of our productivity estimator. Therefore, researchers should carefully consider the assumption of temporally variable relative recovery probabilities (i.e. ratio of juvenile to adults' recovery probability) for different age classes when applying this estimator. Implications: Several national and international management strategies for migratory game birds in North America rely on measures of productivity from harvest survey parts collections, without a justification of the estimator or providing estimates of precision. We derive an estimator of productivity with realistic measures of uncertainty that can be directly incorporated into management plans or ecological studies across large spatial scales.

Wildlife Research

Hawaii StreamStats: A web application for defining drainage-basin characteristics and estimating peak-streamflow statistics

Reliable estimates of the magnitude and frequency of floods are necessary for the safe and efficient design of roads, bridges, water-conveyance structures, and flood-control projects and for the management of flood plains and flood-prone areas. StreamStats provides a simple, fast, and reproducible method to define drainage-basin characteristics and estimate the frequency and magnitude of peak discharges in Hawaii?s streams using recently developed regional regression equations. StreamStats allows the user to estimate the magnitude of floods for streams where data from stream-gaging stations do not exist. Existing estimates of the magnitude and frequency of peak discharges in Hawaii can be improved with continued operation of existing stream-gaging stations and installation of additional gaging stations for areas where limited stream-gaging data are available.

Hawaii

Estimating current and future streamflow characteristics at ungaged sites, central and eastern Montana, with application to evaluating effects of climate change on fish populations

A common statistical procedure for estimating streamflow statistics at ungaged locations is to develop a relational model between streamflow and drainage basin characteristics at gaged locations using least squares regression analysis; however, least squares regression methods are parametric and make constraining assumptions about the data distribution. The random forest regression method provides an alternative nonparametric method for estimating streamflow characteristics at ungaged sites and requires that the data meet fewer statistical conditions than least squares regression methods. Random forest regression analysis was used to develop predictive models for 89 streamflow characteristics using Precipitation-Runoff Modeling System simulated streamflow data and drainage basin characteristics at 179 sites in central and eastern Montana. The predictive models were developed from streamflow data simulated for current (baseline, water years 1982–99) conditions and three future periods (water years 2021–38, 2046–63, and 2071–88) under three different climate-change scenarios. These predictive models were then used to predict streamflow characteristics for baseline conditions and three future periods at 1,707 fish sampling sites in central and eastern Montana. The average root mean square error for all predictive models was about 50 percent. When streamflow predictions at 23 fish sampling sites were compared to nearby locations with simulated data, the mean relative percent difference was about 43 percent. When predictions were compared to streamflow data recorded at 21 U.S. Geological Survey streamflow-gaging stations outside of the calibration basins, the average mean absolute percent error was about 73 percent.

Montana

Fort Collins Science Center: Species and Habitats of Federal Interest

Ecosystem changes directly affect a wide variety of plant and animal species, floral and faunal communities, and groups of species such as amphibians and grassland birds. Appropriate management of public lands plays a crucial role in the conservation and recovery of endangered species and can be a key element in preventing a species from being listed under the Endangered Species Act. The Species and Habitats of Federal Interest Branch of the Fort Collins Science Center (FORT) conducts research on the ecology, habitat requirements, distribution and abundance, population dynamics, and genetics and systematics of many species facing threatened or endangered status or of special concern to resource management agencies. FORT scientists develop reintroduction and restoration techniques, technologies for monitoring populations, and novel methods to analyze data on population trends and habitat requirements. FORT expertise encompasses both traditional and specialized natural resource disciplines within wildlife biology, including population dynamics, animal behavior, plant and community ecology, inventory and monitoring, statistics and computer applications, conservation genetics, stable isotope analysis, and curatorial expertise.

Fact Sheet