Search USGSSearch

SEARCH · Search USGS

Results for “Statistical Science”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Basalt–trachybasalt samples in Gale Crater, Mars

The ChemCam instrument on the Mars Science Laboratory (MSL) rover, Curiosity, observed numerous igneous float rocks and conglomerate clasts, reported previously. A new statistical analysis of single‐laser‐shot spectra of igneous targets observed by ChemCam shows a strong peak at ~55 wt% SiO 2 and 6 wt% total alkalis, with a minor secondary maximum at 47–51 wt% SiO 2 and lower alkali content. The centers of these distributions, together with the rock textures, indicate that many of the ChemCam igneous targets are trachybasalts, Mg# = 27 but with a secondary concentration of basaltic material, with a focus of compositions around Mg# = 54. We suggest that all of these igneous rocks resulted from low‐pressure, olivine‐dominated fractionation of Adirondack (MER) class‐type basalt compositions. This magmatism has subalkaline, tholeiitic affinities. The similarity of the basalt endmember to much of the Gale sediment compositions in the first 1000 sols of the MSL mission suggests that this type of Fe‐rich, relatively low‐Mg#, olivine tholeiite is the dominant constituent of the Gale catchment that is the source material for the fine‐grained sediments in Gale. The similarity to many Gusev igneous compositions suggests that it is a major constituent of ancient Martian magmas, and distinct from the shergottite parental melts thought to be associated with Tharsis and the Northern Lowlands. The Gale Crater catchment sampled a mixture of this tholeiitic basalt along with alkaline igneous material, together giving some analogies to terrestrial intraplate magmatic provinces.

Meteoritics and Planetary Science

Techniques to improve ecological interpretability of black box machine learning models

Statistical modeling of ecological data is often faced with a large number of variables as well as possible nonlinear relationships and higher-order interaction effects. Gradient boosted trees (GBT) have been successful in addressing these issues and have shown a good predictive performance in modeling nonlinear relationships, in particular in classification settings with a categorical response variable. They also tend to be robust against outliers. However, their black-box nature makes it difficult to interpret these models. We introduce several recently developed statistical tools to the environmental research community in order to advance interpretation of these black-box models. To analyze the properties of the tools, we applied gradient boosted trees to investigate biological health of streams within the contiguous USA, as measured by a benthic macroinvertebrate biotic index. Based on these data and a simulation study, we demonstrate the advantages and limitations of partial dependence plots (PDP), individual conditional expectation (ICE) curves and accumulated local effects (ALE) in their ability to identify covariate–response relationships. Additionally, interaction effects were quantified according to interaction strength (IAS) and Friedman’s H 2 "> H 2 statistic. Interpretable machine learning techniques are useful tools to open the black-box of gradient boosted trees in the environmental sciences. This finding is supported by our case study on the effect of impervious surface on the benthic condition, which agrees with previous results in the literature. Overall, the most important variables were ecoregion, bed stability, watershed area, riparian vegetation and catchment slope. These variables were also present in most identified interaction effects. In conclusion, graphical tools (PDP, ICE, ALE) enable visualization and easier interpretation of GBT but should be supported by analytical statistical measures. Future methodological research is needed to investigate the properties of interaction tests. Supplementary materials accompanying this paper appear on-line.

Journal of Agricultural, Biological, and Environme

Improving the analysis of slug tests

This paper examines several techniques that have the potential to improve the quality of slug test analysis. These techniques are applicable in the range from low hydraulic conductivities with overdamped responses to high hydraulic conductivities with nonlinear oscillatory responses. Four techniques for improving slug test analysis will be discussed: use of an extended capability nonlinear model, sensitivity analysis, correction for acceleration and velocity effects, and use of multiple slug tests. The four-parameter nonlinear slug test model used in this work is shown to allow accurate analysis of slug tests with widely differing character. The parameter ?? represents a correction to the water column length caused primarily by radius variations in the wellbore and is most useful in matching the oscillation frequency and amplitude. The water column velocity at slug initiation (V0) is an additional model parameter, which would ideally be zero but may not be due to the initiation mechanism. The remaining two model parameters are A (parameter for nonlinear effects) and K (hydraulic conductivity). Sensitivity analysis shows that in general ?? and V0 have the lowest sensitivity and K usually has the highest. However, for very high K values the sensitivity to A may surpass the sensitivity to K. Oscillatory slug tests involve higher accelerations and velocities of the water column; thus, the pressure transducer responses are affected by these factors and the model response must be corrected to allow maximum accuracy for the analysis. The performance of multiple slug tests will allow some statistical measure of the experimental accuracy and of the reliability of the resulting aquifer parameters. ?? 2002 Elsevier Science B.V. All rights reserved.

Journal of Hydrology

Modeling false positives

Many of the models we are concerned with included explicit descriptions of false negative errors. However, false positive errors can also be commin in practice, especially in citizen science applications where observer skill is highly variable. In addition, new methods which determine detection based on statistical classification or machine learning methods are also prone to false positive errors which must be accounted for. An early treatment of the false positive detection problem by Royle & Link (2006) recognized that false positive errors can be accommodated by a mixture model for detection probability: one value of detection at occupied sites and another non-zero value at unoccupied sites. This model has been extended greatly in recent years to include more informative data about false positives including validation or confirmation data (Miller et al. 2011) and multiple detection methods, among others. A new frontier for the application of false positives models lies in the use of modern technologies such as bioacoustics for efficient automated monitoring. For these technologies to realize their promise there must be improvements in automated processing of the vast quantities of output produced. Statistical classification methods (machine learning) are fallible and necessarily produce false positive detections. Therefore models which account for this process are necessary (Chambert et al. 2017). It stands to reason that false positives will need to be accounted for in other new technologies that rely on automated digital processing, including eDNA, genetic barcoding, and automated detection in remote camera studies. We devise a new occupancy model that integrates data from bioacoustics sampling with an occupancy model. This integrated model allows occupancy probability to inform species classification of samples and vice versa bioacoustics detection data inform occupancy. We provide a proof of concept for this new model in this chapter. As the core hierarchical model for the false positives models covered in this chapter are just ordinary occupancy models, extension of the ideas to open systems poses no technical challenges. We provide a suite of illustrations of these extensions. Perhaps the most prominent mechanism that leads to false positive errors it he mis-classification of species detections, or the confusion of one species for another. Very little work has been done on developing models based on this mechanistic understanding although Chambert et al. (2018) develop this idea as a 2-species occupancy model with error. We believe one important area of future research is to extend these ideas to truly multi-species systems.

Book chapter

Statistical inference from capture data on closed animal populations

The estimation of animal abundance is an important problem in both the theoretical and applied biological sciences. Serious work to develop estimation methods began during the 1950s, with a few attempts before that time. The literature on estimation methods has increased tremendously during the past 25 years (Cormack 1968, Seber 1973). However, in large part, the problem remains unsolved. Past efforts toward comprehensive and systematic estimation of density (D) or population size (N) have been inadequate, in general. While more than 200 papers have been published on the subject, one is generally left without a unified approach to the estimation of abundance of an animal population This situation is unfortunate because a number of pressing research problems require such information. In addition, a wide array of environmental assessment studies and biological inventory programs require the estimation of animal abundance. These needs have been further emphasized by the requirement for the preparation of Environmental Impact Statements imposed by the National Environmental Protection Act in 1970. This publication treats inference procedures for certain types of capture data on closed animal populations. This includes multiple capture-recapture studies (variously called capture-mark-recapture, mark-recapture, or tag-recapture studies) involving livetrapping techniques and removal studies involving kill traps or at least temporary removal of captured individuals during the study. Animals do not necessarily need to be physically trapped; visual sightings of marked animals and electrofishing studies also produce data suitable for the methods described in this monograph. To provide a frame of reference for what follows, we give an exampled of a capture-recapture experiment to estimate population size of small animals using live traps. The general field experiment is similar for all capture-recapture studies (a removal study is, of course, slightly different). A typical field experiment is the following: a number of traps are positioned in the area to be studied, say 144 traps in a 12 X 12 grid, 7 m apart. At the beginning of the study (j=1) a sample size of n 1 is taken from the population, the animals are tagged and marked for future identification, and then returned to the population, usually at the same point where they were trapped. After allowing time of the marked and unmarked animals to mix, a second sample (j=2, often the following day) or n 2 animals is then taken.the second sample normally contains both marked and unmarked animals. The unmarked animals are marked and all captured animals are released back into the population. This procedure continues for t periods where t ≥ 2. The animals should be marked in such a way that the capture-recapture history of each animal caught during the study is known. In practice, toes are often clipped to uniquely identify individual animals (Taber and Cowan 1969) or serially numbered tags are sometimes used on larger animals. Such capture studies are classified by 2 schemes that are directly related to what class of models are appropriate and what parameters can be estimated. The first classification addresses the subject of closure. Closure usually means the size of the population is constant over the priod of investigation, i.e., no recruitment (birth or immigration) or losses (death or emigration). This is a strong assumption and, of course, never completely true in a natural biological population. For greater generality, we define closure to mean there are no unknown changes to the initial population. In practice, this means known losses (trap death), or deliberate removals) do not violate our definition of closure. If the study is properly designed, closure can be met at least approximately. Open or nonclosed populations explicitly allow for one or more types of recruitment or losses to operate during the course of the experiment (Jolly 1965, Seber 1965, Robson 1969, Pollock 1975). Only closed populations will be considered in this monograph. The second classification depends on the type of data collected with 2 possibilities occurring (Pollock 1974, unpublished doctoral dissertation, Cornell University, Ithaca, New York): (1) only information on the recovery of marked animals is available for each sampling occasion, j, j=1, 2, ... t. (2) information on both marked and unmarked animals is available for each sampling occasion, j, j=1, 2, ... t. In case (1), population size (N) is not identifiable, however, other parameters can be estimated (Brownie et al. 1978). In case (2), N can be estimated using a wide variety of approaches depending upon what we wish to assume. Only case (2) will be dealt with here.

Wildlife Monographs

Bayesian statistics for beginners: A step-by-step approach

Bayesian statistics is currently undergoing something of a renaissance. At its heart is a method of statistical inference in which Bayes' theorem is used to update the probability for a hypothesis as more evidence or information becomes available. It is an approach that is ideally suited to making initial assessments based on incomplete or imperfect information; as that information is gathered and disseminated, the Bayesian approach corrects or replaces the assumptions and alters its decision-making accordingly to generate a new set of probabilities. As new data/evidence becomes available the probability for a particular hypothesis can therefore be steadily refined and revised. It is very well-suited to the scientific method in general and is widely used across the social, biological, medical, and physical sciences. Key to this book's novel and informal perspective is its unique pedagogy, a question and answer approach that utilizes accessible language, humor, plentiful illustrations, and frequent reference to on-line resources. Bayesian Statistics for Beginners is an introductory textbook suitable for senior undergraduate and graduate students, professional researchers, and practitioners seeking to improve their understanding of the Bayesian statistical techniques they routinely use for data analysis in the life and medical sciences, psychology, public health, business, and other fields.

Book

Spatially-structured statistical network models for landscape genetics

A basic understanding of how the landscape impedes, or creates resistance to, the dispersal of organisms and hence gene flow is paramount for successful conservation science and management. Spatially structured ecological networks are often used to represent spatial landscape‐genetic relationships, where nodes represent individuals or populations and resistance to movement is represented using non‐binary edge weights. Weights are typically assigned or estimated by the user, rather than observed, and validating such weights is challenging. We provide a synthesis of current methods used to estimate edge weights and an overview of common model types, stressing the advantages and disadvantages of each approach and their ability to model landscape‐genetic data. We further explore a set of spatial‐statistical methods that provide ecologists with alternative approaches for modeling spatially explicit processes that may affect genetic structure. This includes an overview of spatial autoregressive models, with a particular focus on how correlation and partial correlation are used to represent neighborhood structure with the inverse of the covariance matrix (i.e., precision matrix). We then demonstrate how to model resistance by specifying an appropriate statistical model on the nodes, conditioned on the edge weights, through the precision matrix. This integration of network ecology and spatial statistics provides a practical analytical framework for landscape‐genetic studies. The results can be used to make statistical inferences about the relative importance of individual landscape characteristics, such as the vegetative cover, hillslope, or the presence of roads or rivers, on gene flow. In addition, the R code we include allows readers to explore landscape‐genetic structure in their own datasets, which will potentially provide new insights into the evolutionary processes that generated ecological networks, as well as valuable information about the optimal characteristics of conservation corridors.

Ecological Monographs

Statistical analysis of water-quality data containing multiple detection limits II: S-language software for nonparametric distribution modeling and hypothesis testing

Analysis of low concentrations of trace contaminants in environmental media often results in left-censored data that are below some limit of analytical precision. Interpretation of values becomes complicated when there are multiple detection limits in the data-perhaps as a result of changing analytical precision over time. Parametric and semi-parametric methods, such as maximum likelihood estimation and robust regression on order statistics, can be employed to model distributions of multiply censored data and provide estimates of summary statistics. However, these methods are based on assumptions about the underlying distribution of data. Nonparametric methods provide an alternative that does not require such assumptions. A standard nonparametric method for estimating summary statistics of multiply-censored data is the Kaplan-Meier (K-M) method. This method has seen widespread usage in the medical sciences within a general framework termed "survival analysis" where it is employed with right-censored time-to-failure data. However, K-M methods are equally valid for the left-censored data common in the geosciences. Our S-language software provides an analytical framework based on K-M methods that is tailored to the needs of the earth and environmental sciences community. This includes routines for the generation of empirical cumulative distribution functions, prediction or exceedance probabilities, and related confidence limits computation. Additionally, our software contains K-M-based routines for nonparametric hypothesis testing among an unlimited number of grouping variables. A primary characteristic of K-M methods is that they do not perform extrapolation and interpolation. Thus, these routines cannot be used to model statistics beyond the observed data range or when linear interpolation is desired. For such applications, the aforementioned parametric and semi-parametric methods must be used.

Computers & Geosciences

Representing general theoretical concepts in structural equation models: The role of composite variables

Structural equation modeling (SEM) holds the promise of providing natural scientists the capacity to evaluate complex multivariate hypotheses about ecological systems. Building on its predecessors, path analysis and factor analysis, SEM allows for the incorporation of both observed and unobserved (latent) variables into theoretically-based probabilistic models. In this paper we discuss the interface between theory and data in SEM and the use of an additional variable type, the composite. In simple terms, composite variables specify the influences of collections of other variables and can be helpful in modeling heterogeneous concepts of the sort commonly of interest to ecologists. While long recognized as a potentially important element of SEM, composite variables have received very limited use, in part because of a lack of theoretical consideration, but also because of difficulties that arise in parameter estimation when using conventional solution procedures. In this paper we present a framework for discussing composites and demonstrate how the use of partially-reduced-form models can help to overcome some of the parameter estimation and evaluation problems associated with models containing composites. Diagnostic procedures for evaluating the most appropriate and effective use of composites are illustrated with an example from the ecological literature. It is argued that an ability to incorporate composite variables into structural equation models may be particularly valuable in the study of natural systems, where concepts are frequently multifaceted and the influence of suites of variables are often of interest. ?? Springer Science+Business Media, LLC 2007.

Environmental and Ecological Statistics

Remote sensing of global croplands for food security: Way forward

This book opens a new pathway for global mapping that is focused on a specific land use theme, such as irrigated or rain-fed croplands and classes within these themes. Since croplands use most of the water consumed by humans, specific knowledge of irrigated and rain-fed croplands will be critical for precise estimates of water use. At present and in the coming decades, irrigated and rain-fed cropland area mapping is crucial for food security studies. Throughout this book, various subjects pertaining to global croplands are discussed comprehensively.

Book chapter

Geological, hydrological, and biological issues related to the proposed development of a park at the confluence of the Los Angeles River and the Arroyo Seco, Los Angeles County, California

A new park is being considered for the confluence of the Los Angeles River and the Arroyo Seco in Los Angeles County, California. Components of the park development may include creation of a temporary lake on the Los Angeles River, removal of channel lining along part of the Arroyo Seco, restoration of native plants, creation of walking paths, and building of facilities such as a boat ramp and a visitor center. This report, prepared in cooperation with the Mountains Recreation and Conservancy Authority, delineates the geological, hydrological, and biological issues that may have an impact on the park development or result from development at the confluence, and identifies a set a tasks to help address these science issues. Geologic issues of concern relate to surface faulting, earthquake ground motions, liquefaction, landsliding, and induced seismicity. Hydrologic issues of concern relate to the hydraulics and water quality of both surface water and ground water. Biological issues of concern include colonization-extinction dynamics, wildlife corridors, wildlife reintroduction, non-native species, ecotoxicology, and restoration of local habitat and ecology. Potential tasks include (1) basic data collection and follow-up monitoring, and (2) statistical and probabilistic analyses and simulation modeling of the seismic, hydraulic, and ecological processes that may have the greatest impact on the park. The science issues and associated tasks delineated for the proposed confluence park will also have transfer value for river restoration in other urban settings.

Scientific Investigations Report

Preliminary preview for a geographic and monitoring program project; a review of point source-nonpoint source effluent trading/offset systems in watersheds

Watershed-based trading and offset systems are being developed to improve policy-maker?s and regulator?s ability to assess nonpoint source impacts in watersheds and to evaluate the efficacy of using market-incentive programs for preserving environmental quality. An overview of the history of successful and failed trading programs throughout the United States suggests that certain political, economic, and scientific conditions within a temporal and spatial setting help meet water quality standards. The current lack of spontaneous trading among dischargers does not mean that a marketable permit trading system is an inherently inefficient regulatory approach. Rather, its infrequent use is the result of institutional and informational barriers. Improving and refining the earth science information and technologies may help determine whether trading is a suitable policy for improving water quality. However, it is debatable whether or not environmental information is the limiting factor. This paper reviews additional factors affecting the potential for instituting a trading policy. The motivation for investigating and reviewing the history of offsets and trading was inspired by a project in the preliminary stages being developed by U.S. Geological Survey Western Geographic Science Center and the Environmental Protection Agency Region IX. An offset feasibility study will be an integrated, map-based approach that incorporates environmental, economic, and statistical information to investigate the potential for using offsets to meet mercury Total Maximum Daily Loads in the Sacramento River watershed. A regional water-quality offset program is being studied that may help known point sources reduce mercury loading more cost effectively by the remediation of abandoned mines or other diffuse sources as opposed to more costly treatment at their own sites. An efficient offset program requires both a scientific basis and methods to translate that science into a regulatory decision framework.

Open-File Report

Fifteen years of WRTDS for advancing water-quality science: A critical review of methodological developments and global applications

Contamination by nutrients, major ions, and metals poses a major threat to global water sustainability. Understanding how these pollutants vary across time and space requires long-term monitoring and robust statistical approaches. Traditional methods, however, often struggle to account for streamflow variability, seasonality, and nonlinear responses. Introduced in 2010, the Weighted Regressions on Time, Discharge, and Season (WRTDS) method offers a flexible, data-driven framework that generates both observed and flow-normalized estimates of concentration and load. Over the past 15 years, WRTDS has become a state-of-the-art tool for water-quality science and management, with applications spanning a wide range of hydrologic, climatic, and policy contexts─including major watersheds across North America, Europe, Asia, Australia, and the Arctic. In this review of WRTDS, we document the method’s major advancements, examine its expanding geographic and thematic applications, and summarize its relevance to water-quality management programs and policies worldwide. We also discuss its performance relative to other regression and machine-learning approaches. Finally, we identify key priorities for future development to support the continued evolution of WRTDS as a trusted and practical tool for scientists and managers working to protect and sustain water resources.

Environmental Science and Technology

Perils of correlating CUSUM-transformed variables to infer ecological relationships (Breton et al. 2006; Glibert 2010)

We comment on a nonstandard statistical treatment of time-series data first published by Breton et al. (2006) in Limnology and Oceanography and, more recently, used by Glibert (2010) in Reviews in Fisheries Science. In both papers, the authors make strong inferences about the underlying causes of population variability based on correlations between cumulative sum (CUSUM) transformations of organism abundances and environmental variables. Breton et al. (2006) reported correlations between CUSUM-transformed values of diatom biomass in Belgian coastal waters and the North Atlantic Oscillation, and between meteorological and hydrological variables. Each correlation of CUSUM-transformed variables was judged to be statistically significant. On the basis of these correlations, Breton et al. (2006) developed "the first evidence of synergy between climate and human-induced river-based nitrate inputs with respect to their effects on the magnitude of spring Phaeocystis colony blooms and their dominance over diatoms."

Limnology and Oceanography

Missing data in ecology: Syntheses, clarifications, and considerations

In ecology and related sciences, missing data are common and occur in a variety of different contexts. When missing data are not handled properly, subsequent statistical estimates tend to be biased, inefficient, and lack proper confidence interval coverage. Missing data are often grouped into three categories: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). We review each category and compare their benefits and drawbacks. We review several approaches to handling missing data including complete case analysis, imputation, inverse probability weighting, and data augmentation. We clarify what types of variables should accompany imputation methods and how those variables are influenced by the analysis methods. Additionally, we discuss missing data that lack a formal basis for measurement and hence are fundamentally different from MCAR, MAR, and MNAR missing data. Throughout, we introduce concepts and numeric examples using both simulated data and data from the United States Environmental Protection Agency's 2016 National Wetland Condition Assessment. We conclude by providing five considerations for ecologists and other scientists handling missing data.

Ecological Monographs

Landsat 8 on-orbit characterization and calibration system

The Landsat Data Continuity Mission (LDCM) is planning to launch the Landsat 8 satellite in December 2012, which continues an uninterrupted record of consistently calibrated globally acquired multispectral images of the Earth started in 1972. The satellite will carry two imaging sensors: the Operational Land Imager (OLI) and the Thermal Infrared Sensor (TIRS). The OLI will provide visible, near-infrared and short-wave infrared data in nine spectral bands while the TIRS will acquire thermal infrared data in two bands. Both sensors have a pushbroom design and consequently, each has a large number of detectors to be characterized. Image and calibration data downlinked from the satellite will be processed by the U.S. Geological Survey (USGS) Earth Resources Observation and Science (EROS) Center using the Landsat 8 Image Assessment System (IAS), a component of the Ground System. In addition to extracting statistics from all Earth images acquired, the IAS will process and trend results from analysis of special calibration acquisitions, such as solar diffuser, lunar, shutter, night, lamp and blackbody data, and preselected calibration sites. The trended data will be systematically processed and analyzed, and calibration and characterization parameters will be updated using both automatic and customized manual tools. This paper describes the analysis tools and the system developed to monitor and characterize on-orbit performance and calibrate the Landsat 8 sensors and image data products.

Conference Paper

40Ar/39Ar age of Cretaceous-Tertiary boundary tektites from Haiti

40 Ar/ 39 Ar dating of tektites discovered recently in Cretaceous-Tertiary (K-T) boundary marine sedimentary rocks on Haiti indicates that the K-T boundary and impact event are coeval at 64.5 ± 0.1 million years ago. Sanidine from a bentonite that lies directly above the K-T boundary in continental, coal-bearing, sedimentary rocks of Montana was also dated and has a 40 Ar/ 39 Ar age of 64.6 ± 0.2 million years ago, which is indistinguishable statistically from the age of the tektites.

Science

Statistical methods used in research concerning endangered and threatened animal species of Puerto Rico: A meta-study

A concern about statistics in wildlife studies, particularly of endangered and threatened species, is whether the data collected meet the assumptions necessary for the use of parametric statistics. This study identified published papers on the nine endangered and six threatened species found only on Puerto Rico using five different databases. The results from the Zoological Record database identified the most articles, including all identified by the other databases. Of the 222 identified articles, 108 included some form of statistics, 26 used only descriptive statistics, 34 included only parametric statistics, 26 used only nonparametric statistics, and 22 reported both parametric and nonparametric statistical analyses. This meta-study showed that the percentage of articles with no statistical treatment decreased in the most recent 20 years, and that although parametric statistics continue to be the most commonly used in published wildlife studies of Puerto Rican wildlife, there has been a distinct increase in the use of nonparametric statistics over time.

Caribbean Journal of Science