Search USGSSearch

SEARCH · Search USGS

Results for “Statistical Science”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Statistical facilitation in environmental science: Integrating results from complementary statistical analyses can improve ecological interpretations

Professionals working in biological conservation seek to understand, manage, and restore populations of native organisms using many techniques. A common approach for this discipline is using long-term data collections to inform decision making. However, several quantitative issues complicate statistical analysis of monitoring datasets and can reduce the utility of results for conservation decision making. Integrating results from multiple analyses applied to the same dataset (i.e., approaching the same biological problem using different techniques) is one way to address concerns related to field data that violate statistical assumptions. This process allows data analysts, researchers, and managers to assemble insights based on the weight of evidence. Here we tested whether three different statistical techniques [(1) multiple logistic regression on original data, (2) multiple logistic regression on standardized data (i.e., mean of 0 and standard deviation of 1), and (3) random forest analysis] identified a similar hierarchy for selecting natural and anthropogenic habitat regressors. Our examination of how environmental variables affected Plains Minnow ( Hybognathus placitus ), a state-threatened fish, is relevant to other taxa and locations. We gained useful information from redundancies (i.e., agreements across analyses). New directions also emerged by addressing ambiguities (i.e., disagreements among results across analyses). When multiple analyses were integrated into one ecological story, a clearer interpretation emerged. Viewing different statistical tests as facilitators that provide mutual advantages can advance the understanding and application of statistical analyses applied to non-experimental field datasets.

Kansas

Birdwatching preferences reveal synergies and tradeoffs among recreation, carbon, and fisheries ecosystem services in Pacific Northwest estuaries, USA

Coastal ecosystems provide multiple ecosystem services that are valued in diverse ways. The Nisqually River Delta (the Delta), an estuary in Puget Sound, Washington, U.S.A., is co-managed by the Nisqually Indian Tribe and the Billy Frank Jr. Nisqually National Wildlife Refuge. In an ecosystem services assessment, we used different service-appropriate methods including citizen science, statistical and geospatial models, and scenario analysis to evaluate three ecosystem services – recreational birdwatching, soil carbon accumulation and fishery production – indicated as priorities for the Refuge, Nisqually Indian Tribe, and surrounding communities. We developed a generalized additive mixed model set based on eBird mobile application birdwatching observations to understand the biological and landscape features that influence birdwatching and to project birdwatching visitation based on scenarios of Delta habitat change. We evaluated ecosystem service synergies and tradeoffs associated with habitat change for three coastal habitat types using scenario outputs from the birdwatching model and published results on Delta soil carbon accumulation and fishery production. The highest-ranked birdwatching models explained 88 % of the deviance and showed that visitation was greatest in winter months when distance to major cities was approximately 20 km. Recreational birdwatching increased with increasing area of forested wetland, emergent wetland, aquatic vegetation bed, open access, and total estuary. With increasing forested and emergent wetland area, recreational birdwatching, out-migrating juvenile Chinook salmon weight and soil carbon accumulation all increased. With increasing aquatic vegetation bed (resulting from sea level rise), recreational birdwatching increased, but salmon weight and soil carbon accumulation decreased. We identified practical ways in which ecosystem services may be incorporated into adaptive management frameworks that support climate adaptation decision making. This study illustrated how use of ecosystem services can help managers make decisions that have greater benefit for wildlife and people, communicate the societal value of decisions and increase local support and participation.

Washington

Challenges for leveraging citizen science to support statistically robust monitoring programs

Large samples and long time series are often needed for effective broad-scale monitoring of status and trends in wild populations. Obtaining those sample sizes can be more feasible when volunteers contribute to the dataset, but volunteer-selected sites are not always representative of a population. Previous work to account for biased site selection has relied on knowledge of covariates to explain differences between site types, but such knowledge is often unavailable. For cases where relevant covariates have not been defined, we used a simulation study to identify the consequences of including non-probabilistically selected sites (NP sites) in addition to sites selected from a probability-based design (P sites), test modeling frameworks that might correct for biases, and evaluate whether those frameworks could allow NP sites to reduce the sampling requirement for P sites and potentially reduce costs of monitoring. We informed the simulation with pilot data from surveys of monarch butterflies and their obligate larval host plant, milkweed. We found strong biases in NP sites versus P sites in density and trends of monarchs and milkweed. Modeling frameworks that accounted for site type with a group effect or that strongly downweighted NP sites successfully produced unbiased estimates. However, sampling more NP sites typically did not improve accuracy or precision, and adding NP sites sometimes required also adding P sites to prevent biases. Further work on novel modeling frameworks would be useful to allow citizen-science data to contribute useful information to conservation.

Biological Conservation

Seismic hazard assessment: Issues and alternatives

Seismic hazard and risk are two very important concepts in engineering design and other policy considerations. Although seismic hazard and risk have often been used inter-changeably, they are fundamentally different. Furthermore, seismic risk is more important in engineering design and other policy considerations. Seismic hazard assessment is an effort by earth scientists to quantify seismic hazard and its associated uncertainty in time and space and to provide seismic hazard estimates for seismic risk assessment and other applications. Although seismic hazard assessment is more a scientific issue, it deserves special attention because of its significant implication to society. Two approaches, probabilistic seismic hazard analysis (PSHA) and deterministic seismic hazard analysis (DSHA), are commonly used for seismic hazard assessment. Although PSHA has been pro-claimed as the best approach for seismic hazard assessment, it is scientifically flawed (i.e., the physics and mathematics that PSHA is based on are not valid). Use of PSHA could lead to either unsafe or overly conservative engineering design or public policy, each of which has dire consequences to society. On the other hand, DSHA is a viable approach for seismic hazard assessment even though it has been labeled as unreliable. The biggest drawback of DSHA is that the temporal characteristics (i.e., earthquake frequency of occurrence and the associated uncertainty) are often neglected. An alternative, seismic hazard analysis (SHA), utilizes earthquake science and statistics directly and provides a seismic hazard estimate that can be readily used for seismic risk assessment and other applications. ?? 2010 Springer Basel AG.

Pure and Applied Geophysics

And the first one now will later be last: Time-reversal in cormack-jolly-seber models

The models of Cormack, Jolly and Seber (CJS) are remarkable in providing a rich set of inferences about population survival, recruitment, abundance and even sampling probabilities from a seemingly limited data source: a matrix of 1's and 0's reflecting animal captures and recaptures at multiple sampling occasions. Survival and sampling probabilities are estimated directly in CJS models, whereas estimators for recruitment and abundance were initially obtained as derived quantities. Various investigators have noted that just as standard modeling provides direct inferences about survival, reversing the time order of capture history data permits direct modeling and inference about recruitment. Here we review the development of reverse-time modeling efforts, emphasizing the kinds of inferences and questions to which they seem well suited.

Statistical Science

100-Year flood–it's all about chance

In the 1960's, the United States government decided to use the 1-percent annual exceedance probability (AEP) flood as the basis for the National Flood Insurance Program. The 1-percent AEP flood was thought to be a fair balance between protecting the public and overly stringent regulation. Because the 1-percent AEP flood has a 1 in 100 chance of being equaled or exceeded in any 1 year, and it has an average recurrence interval of 100 years, it often is referred to as the '100-year flood'. The term '100-year flood' is part of the national lexicon, but is often a source of confusion by those not familiar with flood science and statistics. This poster is an attempt to explain the concept, probabilistic nature, and inherent uncertainties of the '100-year flood' to the layman.

General Information Product

Revised recommended methods for analyzing crater size-frequency distributions

Impact crater populations crucially help us to understand solar system dynamics, planetary surface histories, and surface modification processes. A single previous effort to standardize how crater data are displayed in graphs, tables, and archives, was in a 1978 NASA report by the Crater Analysis Techniques Working Group, published in 1979 in Icarus . The report had a significant lasting effect, but later decades brought major advances in statistical and computer sciences while the crater field has remained fairly stagnant. In this new work, we revisit the fundamental techniques for displaying and analyzing crater population data and demonstrate better statistical methods that can be used. Specifically, we address (1) how crater size-frequency distributions (SFDs) are constructed, (2) how error bars are assigned to SFDs, and (3) how SFDs are fit to power laws and other models. We show how the new methods yield results similar to those of previous techniques in that the SFDs have familiar shapes but better account for multiple sources of uncertainty. We also recommend graphic, display, and archiving methods that reflect computers' capabilities and fulfill NASA's current requirements for Data Management Plans.

Meteoritics and Planetary Science

Opportunities and challenges of macrogenetic studies

The rapidly emerging field of macrogenetics focuses on analysing publicly accessible genetic datasets from thousands of species to explore large-scale patterns and predictors of intraspecific genetic variation. Facilitated by advances in evolutionary biology, technology, data infrastructure, statistics and open science, macrogenetics addresses core evolutionary hypotheses (such as disentangling environmental and life-history effects on genetic variation) with a global focus. Yet, there are important, often overlooked, limitations to this approach and best practices need to be considered and adopted if macrogenetics is to continue its exciting trajectory and reach its full potential in fields such as biodiversity monitoring and conservation. Here, we review the history of this rapidly growing field, highlight knowledge gaps and future directions, and provide guidelines for further research.

Nature Reviews Genetics

Basic Statistical Concepts and Methods for Earth Scientists

INTRODUCTION Statistics is the science of collecting, analyzing, interpreting, modeling, and displaying masses of numerical data primarily for the characterization and understanding of incompletely known systems. Over the years, these objectives have lead to a fair amount of analytical work to achieve, substantiate, and guide descriptions and inferences.

Open-File Report

CORSSA: The Community Online Resource for Statistical Seismicity Analysis

Statistical seismology is the application of rigorous statistical methods to earthquake science with the goal of improving our knowledge of how the earth works. Within statistical seismology there is a strong emphasis on the analysis of seismicity data in order to improve our scientific understanding of earthquakes and to improve the evaluation and testing of earthquake forecasts, earthquake early warning, and seismic hazards assessments. Given the societal importance of these applications, statistical seismology must be done well. Unfortunately, a lack of educational resources and available software tools make it difficult for students and new practitioners to learn about this discipline. The goal of the Community Online Resource for Statistical Seismicity Analysis (CORSSA) is to promote excellence in statistical seismology by providing the knowledge and resources necessary to understand and implement the best practices, so that the reader can apply these methods to their own research. This introduction describes the motivation for and vision of CORRSA. It also describes its structure and contents.

Community Online Resource for Statistical Seismici

A method for assigning species into groups based on generalized Mahalanobis distance between habitat model coefficients

Habitat association models are commonly developed for individual animal species using generalized linear modeling methods such as logistic regression. We considered the issue of grouping species based on their habitat use so that management decisions can be based on sets of species rather than individual species. This research was motivated by a study of western landbirds in northern Idaho forests. The method we examined was to separately fit models to each species and to use a generalized Mahalanobis distance between coefficient vectors to create a distance matrix among species. Clustering methods were used to group species from the distance matrix, and multidimensional scaling methods were used to visualize the relations among species groups. Methods were also discussed for evaluating the sensitivity of the conclusions because of outliers or influential data points. We illustrate these methods with data from the landbird study conducted in northern Idaho. Simulation results are presented to compare the success of this method to alternative methods using Euclidean distance between coefficient vectors and to methods that do not use habitat association models. These simulations demonstrate that our Mahalanobis-distance- based method was nearly always better than Euclidean-distance-based methods or methods not based on habitat association models. The methods used to develop candidate species groups are easily explained to other scientists and resource managers since they mainly rely on classical multivariate statistical methods. ?? 2008 Springer Science+Business Media, LLC.

Environmental and Ecological Statistics

Ten quick tips to get you started with Bayesian statistics

Bayesian statistics is a framework in which our knowledge about unknown quantities of interest (especially parameters) is updated with the information in observed data, though it can also be viewed as simply another method to fit a statistical model. It has become popular in many branches of biology. For context, five of the ten most cited papers in Web of Science with keywords 'Bayesian statistics' are related to biology (as of August 19, 2024). Bayesian statistics is particularly valuable for biology because it allows researchers to incorporate prior knowledge, handle complex systems, and work effectively with limited or messy data. However, most biologists are trained in frequentist techniques, and the learning curve to become fluent in Bayesian statistics may be perceived as too time-consuming to undertake, or the prospect of adopting an unfamiliar statistical framework can simply appear too daunting. We provide a list of 10 tips to help you get started with Bayesian statistics. You can also refer to the Glossary for definitions of the technical terms. This paper isn’t just for newcomers; even those with some experience in Bayesian methods may find it a useful roadmap to design, conduct, and publish Bayesian analyses. We’ve drawn mainly on our experience teaching and working with ecologists, but we hope these tips will be relevant to a broader audience of biologists. For those seeking to deepen their understanding, we point to more comprehensive resources that offer in-depth exploration of Bayesian statistics.

HAL Open Science

Using assemblage data in ecological indicators: A comparison and evaluation of commonly available statistical tools

Ecological indicators are science-based tools used to assess how human activities have impacted environmental resources. For monitoring and environmental assessment, existing species assemblage data can be used to make these comparisons through time or across sites. An impediment to using assemblage data, however, is that these data are complex and need to be simplified in an ecologically meaningful way. Because multivariate statistics are mathematical relationships, statistical groupings may not make ecological sense and will not have utility as indicators. Our goal was to define a process to select defensible and ecologically interpretable statistical simplifications of assemblage data in which researchers and managers can have confidence. For this, we chose a suite of statistical methods, compared the groupings that resulted from these analyses, identified convergence among groupings, then we interpreted the groupings using species and ecological guilds. When we tested this approach using a statewide stream fish dataset, not all statistical methods worked equally well. For our dataset, logistic regression (Log), detrended correspondence analysis (DCA), cluster analysis (CL), and non-metric multidimensional scaling (NMDS) provided consistent, simplified output. Specifically, the Log, DCA, CL-1, and NMDS-1 groupings were ≥60% similar to each other, overlapped with the fluvial-specialist ecological guild, and contained a common subset of species. Groupings based on number of species (e.g., Log, DCA, CL and NMDS) outperformed groupings based on abundance [e.g., principal components analysis (PCA) and Poisson regression]. Although the specific methods that worked on our test dataset have generality, here we are advocating a process (e.g., identifying convergent groupings with redundant species composition that are ecologically interpretable) rather than the automatic use of any single statistical tool. We summarize this process in step-by-step guidance for the future use of these commonly available ecological and statistical methods in preparing assemblage data for use in ecological indicators.

Ecological Indicators

Combining statistical inference and decisions in ecology

Statistical decision theory (SDT) is a sub-field of decision theory that formally incorporates statistical investigation into a decision-theoretic framework to account for uncertainties in a decision problem. SDT provides a unifying analysis of three types of information: statistical results from a data set, knowledge of the consequences of potential choices (i.e., loss), and prior beliefs about a system. SDT links the theoretical development of a large body of statistical methods including point estimation, hypothesis testing, and confidence interval estimation. The theory and application of SDT have mainly been developed and published in the fields of mathematics, statistics, operations research, and other decision sciences, but have had limited exposure in ecology. Thus, we provide an introduction to SDT for ecologists and describe its utility for linking the conventionally separate tasks of statistical investigation and decision making in a single framework. We describe the basic framework of both Bayesian and frequentist SDT, its traditional use in statistics, and discuss its application to decision problems that occur in ecology. We demonstrate SDT with two types of decisions: Bayesian point estimation, and an applied management problem of selecting a prescribed fire rotation for managing a grassland bird species. Central to SDT, and decision theory in general, are loss functions. Thus, we also provide basic guidance and references for constructing loss functions for an SDT problem.

Ecological Applications

A practical guide to understanding and validating complex models using data simulations

Biologists routinely fit novel and complex statistical models to push the limits of our understanding. Examples include, but are not limited to, flexible Bayesian approaches (e.g. BUGS, stan), frequentist and likelihood-based approaches (e.g. packages lme4 ) and machine learning methods. These software and programs afford the user greater control and flexibility in tailoring complex hierarchical models. However, this level of control and flexibility places a higher degree of responsibility on the user to evaluate the robustness of their statistical inference. To determine how often biologists are running model diagnostics on hierarchical models, we reviewed 50 recently published papers in 2021 in the journal Nature Ecology & Evolution , and we found that the majority of published papers did not report any validation of their hierarchical models, making it difficult for the reader to assess the robustness of their inference. This lack of reporting likely stems from a lack of standardized guidance for best practices and standard methods. Here, we provide a guide to understanding and validating complex models using data simulations. To determine how often biologists use data simulation techniques, we also reviewed 50 recently published papers in 2021 in the journal Methods Ecology & Evolution . We found that 78% of the papers that proposed a new estimation technique, package or model used simulations or generated data in some capacity (18 of 23 papers); but very few of those papers (5 of 23 papers) included either a demonstration that the code could recover realistic estimates for a dataset with known parameters or a demonstration of the statistical properties of the approach. To distil the variety of simulations techniques and their uses, we provide a taxonomy of simulation studies based on the intended inference. We also encourage authors to include a basic validation study whenever novel statistical models are used, which in general, is easy to implement. Simulating data helps a researcher gain a deeper understanding of the models and their assumptions and establish the reliability of their estimation approaches. Wider adoption of data simulations by biologists can improve statistical inference, reliability and open science practices.

Methods in Ecology and Evolution

Resource materials for a GIS spatial analysis course

This report consists of materials prepared for a GIS spatial analysis course offered as part of the Geography curriculum at the University of Nevada, Reno and the University of California at Santa Barbara in the spring of 2000. The report is intended to share information with instructors preparing spatial-modeling training and scientists with advanced GIS expertise. The students taking this class had completed each universities GIS curriculum and had a foundation in statistics as part of a science major. This report is organized into chapters that contain the following: Slides used during lectures, Guidance on the use of Arcview, Introduction to filtering in Arcview, Conventional and spatial correlation in Arcview, Tools for fuzzification in Arcview, Data and instructions for creating using ArcSDM for simple weights-of-evidence, fuzzy logic, and neural network models for Carlin-type gold deposits in central Nevada, Reading list on spatial modeling, and Selected student spatial-modeling posters from the laboratory exercises.

Open-File Report

Concentrations of tritium and strontium-90 in water from selected wells at the Idaho National Engineering Laboratory after purging one, two, and three borehole volumes

Water from 11 wells completed in the Snake River Plain aquifer at the Idaho National Engineering Laboratory was sampled as part of the U.S. Geological Survey's quality assurance program to determine the effect of purging different borehole volumes on tritium and strontium-90 concentrations. Wells were selected for sampling on the basis of the length of time it took to purge a borehole volume of water. Samples were collected after purging one, two, and three borehole volumes. The U.S. Department of Energy's Radiological and Environmental Sciences Laboratory provided analytical services. Statistics were used to determine the reproducibility of analytical results. The comparison between tritium and strontium-90 concentrations after purging one and three borehole volumes and two and three borehole volumes showed that all but two sample pairs with defined numbers were in statistical agreement. Results indicate that concentrations of tritium and strontium-90 are not affected measurably by the number of borehole volumes purged.

Water-Resources Investigations Report