Search USGSSearch

SEARCH · Search USGS

Results for “Statistical Methodology”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Thematic accuracy of the 1992 National Land-Cover Data for the eastern United States: Statistical methodology and regional results

The accuracy of the 1992 National Land-Cover Data (NLCD) map is assessed via a probability sampling design incorporating three levels of stratification and two stages of selection. Agreement between the map and reference land-cover labels is defined as a match between the primary or alternate reference label determined for a sample pixel and a mode class of the mapped 3×3 block of pixels centered on the sample pixel. Results are reported for each of the four regions comprising the eastern United States for both Anderson Level I and II classifications. Overall accuracies for Levels I and II are 80% and 46% for New England, 82% and 62% for New York/New Jersey (NY/NJ), 70% and 43% for the Mid-Atlantic, and 83% and 66% for the Southeast.

Remote Sensing of Environment

Spatio-temporal ecological models via physics-informed neural networks for studying chronic wasting disease

To mitigate the negative effects of emerging wildlife diseases in biodiversity and public health it is critical to accurately forecast pathogen dissemination while incorporating relevant spatio-temporal covariates. Forecasting spatio-temporal processes can often be improved by incorporating scientific knowledge about the dynamics of the process using physical models. Ecological diffusion equations are often used to model epidemiological processes of wildlife diseases where environmental factors play a role in disease spread. Physics-informed neural networks (PINN) are deep learning algorithms that constrain neural network predictions based on physical laws and therefore are powerful forecasting models useful even in cases of limited and imperfect training data. In this paper, we develop a novel ecological modeling tool using PINNs, which fits a feedforward neural network and simultaneously performs parameter identification in a partial differential equation (PDE) with varying coefficients. We demonstrate the applicability of our model by comparing it with the commonly used Bayesian stochastic partial differential equation method and traditional machine learning approaches, showing that our proposed model exhibits superior prediction and forecasting performance when modeling chronic wasting disease in deer in Wisconsin. Furthermore, our model provides the opportunity to obtain scientific insights into spatiotemporal covariates affecting spread and growth of diseases. This work contributes to future machine learning and statistical methodology development by studying spatio-temporal processes enhanced by prior physical knowledge.

Spatial Statistics

Basis function models for animal movement

Advances in satellite-based data collection techniques have served as a catalyst for new statistical methodology to analyze these data. In wildlife ecological studies, satellite-based data and methodology have provided a wealth of information about animal space use and the investigation of individual-based animal–environment relationships. With the technology for data collection improving dramatically over time, we are left with massive archives of historical animal telemetry data of varying quality. While many contemporary statistical approaches for inferring movement behavior are specified in discrete time, we develop a flexible continuous-time stochastic integral equation framework that is amenable to reduced-rank second-order covariance parameterizations. We demonstrate how the associated first-order basis functions can be constructed to mimic behavioral characteristics in realistic trajectory processes using telemetry data from mule deer and mountain lion individuals in western North America. Our approach is parallelizable and provides inference for heterogenous trajectories using nonstationary spatial modeling techniques that are feasible for large telemetry datasets. Supplementary materials for this article are available online.

Journal of the American Statistical Association

A general science-based framework for dynamical spatio-temporal models

Spatio-temporal statistical models are increasingly being used across a wide variety of scientific disciplines to describe and predict spatially-explicit processes that evolve over time. Correspondingly, in recent years there has been a significant amount of research on new statistical methodology for such models. Although descriptive models that approach the problem from the second-order (covariance) perspective are important, and innovative work is being done in this regard, many real-world processes are dynamic, and it can be more efficient in some cases to characterize the associated spatio-temporal dependence by the use of dynamical models. The chief challenge with the specification of such dynamical models has been related to the curse of dimensionality. Even in fairly simple linear, first-order Markovian, Gaussian error settings, statistical models are often over parameterized. Hierarchical models have proven invaluable in their ability to deal to some extent with this issue by allowing dependency among groups of parameters. In addition, this framework has allowed for the specification of science based parameterizations (and associated prior distributions) in which classes of deterministic dynamical models (e. g., partial differential equations (PDEs), integro-difference equations (IDEs), matrix models, and agent-based models) are used to guide specific parameterizations. Most of the focus for the application of such models in statistics has been in the linear case. The problems mentioned above with linear dynamic models are compounded in the case of nonlinear models. In this sense, the need for coherent and sensible model parameterizations is not only helpful, it is essential. Here, we present an overview of a framework for incorporating scientific information to motivate dynamical spatio-temporal models. First, we illustrate the methodology with the linear case. We then develop a general nonlinear spatio-temporal framework that we call general quadratic nonlinearity and demonstrate that it accommodates many different classes of scientific-based parameterizations as special cases. The model is presented in a hierarchical Bayesian framework and is illustrated with examples from ecology and oceanography. ?? 2010 Sociedad de Estad??stica e Investigaci??n Operativa.

Test

Microsatellites: Evolutionary and methodological background and empirical applications at individual, population, and phylogenetic levels

The recent proliferation and greater accessibility of molecular genetic markers has led to a growing appreciation of the ecological and evolutionary inferences that can be drawn from molecular characterizations of individuals and populations (Burke et al. 1992, Avise 1994). Different techniques have the ability to target DNA sequences which have different patterns of inheritance, different modes and rates of evolution and, concomitantly, different levels of variation. In the quest for 'the right marker for the right job', microsatellites have been widely embraced as the marker of choice for many empirical genetic studies. The proliferation of microsatellite loci for various species and the voluminous literature compiled in very few years associated with their evolution and use in various research applications, exemplifies their growing importance as a research tool in the biological sciences. The ability to define allelic states based on variation at the nucleotide level has afforded unparalleled opportunities to document the actual mutational process and rates of evolution at individual microsatellite loci. The scrutiny to which these loci have been subjected has resulted in data that raise issues pertaining to assumptions formerly stated, but largely untestable for other marker classes. Indeed this is an active arena for theoretical and empirical work. Given the extensive and ever-increasing literature on various statistical methodologies and cautionary notes regarding the uses of microsatellites, some consideration should be given to the unique characteristics of these loci when determining how and under what conditions they can be employed.

Book chapter

Caution is warranted when using animal space-use and movement to infer behavioral states

Background Identifying the behavioral state for wild animals that can’t be directly observed is of growing interest to the ecological community. Advances in telemetry technology and statistical methodologies allow researchers to use space-use and movement metrics to infer the underlying, latent, behavioral state of an animal without direct observations. For example, researchers studying ungulate ecology have started using these methods to quantify behaviors related to mating strategies. However, little work has been done to determine if assumed behaviors inferred from movement and space-use patterns correspond to actual behaviors of individuals. Methods Using a dataset with male and female white-tailed deer location data, we evaluated the ability of these two methods to correctly identify male-female interaction events (MFIEs). We identified MFIEs using the proximity of their locations in space as indicators of when mating could have occurred. We then tested the ability of utilization distributions (UDs) and hidden Markov models (HMMs) rendered with single sex location data to identify these events. Results For white-tailed deer, male and female space-use and movement behavior did not vary consistently when with a potential mate. There was no evidence that a probability contour threshold based on UD volume applied to an individual’s UD could be used to identify MFIEs. Additionally, HMMs were unable to identify MFIEs, as single MFIEs were often split across multiple states and the primary state of each MFIE was not consistent across events. Conclusions Caution is warranted when interpreting behavioral insights rendered from statistical models applied to location data, particularly when there is no form of validation data. For these models to detect latent behaviors, the individual needs to exhibit a consistently different type of space-use and movement when engaged in the behavior. Unvalidated assumptions about that relationship may lead to incorrect inference about mating strategies or other behaviors.

Movement Ecology

Effects of environmental variables on invasive amphibian activity: Using model selection on quantiles for counts

Many different factors influence animal activity. Often, the value of an environmental variable may influence significantly the upper or lower tails of the activity distribution. For describing relationships with heterogeneous boundaries, quantile regressions predict a quantile of the conditional distribution of the dependent variable. A quantile count model extends linear quantile regression methods to discrete response variables, and is useful if activity is quantified by trapping, where there may be many tied (equal) values in the activity distribution, over a small range of discrete values. Additionally, different environmental variables in combination may have synergistic or antagonistic effects on activity, so examining their effects together, in a modeling framework, is a useful approach. Thus, model selection on quantile counts can be used to determine the relative importance of different variables in determining activity, across the entire distribution of capture results. We conducted model selection on quantile count models to describe the factors affecting activity (numbers of captures) of cane toads ( Rhinella marina ) in response to several environmental variables (humidity, temperature, rainfall, wind speed, and moon luminosity) over eleven months of trapping. Environmental effects on activity are understudied in this pest animal. In the dry season, model selection on quantile count models suggested that rainfall positively affected activity, especially near the lower tails of the activity distribution. In the wet season, wind speed limited activity near the maximum of the distribution, while minimum activity increased with minimum temperature. This statistical methodology allowed us to explore, in depth, how environmental factors influenced activity across the entire distribution, and is applicable to any survey or trapping regime, in which environmental variables affect activity.

Ecosphere

Assessing giant sequoia mortality and regeneration following high-severity wildfire

Fire is a critical driver of giant sequoia ( Sequoiadendron giganteum [Lindl.] Buchholz) regeneration. However, fire suppression combined with the effects of increased temperature and severe drought has resulted in fires of an intensity and size outside of the historical norm. As a result, recent mega-fires have killed a significant portion of the world's sequoia population (13%–19%), and uncertainty surrounds whether severely affected groves will be able to recover naturally, potentially leading to a loss of grove area. To assess the likelihood of natural recovery, we collected spatially explicit data assessing mortality, crown condition, and regeneration within four giant sequoia groves that were severely impacted by the SQF- (2020) and KNP-Complex (2021) wildfires within Sequoia and Kings Canyon National Parks. In total, we surveyed 5.9 ha for seedlings and assessed the crown condition of 1104 giant sequoias. To inform management, we used a statistical methodology that robustly quantifies the uncertainty in inherently “noisy” seedling data and takes advantage of readily available remote sensing metrics that would make our findings applicable to other recently burned groves. A loss of giant sequoia grove area would be a consequence of giant sequoia tree mortality followed by a failure of natural regeneration. We found that areas that experienced very high-severity fire (above ~800 RdNBR) are at substantial risk for the loss of grove area, with tree mortality rapidly increasing and giant sequoia seedling density simultaneously decreasing with fire severity. Such high-severity areas comprised 17.8, 142.0, 14.6, 1.6 ha and ~90%, ~14%, ~53%, and ~27% of Board Camp, Redwood Mountain, Suwanee, and New Oriole Lake groves, respectively. In all sampling areas, we found that seedling densities fell far below the average density measured after prescribed fires, where seedling numbers were almost certainly adequate to maintain giant sequoia populations and postfire conditions were more in keeping with historical norms. Importantly, spatial pattern is also important in assessing the risk of grove loss, and in two groves, Suwanee and New Oriole Lake, the high-severity patches were not always contiguous, potentially making some areas more resilient to regeneration failure due to the proximity of surviving trees.

Ecosphere

Quantifying the relative importance of biotic and abiotic factors in landscape-based models of stream fish distributions

Lotic fish species distributions are frequently predicted using remotely sensed habitat variables that characterize the adjacent landscape and serve as proxies for instream habitat. Recent advancements in statistical methodology, however, allow for leveraging fish assemblage data when predicting distributions. This is important because assemblage composition likely provides better information about instream habitat compared to landscape-derived metrics and therefore may improve predictions. To better understand the value of using multi-species fish data in species distribution modeling, we fit two conditional random fields (CRF) models to quantify the relative importance of fish assemblage co-occurrence, landscape-derived habitat variables, and interactions between these two predictor groups (i.e., effects of co-occurrence could be context-dependent) at over 1200 stream catchments in Pennsylvania, USA. We first compared predictive performance of CRF models against traditionally used single-species logistic regressions (generalized linear models; GLMs) and found that inclusion of fish assemblage data often improved predictive performance. The multi-species CRF models performed significantly better at predicting occurrence for 63% of species with an average percent increase in AUC of 25% compared to GLMs. Furthermore, the CRF identified species co-occurrences as more informative, and thus relatively more important, at predicting occurrence than the other effect types. The CRF also suggested that allowing these biotic effects to be context-dependent was important for predicting occurrence of many species. These findings illustrate the value of fish assemblage data for landscape-scale species distribution modeling and leveraging this information can improve predictions and inferences to help inform the management and conservation of freshwater fishes.

Pennsylvania

Combining terrestrial lidar with single line transects to investigate geomorphic change: A case study on the Upper Verde River, Arizona

The Upper Verde River in northern Arizona, USA is a vital resource for the wildlife and humans that rely on its waters. We characterize the riparian corridor topography using terrestrial laser scanner (TLS) data from 2021 to 2022. We also quantify geomorphic changes associated with human and climate-driven alterations in river flow and vegetation changes by combining the contemporary lidar surveys with legacy measurements from single line geomorphology transects measured by the United States Forest Service (USFS) in 2009. Seventeen plots along the Upper Verde River were surveyed with the TLS and the data were coregistered within individual plots with a Root Mean Square Error of <0.03 m among scan positions. Digital Elevation Models (DEM) were derived for each plot from the TLS data at 10 cm resolution and compared to the 2009 USFS cross-section data to quantify elevation changes. In areas with statistically significant change, we detected maximum changes in elevation due to erosion and deposition of −0.37 m and + 0.97 m, respectively. Topographic changes over the 13-year period were predominately aggradation and associated with sediment deposition, which we hypothesize might have resulted from altered river flow and vegetation encroachment. This study also demonstrates a quantitative and statistical methodology to fuse traditional single line cross-section data with contemporary lidar data to quantify geomorphic change. The novel approach demonstrated here is broadly applicable to natural resource managers for integrating and contextualizing legacy topographic data for understanding past, present, and future landscape and habitat changes.

Arizona

On the estimation of dispersal and movement of birds

The estimation of dispersal and movement is important to evolutionary and population ecologists, as well as to wildlife managers. We review statistical methodology available to estimate movement probabilities. We begin with cases where individual birds can be marked and their movements estimated with the use of multisite capture-recapture methods. Movements can be monitored either directly, using telemetry, or by accounting for detection probability when conventional marks are used. When one or more sites are unobservable, telemetry, band recoveries, incidental observations, a closed- or open-population robust design, or partial determinism in movements can be used to estimate movement. When individuals cannot be marked, presence-absence data can be used to model changes in occupancy over time, providing indirect inferences about movement. Where abundance estimates over time are available for multiple sites, potential coupling of their dynamics can be investigated using linear cross-correlation or nonlinear dynamic tools.

Condor

A functional model for characterizing long-distance movement behaviour

Advancements in wildlife telemetry techniques have made it possible to collect large data sets of highly accurate animal locations at a fine temporal resolution. These data sets have prompted the development of a number of statistical methodologies for modelling animal movement. Telemetry data sets are often collected for purposes other than fine-scale movement analysis. These data sets may differ substantially from those that are collected with technologies suitable for fine-scale movement modelling and may consist of locations that are irregular in time, are temporally coarse or have large measurement error. These data sets are time-consuming and costly to collect but may still provide valuable information about movement behaviour. We developed a Bayesian movement model that accounts for error from multiple data sources as well as movement behaviour at different temporal scales. The Bayesian framework allows us to calculate derived quantities that describe temporally varying movement behaviour, such as residence time, speed and persistence in direction. The model is flexible, easy to implement and computationally efficient. We apply this model to data from Colorado Canada lynx ( Lynx canadensis ) and use derived quantities to identify changes in movement behaviour.

Methods in Ecology and Evolution

Time-series model, statistical methods, and software documentation for R–QWTREND—An R package for analyzing trends in stream-water quality

As part of a U.S. Geological Survey water-quality study started in 2018, in cooperation with the International Joint Commission, North Dakota Department of Environmental Quality, and Minnesota Pollution Control Agency, a publicly available software package called R–QWTREND was developed for analyzing trends in stream-water quality. The R–QWTREND package is a collection of functions written in R, an open source language and a general environment for statistical computing and graphics. The package uses a parametric time-series model to express logarithmically transformed concentration in terms of flow-related variability, trend, and serially correlated model errors. Flow-related variability captures natural variability in concentration on the basis of concurrent and antecedent streamflow. The trend identifies systematic changes in concentration in terms of potential step trends, piecewise monotonic trends, or user-specified trends. Maximum likelihood estimation is used to estimate model parameters and determine the best-fit trend model. This report describes the time-series model and statistical methodology behind R–QWTREND and provides formal documentation for installing and using the package.

Open-File Report

Decision-support framework for linking regional-scale management actions to continental-scale conservation of wide-ranging species

Anas acuta (Northern pintail; hereafter pintail) was selected as a model species on which to base a decision-support framework linking regional actions to continental-scale population and harvest objectives. This framework was then used to engage stakeholders, such as Landscape Conservation Cooperatives’ (LCCs’) habitat management partners within areas of importance to pintails, while maximizing cross-taxa effects from the framework. The mathematical framework for the model had been previously developed for pintails. A key assumption incorporated into the model is that density dependence in survival occurs during the post-hunting (winter) period, where resources are hypothesized to be limiting. Because few data are available to directly inform this process, the approach used was to build a hierarchical Bayesian integrated population model (IPM) that simultaneously uses data from bird-band recoveries, breeding population counts, and harvest surveys to estimate values of parameters of an annual population projection model, including population size, survival rate, reproductive rate, and process and observation error variances, that are logically consistent with each other, given the mathematical structure imposed through the IPM. The main accomplishments of this study are (1) development of an IPM for pintail to guide harvest and habitat management, (2) development of a Prairie Parkland Region breeding submodel to predict pintail productivity, (3) development of statistical methodology to estimate pintail productivity (as measured by the ratio of juvenile to adults in hunter-collected wing samples) and winter survival and to relate these estimates to covariates, and (4) illustration of how to use a model and estimated parameter values to predict pintail population size and sustainable harvest as a function of habitat. Estimation of pintail survival from bird-banding data shows that there has been relatively little variation in survival over the period 1960–2013. A productivity model showed strong effects of breeding ground conditions, wintering-ground precipitation, and density dependence on pintail productivity. Thus, most temporal variation in pintail demographic rates has been due to effects on reproduction and not survival, including effects of breeding or wintering-ground habitat. These results indicate that habitat conservation efforts may be most effective if they focus on maintaining or increasing breeding and wintering-ground habitat to increase pintail productivity rather than pintail survival. Environmental perturbations in excess of historical experience, such as what could occur under climate change, might have meaningful effects on survival but cannot be estimated with current data. Direct effects of climate, land use, or management are likely to be greater on productivity than survival, but substantial uncertainty remains about predictions of equilibrium population size and sustainable yield.

Open-File Report

Estimating total population size for adult female sea turtles: Accounting for non-nesters

Assessment of population size and changes therein is important to sea turtle management and population or life history research. Investigators might be interested in testing hypotheses about the effect of current population size or density (number of animals per unit resource) on future population processes. Decision makers might want to determine a level of allowable take of individual turtles of specified life stage. Nevertheless, monitoring most stages of sea turtle life histories is difficult, because obtaining access to individuals is difficult. Although in-water assessments are becoming more common, nesting females and their hatchlings remain the most accessible life stages. In some cases adult females of a given nesting population are sufficiently philopatric that the population itself can be well defined. If a well designed tagging study is conducted on this population, survival, breeding probability, and the size of the nesting population in a given year can be estimated. However, with published statistical methodology the size of the entire breeding population (including those females skipping nesting in that year) cannot be estimated without assuming that each adult female in this population has the same probability of nesting in a given year (even those that had just nested in the previous year). We present a method for estimating the total size of a breeding population (including nesters those skipping nesting) from a tagging study limited to the nesting population, allowing for the probability of nesting in a given year to depend on an individual's nesting status in the previous year (i.e., a Markov process). From this we further develop estimators for rate of growth from year to year in both nesting population and total breeding population, and the proportion of the breeding population that is breeding in a given year. We also discuss assumptions and apply these methods to a breeding population of hawksbill sea turtles (Eretmochelys imbricata) from the Caribbean. We anticipate that this method could also be useful for in-water studies of well defined populations.

Book chapter

Improved population estimates through the use of auxiliary information

When estimating the size of a population of birds, the investigator may have, in addition to an estimator based on a statistical sample, information on one of several auxiliary variables, such as: (1) estimates of the population made on previous occasions, (2) measures of habitat variables associated with the size of the population, and (3) estimates of the population sizes of other species that correlate with the species of interest. Although many studies have described the relationships between each of these kinds of data and the population size to be estimated, very little work has been done to improve the estimator by incorporating such auxiliary information. A statistical methodology termed 'empirical Bayes' seems to be appropriate to these situations. The potential that empirical Bayes methodology has for improved estimation of the population size of the Mallard (Anas platyrhynchos) is explored. In the example considered, three empirical Bayes estimators were found to reduce the error by one-fourth to one-half of that of the usual estimator.

Book chapter