Search USGSSearch

SEARCH · Search USGS

Results for “Journal of Statistical Software”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Unmarked: An R package for fitting hierarchical models of wildlife occurrence and abundance

Ecological research uses data collection techniques that are prone to substantial and unique types of measurement error to address scientific questions about species abundance and distribution. These data collection schemes include a number of survey methods in which unmarked individuals are counted, or determined to be present, at spatially- referenced sites. Examples include site occupancy sampling, repeated counts, distance sampling, removal sampling, and double observer sampling. To appropriately analyze these data, hierarchical models have been developed to separately model explanatory variables of both a latent abundance or occurrence process and a conditional detection process. Because these models have a straightforward interpretation paralleling mechanisms under which the data arose, they have recently gained immense popularity. The common hierarchical structure of these models is well-suited for a unified modeling interface. The R package unmarked provides such a unified modeling framework, including tools for data exploration, model fitting, model criticism, post-hoc analysis, and model comparison.

Journal of Statistical Software

colorspace: A toolbox for manipulating and assessing colors and palettes

The R package colorspace provides a flexible toolbox for selecting individual colors or color palettes, manipulating these colors, and employing them in statistical graphics and data visualizations. In particular, the package provides a broad range of color palettes based on the HCL (hue-chroma-luminance) color space. The three HCL dimensions have been shown to match those of the human visual system very well, thus facilitating intuitive selection of color palettes through trajectories in this space. Using the HCL color model, general strategies for three types of palettes are implemented: (1) Qualitative for coding categorical information, i.e., where no particular ordering of categories is available. (2) Sequential for coding ordered/numeric information, i.e., going from high to low (or vice versa). (3) Diverging for coding ordered/numeric information around a central neutral value, i.e., where colors diverge from neutral to two extremes. To aid selection and application of these palettes, the package also contains scales for use with ggplot2, shiny and tcltk apps for interactive exploration, visualizations of palette properties, accompanying manipulation utilities (like desaturation and lighten/darken), and emulation of color vision deficiencies. The shiny apps are also hosted online at http://hclwizard.org/.

Journal of Statistical Software

Hierarchical models of animal abundance and occurrence

Much of animal ecology is devoted to studies of abundance and occurrence of species, based on surveys of spatially referenced sample units. These surveys frequently yield sparse counts that are contaminated by imperfect detection, making direct inference about abundance or occurrence based on observational data infeasible. This article describes a flexible hierarchical modeling framework for estimation and inference about animal abundance and occurrence from survey data that are subject to imperfect detection. Within this framework, we specify models of abundance and detectability of animals at the level of the local populations defined by the sample units. Information at the level of the local population is aggregated by specifying models that describe variation in abundance and detection among sites. We describe likelihood-based and Bayesian methods for estimation and inference under the resulting hierarchical model. We provide two examples of the application of hierarchical models to animal survey data, the first based on removal counts of stream fish and the second based on avian quadrat counts. For both examples, we provide a Bayesian analysis of the models using the software WinBUGS.

Journal of Agricultural, Biological, and Environme

Fish size structure analysis via ordination: A visualization aid

Objective Visual aids like length-frequency histograms are widely used to examine fish population status and trends; however, comparing multiple histograms simultaneously becomes cumbersome and inefficient. Complicating matters further, overlaying covariates on histograms to highlight connections with length frequencies can be challenging. An alternative, and the subject of this Perspective, is to display length distributions as an ordination using similarity indexes; in many cases, this allows for improved visual organization and representation of relationships with covariates. Methods I review the application of ordination methods for analysis of size structures using alternative visualizations that may facilitate the identification of connections that are concealed when analyzing a series of histograms. After a brief introduction to similarity indexes, types of ordinations, and sample sizes, I examine four case studies to illustrate size structure analysis via similarity indices: (1) unconstrained ordination to identify “bass-crowded” populations in a set of 34 small fishing lakes, (2) unconstrained ordination to evaluate the effect of three consecutive length limits on a Largemouth Bass Micropterus nigricans population over a span of 28 years, (3) constrained ordination to assess the relationships between fish community size structure and in-lake and off-lake environmental descriptors in 30 oxbow lakes, and (4) constrained ordination to identify what aspects of Largemouth Bass size structure were related to six types of reservoir habitats. Result Size structure analysis via similarity indexes enabled the exploration of extensive length-frequency data. It is important to acknowledge that ordinations serve solely as a visual aid for assessing size structure—no statistical testing is involved. Conclusion Ordination techniques and software are advancing at a quick pace, holding great promise for the future of size structure analysis via similarity indices.

North American Journal of Fisheries Management

Using a Geographic Information System to determine the relation between stream quality and geology in the Roberts Creek watershed, Clayton County, Iowa

A geographic information system (GIS) was used to determine the relation between the stream-water quality and underlying geology in Roberts Creek watershed, Clayton County, Iowa, for base-flow conditions during the spring and summer of 1988–90. Geologic, stream, basin and subbasin boundaries, and water-quality sampling-site coverages were created by digitizing available maps. A contour coverage was created from digital line-graph data. The areal extent of geologic units subcropping in each subbasin was quantified with GIS, and the results then were output and joined with the discharge and water-quality data for statistical analyses. Illustrations showing the geology of the study area and the results of the study were prepared using GIS. By using GIS and a statistical software package, a weak but statistically significant relation was found between the water temperature, pH, and nitrogen concentrations in Roberts Creek and the underlying geology during base-flow conditions.

Iowa

A practical guide to understanding and validating complex models using data simulations

Biologists routinely fit novel and complex statistical models to push the limits of our understanding. Examples include, but are not limited to, flexible Bayesian approaches (e.g. BUGS, stan), frequentist and likelihood-based approaches (e.g. packages lme4 ) and machine learning methods. These software and programs afford the user greater control and flexibility in tailoring complex hierarchical models. However, this level of control and flexibility places a higher degree of responsibility on the user to evaluate the robustness of their statistical inference. To determine how often biologists are running model diagnostics on hierarchical models, we reviewed 50 recently published papers in 2021 in the journal Nature Ecology & Evolution , and we found that the majority of published papers did not report any validation of their hierarchical models, making it difficult for the reader to assess the robustness of their inference. This lack of reporting likely stems from a lack of standardized guidance for best practices and standard methods. Here, we provide a guide to understanding and validating complex models using data simulations. To determine how often biologists use data simulation techniques, we also reviewed 50 recently published papers in 2021 in the journal Methods Ecology & Evolution . We found that 78% of the papers that proposed a new estimation technique, package or model used simulations or generated data in some capacity (18 of 23 papers); but very few of those papers (5 of 23 papers) included either a demonstration that the code could recover realistic estimates for a dataset with known parameters or a demonstration of the statistical properties of the approach. To distil the variety of simulations techniques and their uses, we provide a taxonomy of simulation studies based on the intended inference. We also encourage authors to include a basic validation study whenever novel statistical models are used, which in general, is easy to implement. Simulating data helps a researcher gain a deeper understanding of the models and their assumptions and establish the reliability of their estimation approaches. Wider adoption of data simulations by biologists can improve statistical inference, reliability and open science practices.

Methods in Ecology and Evolution

Methods to quality assure, plot, summarize, interpolate, and extend groundwater-level information—Examples for the Mississippi River Valley alluvial aquifer

Large-scale computational investigations of groundwater levels are proposed to accelerate science delivery through a workflow spanning database assembly, statistics, and information synthesis and packaging. A water-availability study of the Mississippi River alluvial plain, and particularly the Mississippi River Valley alluvial aquifer (MRVA), is ongoing. Software (visGWDBmrva) has been released as part of the study that demonstrates groundwater informatics for the aquifer. Considerable water-level data collected by multiple agencies over a seven-state area exist (18,903 wells; 287,272 measurements [April 22, 2019]). Data and metadata quality assurance methods, basic statistics, hydrograph visualization, outlier identification, hypothesis testing, and time-series modeling are described. Two approaches (generalized additive models [GAMs] and support vector machines [SVMs]) are used for data interpolation and extension to monthly water-level estimates. Numerical congruence between GAM and SVM estimates will be useful to limit inclusion of monthly estimates from subsequent science activities.

Mississippi River valley

Using the beta distribution to analyze plant cover data

Most plant species are spatially aggregated. Local demographic and ecological processes (e.g. vegetative growth and limited seed dispersal) result in a clustered spatial pattern within an environmentally homogenous area. Spatial aggregation should be considered when modelling plant abundance data. Commonly, plant abundance is quantified by measuring cover within multiple areal plots or along multiple lines randomly placed within a study area. A common practice for analyzing plant cover is to use statistical methods that rely on the normal distribution for quantifying uncertainty. This is problematic because plant cover data tend to be left‐skewed (J‐shaped), right skewed (L‐shaped) or U‐shaped and, therefore, commonly violate classic statistical assumptions, such as normality. We outline statistical analyses that explicitly account for spatial aggregation by assuming that plant cover is beta‐distributed. The beta distribution is a flexible choice because within the open unit interval it can take on a wide range of shapes (L, J, U, or a bell‐shaped). We discuss and introduce extensions to the beta distribution that address common analysis issues encountered in plant cover datasets, such as i) the treatment of zero and one cover values, ii) hierarchical data structures, and iii) observations errors. For heuristic purposes, we focus on single species analyses, but we demonstrate how the outlined methods can be generalized to more species. The assumption that plant cover is beta‐distributed allows us to estimate the degree of spatial aggregation, and the ecological significance of this new knowledge is discussed. We provide a summary of available software for analyses (emphasizing standard R packages) and include worked examples and a simulation study comparing analysis options as supplemental information. Synthesis . Previously, the state of the statistical software made it practically difficult for empirical plant ecologists to analyze their cover data correctly, but new theory and R‐packages have been developed, and this difficulty no longer exists. We recommend that empirical plant ecologists embrace the new statistical possibilities for exploring the exciting ecological features in spatial variation of plant cover.

Journal of Ecology

Subsampling large-scale digital elevation models to expedite geospatial analyses in coastal regions

Large-area, high-resolution digital elevation models (DEMs) created from light detection and ranging (LIDAR) and/or multibeam echosounder data sets are commonly used in many scientific disciplines. These DEMs can span thousands of square kilometers, typically with a spatial resolution of 1 m or finer, and can be difficult to process and analyze without specialized computers and software. Such DEMs often can be subsampled to expedite analysis with negligible impact on results for large-scale geospatial analyses. Subsampling can be achieved by creating a grid of points that specify the locations from which to extract elevation values from the DEM. This paper presents a method that can be used to accurately perform subsampling of large-scale, high-resolution DEMs using GIS software. This subsampling method was applied to two LIDAR-derived DEMs encompassing 242 km 2 of the northern Florida Reef Tract as an example application and to test subsampling accuracy. Results indicate that subsampling 1-m-resolution DEMs using a 2-m-spaced grid results in no significant difference in mean elevation or other basic statistics for analyses performed over multiple spatial scales ranging from 1 km 2 to 242 km 2 .

Journal of Coastal Research

Journal news

Statistical power (and conversely, Type II error) is often ignored by biologists. Power is important to consider in the design of studies, to ensure that sufficient resources are allocated to address a hypothesis under examination. Deter- mining appropriate sample size when designing experiments or calculating power for a statistical test requires an investigator to consider the importance of making incorrect conclusions about the experimental hypothesis and the biological importance of the alternative hypothesis (or the biological effect size researchers are attempting to measure). Poorly designed studies frequently provide results that are at best equivocal, and do little to advance science or assist in decision making. Completed studies that fail to reject Ho should consider power and the related probability of a Type II error in the interpretation of results, particularly when implicit or explicit acceptance of Ho is used to support a biological hypothesis or management decision. Investigators must consider the biological question they wish to answer (Tacha et al. 1982) and assess power on the basis of biologically significant differences (Taylor and Gerrodette 1993). Power calculations are somewhat subjective, because the author must specify either f or the minimum difference that is biologically important. Biologists may have different ideas about what values are appropriate. While determining biological significance is of central importance in power analysis, it is also an issue of importance in wildlife science. Procedures, references, and computer software to compute power are accessible; therefore, authors should consider power. We welcome comments or suggestions on this subject.

Journal of Wildlife Management

Validation of streamflow measurements made with M9 and RiverRay acoustic Doppler current profilers

The U.S. Geological Survey (USGS) Office of Surface Water (OSW) previously validated the use of Teledyne RD Instruments (TRDI) Rio Grande (in 2007), StreamPro (in 2006), and Broadband (in 1996) acoustic Doppler current profilers (ADCPs) for streamflow (discharge) measurements made by the USGS. Two new ADCPs, the SonTek M9 and the TRDI RiverRay, were first used in the USGS Water Mission Area programs in 2009. Since 2009, the OSW and USGS Water Science Centers (WSCs) have been conducting field measurements as part of their stream-gaging program using these ADCPs. The purpose of this paper is to document the results of USGS OSW analyses for validation of M9 and RiverRay ADCP streamflow measurements. The OSW required each participating WSC to make comparison measurements over the range of operating conditions in which the instruments were used until sufficient measurements were available. The performance of these ADCPs was evaluated for validation and to identify any present and potential problems. Statistical analyses of streamflow measurements indicate that measurements made with the SonTek M9 ADCP using firmware 2.00–3.00 or the TRDI RiverRay ADCP using firmware 44.12–44.15 are unbiased, and therefore, can continue to be used to make streamflow measurements in the USGS stream-gaging program. However, for the M9 ADCP, there are some important issues to be considered in making future measurements. Possible future work may include additional validation of streamflow measurements made with these instruments from other locations in the United States and measurement validation using updated firmware and software.

Journal of Hydraulic Engineering

An approach to modeling abundance of marine wildlife over space and time using unstructured aerial surveys

Estimating spatial and temporal patterns in abundance is often a goal of ecological studies and can be useful for informing management decisions, such as determining the optimal placement of wildlife protection zones. However, estimating abundance can be difficult in practice, especially over large areas, because of imperfect detection, where individuals are present but not detected because of either availability or observer error. Several methods for estimating abundance that account for imperfect detection exist but can be logistically challenging to implement. We present a simpler approach to some of the more commonly used techniques for estimating the abundance of marine wildlife over space and time from unstructured aerial surveys. This approach combines a spatial model for count data with auxiliary information on detection probability obtained from small-scale or previous studies. We employ generalized linear models and generalized additive models with spatial habitat covariates to illustrate this approach using maximum-likelihood with free, open-source statistical software. This framework is intended to be accessible and flexible, requiring lower survey costs and less computation time than other alternatives for estimating abundance. Indeed, our simulation results show that this approach can reduce computation times, while appropriately characterizing uncertainty, compared to a Bayesian approach. We also present R code for our approach using an example of estimating Florida manatee ( Trichechus manatus latirostris ) abundance in Indian River County in Florida, USA. This approach could be applied to other study systems and marine wildlife species using unstructured aerial surveys.

Florida

FishStan: Hierarchical Bayesian models for fisheries

Fisheries managers and ecologists use statistical models to estimate population-level relations and demographic rates (e.g., length-maturity curves, growth curves, and mortality rates). These relations and rates provide insight into populations and inputs for other models. For example, growth curves may vary across lakes showing fish populations differ due to management actions or underlying environmental conditions. A fisheries manager could use this information to set lake-specific harvest limits or an ecologist could use this information to test scientific hypotheses about fish populations. The above example also demonstrates how populations exist within hierarchical structures where sub-populations may be nested within a meta-population. More generally, these hierarchical structures may be both biological (e.g., different lakes or river pools) and statistical (e.g., correlated error structures). Currently, limited options exist for fitting these hierarchical models and people seeking to use them often must program their own implementations. Furthermore, many fisheries managers and researchers may not have Bayesian programming skills, but many can use interactive languages such as R. Additionally, programs such as JAGS often require long run times (e.g., hours if not days) to fit hierarchical models and programs such as Stan can be more difficult to program because it is a compiled language. We created fishStan to share hierarchical models for fisheries and ecology in an easy-to-use R package.

Journal of Open Source Software

Characterizing groundwater/surface-water interaction using hydrograph-separation techniques and groundwater-level data throughout the Mississippi Delta, USA

The Mississippi Delta, located in northwest Mississippi, is an area dense with industrial-level agriculture sustained by groundwater-dependent irrigation supplied by the Mississippi River Valley Alluvial aquifer (alluvial aquifer). The Delta provides agricultural commodities across the United States and around the world. Observed declines in groundwater altitudes and streamflow contemporaneous with increases in irrigation have raised concerns about future groundwater availability and the effects of groundwater withdrawals on streamflow. To quantify the impacts of groundwater withdrawals on streamflow and increase understanding of groundwater and surface-water interaction, hydrograph-separation techniques were used to estimate baseflow and identify statistical streamflow trends. The analysis was conducted using the U.S. Geological Survey Groundwater Toolbox open-source software and daily hydrologic data provided by a spatially-distributed network of paired groundwater wells and streamgaging sites. This study found that effects of groundwater withdrawals on streamflow were observed as statistically significant reductions in baseflow in areas with substantial groundwater-altitude declines. Hydrograph-separation and trend analyses may be applicable to assess the impacts of groundwater withdrawals in altered environments and streamflow may be used as a proxy for changes in groundwater availability. Characterizing and defining hydrologic relations between groundwater and surface water will help scientists and water-resource managers refine a regional groundwater-flow model that includes the Mississippi Delta that will be used to aid water-resource managers in future decisions concerning the alluvial aquifer.

Arkansas, Illinois, Kentucky, Louisiana, Mississip

Estimating animal resource selection from telemetry data using point process models

Analyses of animal resource selection functions (RSF) using data collected from relocations of individuals via remote telemetry devices have become commonplace. Increasing technological advances, however, have produced statistical challenges in analysing such highly autocorrelated data. Weighted distribution methods have been proposed for analysing RSFs with telemetry data. However, they can be computationally challenging due to an intractable normalizing constant and cannot be aggregated (i.e. collapsed) over time to make space-only inference. In this study, we take a conceptually different approach to modelling animal telemetry data for making RSF inference. We consider the telemetry data to be a realization of a space–time point process. Under the point process paradigm, the times of the relocations are also considered to be random rather than fixed. We show the point process models we propose are a generalization of the weighted distribution telemetry models. By generalizing the weighted model, we can access several numerical techniques for evaluating point process likelihoods that make use of common statistical software. Thus, the analysis methods can be readily implemented by animal ecologists. In addition to ease of computation, the point process models can be aggregated over time by marginalizing over the temporal component of the model. This allows a full range of models to be constructed for RSF analysis at the individual movement level up to the study area level. To demonstrate the analysis of telemetry data with the point process approach, we analysed a data set of telemetry locations from northern fur seals (Callorhinus ursinus) in the Pribilof Islands, Alaska. Both a space–time and an aggregated space-only model were fitted. At the individual level, the space–time analysis showed little selection relative to the habitat covariates. However, at the study area level, the space-only model showed strong selection relative to the covariates.

Alaska

Of bugs and birds: Markov Chain Monte Carlo for hierarchical modeling in wildlife research

Markov chain Monte Carlo (MCMC) is a statistical innovation that allows researchers to fit far more complex models to data than is feasible using conventional methods. Despite its widespread use in a variety of scientific fields, MCMC appears to be underutilized in wildlife applications. This may be due to a misconception that MCMC requires the adoption of a subjective Bayesian analysis, or perhaps simply to its lack of familiarity among wildlife researchers. We introduce the basic ideas of MCMC and software BUGS (Bayesian inference using Gibbs sampling), stressing that a simple and satisfactory intuition for MCMC does not require extraordinary mathematical sophistication. We illustrate the use of MCMC with an analysis of the association between latent factors governing individual heterogeneity in breeding and survival rates of kittiwakes ( Rissa tridactyla ). We conclude with a discussion of the importance of individual heterogeneity for understanding population dynamics and designing management plans.

Journal of Wildlife Management

Classifying physiographic regimes on terrain and hydrologic factors for adaptive generalization of stream networks

Automated generalization software must accommodate multi-scale representations of hydrographic networks across a variety of geographic landscapes, because scale-related hydrography differences are known to vary in different physical conditions. While generalization algorithms have been tailored to specific regions and landscape conditions by several researchers in recent years, the selection and characterization of regional conditions have not been formally defined nor statistically validated. This paper undertakes a systematic classification of landscape types in the conterminous United States to spatially subset the country into workable units, in preparation for systematic tailoring of generalization workflows that preserve hydrographic characteristics. The classification is based upon elevation, standard deviation of elevation, slope, runoff, drainage and bedrock density, soil and bedrock permeability, area of inland surface water, infiltration-excess of overland flow, and a base flow index. A seven class solution shows low misclassification rates except in areas of high landscape diversity such as the Appalachians, Rocky Mountains, and Western coastal regions.

International Journal of Cartography

Incorporating water quality analysis into navigation assessments as demonstrated in the Mississippi River Basin

A description of historical and ambient water quality conditions is often required as part of navigational studies. This paper describes a series of tools developed by the USGS that can aid navigation managers in developing water quality assessments. The tools use R, a statistical software program, and provide methods to retrieve historical streamflow and water quality data, summarize observations, model concentrations and fluxes, and estimate seasonal, annual, and decadal trends. The utility of these tools is demonstrated by providing an analysis of the seasonal variability and long-term trends of nitrate plus nitrite, orthophosphate, and suspended sediment concentrations and fluxes at nine sites in the Mississippi River Basin. Trends in annual mean concentration and flux showed fairly stable nitrate plus nitrite at most of the nine sites, with increases in the Upper Mississippi and Missouri Rivers and decreases on the Illinois River over a 40-year period beginning in 1980. Orthophosphate concentration or flux increased at almost all sites over a similar time period. Conversely, a concurrent steady decline in suspended sediment concentrations and fluxes was noted at sites throughout the basin.

Mississippi River Basin