Search USGSSearch

Geology topics

Andrew Hoegh

Publications and source records attributed to Andrew Hoegh.

6 recordsLinked to original sources

Estimating disease prevalence from preferentially sampled, pooled data

After the onset of the COVID-19 pandemic, scientific interest in coronaviruses endemic in animal populations has increased dramatically. However, investigating the prevalence of disease in animal populations across the landscape, which requires finding and capturing animals can be difficult. Spatial random sampling over a grid could be extremely inefficient because animals can be hard to locate, and the total number of samples may be small. Alternatively, preferential sampling, using existing knowledge to inform sample location, can guarantee larger numbers of samples, but estimates derived from this sampling scheme may exhibit bias if there is a relationship between higher probability sampling locations and the disease prevalence. Sample specimens are commonly grouped and tested in pools which can also be an added challenge when combined with preferential sampling. Here we present a Bayesian method for estimating disease prevalence with preferential sampling in pooled presence-absence data motivated by estimating factors related to coronavirus infection among Mexican free-tailed bats ( Tadarida brasiliensis ) in California. We demonstrate the efficacy of our approach in a simulation study, where a naive model, not accounting for preferential sampling, returns biased estimates of parameter values; however, our model returns unbiased results regardless of the degree of preferential sampling. Our model framework is then applied to data from California to estimate factors related to coronavirus prevalence. After accounting for preferential sampling impacts, our model suggests small prevalence differences between male and female bats.

California

Bayesian model selection to investigate meaningful spatial scales

Ecologists and other statistical practitioners with access to high-resolution spatial data lack guidance on best approaches for discerning meaningful spatial scales for environmental covariates which is necessary when spatial factors influence environmental processes. Recently developed methods have attempted to automate investigating spatial scales for covariates by evaluating models for which potential explanatory variables are derived from concentric circles of increasing size centered at survey locations. However, these methods make a strong assumption on the inclusion of the covariate and do not help discern whether a covariate should be included in the model. We present an approach that utilizes researcher guidance to create informative priors on the model space that, along with parallelizable Reversible Jump MCMC techniques, enables efficient estimation of posterior model probabilities to assist with the choice of meaningful spatial scales for environmental covariates.

Authorea

Leveraging an observed-data likelihood improves the use of machine learning labels in a Bayesian hierarchical model for bioacoustic data

Classification of massive datasets by machine learning (ML) algorithms is promising for many scientific domains, especially wildlife monitoring programs that rely on passive acoustic surveys for detecting species. However, treating ML-predicted class labels (e.g., species identity) as truth biases inferences of focal parameters within common modeling frameworks. One solution is to model the misclassification process explicitly using human-validated true-class labels for a subset of observations. Validation by experts can present a substantial bottleneck in otherwise efficient workflows that use ML predictions. Bioacoustics practitioners seek guidance on both the quantity and process for selecting ML-labeled data to validate by an expert. We derive an alternative model formulation that jointly models human-validated and ML-predicted class labels with an observed-data likelihood (ODL) and use empirically informed simulations motivated by a real-data application to explore different probability designs for selecting class labels for validation. Simulation results suggest that with smaller validation sets the ODL formulation increases computational speed and reduces estimation error compared to a default MCMC data augmentation routine. Our methodology is transferable to applications that treat predictions from classification algorithms as the response variable of interest.

Annals of Applied Statistics

Clustering and unconstrained ordination with Dirichlet process mixture models

Assessment of similarity in species composition or abundance across sampled locations is a common goal in multi-species monitoring programs. Existing ordination techniques provide a framework for clustering sample locations based on species composition by projecting high-dimensional community data into a low-dimensional, latent ecological gradient representing species composition. However, these techniques require specification of the number of distinct ecological communities present in the latent space, which can be difficult to determine in advance. We develop an ordination model capable of simultaneous clustering and ordination that allows for estimation of the number of clusters present in the latent ecological gradient. This model draws latent coordinates for each sample location from a Dirichlet process mixture model, affording researchers with probabilistic statements about the number of clusters present in the latent ecological gradient. The model is compared to existing methods for simultaneous clustering and ordination via simulation and applied to two empirical datasets; JAGS code to fit the proposed model is provided in an appendix. The first dataset concerns presence-absence records of fish in the Doubs river in eastern France and the second dataset describes presence-absence records of plant species in Craters of the Moon National Monument and Preserve (CRMO) in Idaho, USA. Results from both analyses align with existing ecological gradients at each location. Development of the Dirichlet process ordination model provides wildlife managers with data-driven inferences about the number of distinct communities present across monitored locations, allowing for more cost-effective monitoring and reliable decision-making for conservation management.

Methods in Ecology and Evolution

An initial assessment of plankton tow detection probabilities for dreissenid mussels in the western United States

Early detection of dreissenid mussels ( Dreissena polymorpha and D. rostriformis bugensis ) is crucial to mitigating the economic and environmental impacts of an infestation. Plankton tow sampling is a common method used for early detection of dreissenid mussels, but little is known about the sampling intensity required for a high probability of early detection using the method. We used implicit dynamic occupancy models to estimate plankton tow detection probabilities of dreissenid mussels from a long-term data set containing plankton tow samples collected across central and western United States. We fit models using a) the entire data set, including water bodies with unknown occupancy status in addition to heavily infested water bodies, b) a data subset that included water bodies with paired water temperature data, and c) a data subset that included water bodies with lower dreissenid densities. For the entire data set, we found that estimated detection probabilities varied by water body size and ranged from approximately 0.10 to 0.86. For the water temperature subset, we observed the same pattern between detection probability and water body size as we did for the full data but additionally found that the estimated detection probabilities were much higher when water temperatures were above 12 °C. For the lower dreissenid density subset, we found that the estimated probability of detecting dreissenid mussels with a single aggregated plankton tow sample was near zero. Given these estimates, we conclude that the number of aggregated plankton tow samples taken per water body in the data is far fewer than the number needed to ensure a high probability of detecting dreissenid mussels, especially if they are at low densities. We summarize the analyses with a discussion of plankton tow sampling protocol changes needed to improve estimates of dreissenid detection probabilities.

western United States

Assessing spatial and temporal patterns in sagebrush steppe vegetation communities 2012-2018: Grand Teton National Park

Visual cover class data were collected on over 80 species across 30 permanent sampling frames in sagebrush steppe vegetation communities in Grand Teton National Park from 2012 to 2018. In this report, temporal and spatial patterns in species composition were assessed and used to inform potential sampling strategies for future monitoring. Specifically, the viability of a reduction in sampling effort was evaluated based on the similarity in species composition within each frame over time and among frames within each year. Using distance-based ordination techniques, we found little to no evidence of differences in species composition within each frame over time. Furthermore, there was little evidence of heterogeneity in species composition among frames within each year, though there was some evidence of differences in composition between the two principle sagebrush community types (sagebrush dry shrubland and sagebrush-bitterbrush) aggregated across frames. Based on these results, we propose that a reduction in sampling effort is viable and suggest a new monitoring schedule.

Wyoming