Search USGSSearch

USGS · 70250311

A practical guide to understanding and validating complex models using data simulations

Abstract

Biologists routinely fit novel and complex statistical models to push the limits of our understanding. Examples include, but are not limited to, flexible Bayesian approaches (e.g. BUGS, stan), frequentist and likelihood-based approaches (e.g. packages lme4 ) and machine learning methods. These software and programs afford the user greater control and flexibility in tailoring complex hierarchical models. However, this level of control and flexibility places a higher degree of responsibility on the user to evaluate the robustness of their statistical inference. To determine how often biologists are running model diagnostics on hierarchical models, we reviewed 50 recently published papers in 2021 in the journal Nature Ecology & Evolution , and we found that the majority of published papers did not report any validation of their hierarchical models, making it difficult for the reader to assess the robustness of their inference. This lack of reporting likely stems from a lack of standardized guidance for best practices and standard methods. Here, we provide a guide to understanding and validating complex models using data simulations. To determine how often biologists use data simulation techniques, we also reviewed 50 recently published papers in 2021 in the journal Methods Ecology & Evolution . We found that 78% of the papers that proposed a new estimation technique, package or model used simulations or generated data in some capacity (18 of 23 papers); but very few of those papers (5 of 23 papers) included either a demonstration that the code could recover realistic estimates for a dataset with known parameters or a demonstration of the statistical properties of the approach. To distil the variety of simulations techniques and their uses, we provide a taxonomy of simulation studies based on the intended inference. We also encourage authors to include a basic validation study whenever novel statistical models are used, which in general, is easy to implement. Simulating data helps a researcher gain a deeper understanding of the models and their assumptions and establish the reliability of their estimation approaches. Wider adoption of data simulations by biologists can improve statistical inference, reliability and open science practices.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Graziella Vittoria DiRenzo, Ephraim Hanks, David A. W. Miller. 2022-11-18. A practical guide to understanding and validating complex models using data simulations. https://doi.org/10.1111/2041-210x.14030

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Simulated soundscapes and transfer learning boost the performance of acoustic classifiers under data scarcity

1. The biodiversity crisis necessitates spatially extensive methods to monitor multiple taxonomic groups for evidence of change in response to evolving environmental conditions. Programs that combine passive acoustic monitoring and machine learning are increasingly used to meet this need. These methods require large, annotated datasets, which are time-consuming and expensive to produce, creating potential barriers to adoption in data- and funding-poor regions. Recently released pre-trained avian acoustic classification models provide opportunities to reduce the need for manual labelling and accelerate the development of new acoustic classification algorithms through transfer learning. Transfer learning is a strategy for developing algorithms under data scarcity that uses pre-trained models from related tasks to adapt to new tasks. 2. Our primary objective was to develop a transfer learning strategy using the feature embeddings of a pre-trained avian classification model to train custom acoustic classification models in data-scarce contexts. We used three annotated avian acoustic datasets to test whether transfer learning and soundscape simulation-based data augmentation could substantially reduce the annotated training data necessary to develop performant custom acoustic classifiers. We also conducted a sensitivity analysis for hyperparameter choice and model architecture. We then assessed the generalizability of our strategy to increasingly novel non-avian classification tasks. 3. With as few as two training examples per class, our soundscape simulation data augmentation approach consistently yielded new classifiers with improved performance relative to the pre-trained classification model and transfer learning classifiers trained with other augmentation approaches. Performance increases were evident for three avian test datasets, including single-class and multi-label contexts. We observed that the relative performance among our data augmentation approaches varied for the avian datasets and nearly converged for one dataset when we included more training examples. 4. We demonstrate an efficient approach to developing new acoustic classifiers leveraging open-source sound repositories and pre-trained networks to reduce manual labelling. With very few examples, our soundscape simulation approach to data augmentation yielded classifiers with performance equivalent to those trained with many more examples, showing it is possible to reduce manual label-ling while still achieving high-performance classifiers and, in turn, expanding the potential for passive acoustic monitoring to address rising biodiversity monitoring needs.

Methods in Ecology and Evolution

Spatial close-kin mark-recapture models applied to terrestrial species with continuous natal dispersal

Close-kin mark–recapture (CKMR) methods use information on genetic relatedness among individuals to estimate demographic parameters. An individual's genotype can be considered a ‘recapture’ of each of its parent's genotype, and the frequency of kin-pair matches detected in a population sample can directly inform estimates of abundance. CKMR inference procedures require analysts to define kinship probabilities in functional forms, which inevitably involve simplifying assumptions. Among others, population structure can have a strong influence on how kinship probabilities are formulated. Many terrestrial species are philopatric or face barriers to dispersal, and not accounting for dispersal limitation in kinship probabilities, can create substantial bias if sampling is also spatially structured (e.g. via harvest). We present a spatially explicit formulation of CKMR that corrects for incomplete mixing by incorporating natal dispersal distances and spatial distribution of individuals into the kinship probabilities. We used individual-based simulations to evaluate the accuracy of abundance estimates obtained with one spatially naïve and two spatially explicit CKMR models across six scenarios with distinct spatial patterns of relative abundance and sampling probability. Estimates of abundance obtained with a CKMR model naïve to spatial structure were negatively biased when sampling was spatially biased. Incorporating patterns of natal dispersal in the kinship probabilities helped address this bias, but estimates were not always accurate depending on the model used and the scenario considered. Incorporating natal dispersal into spatially structured CKMR models can address the bias created by population structure and heterogeneous sampling but will often require additional assumptions and auxiliary data (e.g. relative abundance indices). The models shown here were designed for terrestrial species with continuous patterns of natal dispersal and high year-to-year site fidelity but could be extended to other species.

Methods in Ecology and Evolution

Sampling mass mortality events to enable diagnoses: A protocol using freshwater mussels

Many taxa around the globe are threatened by often unexplained mass mortality events (MMEs), which can decimate populations and compromise key ecosystem functions. One example of a highly threatened taxon facing frequent MMEs is freshwater mussels (Unionida). There has been a recent increase in interest in understanding the causes of freshwater mussel MMEs, but standardised methodologies for how best to respond to them to facilitate diagnoses are unavailable. When an MME is observed, swift and appropriate sample collection is imperative owing to the transient nature of these phenomena. Here we provide structured guidance that will facilitate rapid and appropriate sampling of MMEs, using freshwater mussels as an example. We set out standardised procedures for sample collection, preparation and preservation. The procedures we outline will improve our capacity for diagnostic investigations of MMEs and other mortality events, not only in freshwater mussels but also across many other taxa. This, in turn, can inform appropriate management responses.

Methods in Ecology and Evolution