Search USGSSearch

SEARCH · Search USGS

Results for “Ecological Informatics”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Promoting synergy in the innovative use of environmental data—Workshop summary

From December 2 to 4, 2015, NatureServe and the U.S. Geological Survey organized and hosted a biodiversity and ecological informatics workshop at the U.S. Department of the Interior in Washington, D.C. The workshop objective was to identify user-driven future directions and areas of collaboration in advanced applications of environmental data applied to forecasting and decision making for the sustainability of biodiversity and ecosystem services. Substantial effort to recruit attendees from diverse Federal, State, and private sector organizations successfully attracted participants from 20 Federal agencies and 48 different institutions in the academic, nonprofit, State government, and commercial sectors; the total number of attendees ranged from 100 to 144 during the 3-day workshop. The first one-half of the workshop was divided into 7 plenary sessions and 3 sets of lightning talk sessions organized by sector, providing 48 oral and visual plenary presentations that shared diverse perspectives on biodiversity and ecological informatics, including original biospatial analyses from 6 graduate student map contest winners. The second one-half of the workshop focused on 10 breakout sessions with participant-driven themes from the environmental data sphere and concluded with an address by the Director of the U.S. Fish and Wildlife Service. The workshop was structured to encourage interactivity. About 80–90 percent of attendees provided direct feedback using clicker devices for specific questions related to biodiversity and ecological data uses and needs, and 10 breakout session leaders shared the highlights of their group discussions during the final workshop plenary sessions. Participants were encouraged to use the Twitter hashtag #ShareUrData. Over lunch on day 2 there were 20 simultaneous presentations of tools and apps during a special “Tools Café” session. The 10 participant-defined breakout session topics are listed below: Ecosystem services and ecological indicators Inventory and monitoring Biogeographic map of the Nation Pollinators Invasive species Remote sensing Drivers of agricultural change Citizen science Climate Hydrology and watersheds Numerous common themes that emerged from the workshop include the following: The vital importance of completing foundational environmental datasets that are nationally consistent and are essential to multiple sectors, such as the Soil Survey Geographic database high-resolution soils data, a minimum 5-meter resolution digital elevation model, national hydrographic data, high-resolution land cover data, time series high-resolution spatial climate data from historical to future time steps, and a national wetland inventory. Improved, nationally consistent environmental datasets (integrated with targeted observations) will dramatically advance forecasting capacity and support early warning systems (that is, drought, forest disease); however, multiagency coordination should focus on decision support tools that convey appropriate actions and responses to adapt to, and mitigate, potential negative consequences. Digitizing and providing access to the vast stores of underused historical data that can be leveraged for this purpose is of national importance. Modern computational techniques and the ever-increasing flow of environmental data from ground and remote observations can support improved understanding of environmental change. Success of understanding patterns of change for decision making requires establishing baselines from which change can be measured. The value of digitized historical data is greater than ever before. There is a need to recognize the multifaceted potential of citizen science to engage the public in resource stewardship, to create the next generation of science, technology, engineering, math, and environmental leaders, and to have sufficient field personnel to monitor environmental trends, including early detection of alien invasive species, phenological shifts, shifting distribution and abundance of indicator species, and species inventories. The Federal government has an essential role in creating the infrastructure to dramatically improve mobilization of citizen science (and other) data by fostering the following: creation of data standards, creation of nationally consistent framework datasets, vertical integration of observation data, visualization and dissemination of aggregated datasets, and calculation and communication of derived trends. Current and near future trends in the availability of remotely sensed data (rapid expansion of satellite fleets and drones) is revolutionizing access to near-real-time ecological data. Targeted integration with ground-based observations and instrumentation has an extremely valuable role in validating remotely sensed data, filling data gaps, improving data quality, and fully realizing the potential of the near-real-time monitoring of environmental indicator trends. Integrated management of environmental data at the landscape scale is required even as specific actions on the ground are largely local in nature. The workshop highlighted numerous success stories; however, almost every breakout group pointed out the still-too-fragmented nature of the current data landscape. Management and delivery of the necessary data, tools, and analyses to sustain our Nation’s environmental capital must be a collaborative effort between Federal, State, and local governments, academia, nonprofits, and the commercial sector, even though the responsibilities of each sector are different.

Open-File Report

Synthesizing and analyzing long-term monitoring data: A greater sage-grouse case study

Long-term monitoring of natural resources is imperative for increasing the understanding of ecosystem processes, services, and how to manage those ecosystems to maintain or improve function. Challenges with using these data may occur because methods of monitoring changed over time, multiple organizations collect and manage data differently, and monetary resources fluctuate, affecting many aspects of data. Because many species respond to changes in habitat conditions and predator-prey relationships across different spatial scales that span management boundaries, greater efforts for collaborating are essential. We demonstrate the challenges and methods for standardizing greater sage-grouse ( Centrocercus urophasianus ) long-term monitoring data across the species range in the western United States to inform population modeling needs identified by the Western Association of Fish and Wildlife Agencies. We used automated and repeatable methods of standardizing data via custom open-source software ( grsg_lekdb ) to improve the scientific integrity of future sage-grouse population assessments within and among states. Data standardization included reconciling uses of different terminology and expunging unusable data, resulting in the removal of 26% of data records due to database insertion errors and modifications to >1 million values to correct formatting and typing errors. Our approaches maximized the inclusion of usable data and identified data that could inform detection probabilities, population trends, and monitoring guidelines. Using sage-grouse databases as an example, we identified the importance of data management and how quality assurance and quality control measures can improve the usefulness of these data for future research needs. Our methods of using informatics and concluding recommendations can support similar endeavors of flora and fauna monitoring programs, whether those efforts are to use existing data or support new monitoring programs.

Ecological Informatics

Development of a generic auto-calibration package for regional ecological modeling and application in the Central Plains of the United States

Process-oriented ecological models are frequently used for predicting potential impacts of global changes such as climate and land-cover changes, which can be useful for policy making. It is critical but challenging to automatically derive optimal parameter values at different scales, especially at regional scale, and validate the model performance. In this study, we developed an automatic calibration (auto-calibration) function for a well-established biogeochemical model—the General Ensemble Biogeochemical Modeling System (GEMS)-Erosion Deposition Carbon Model (EDCM)—using data assimilation technique: the Shuffled Complex Evolution algorithm and a model-inversion R package—Flexible Modeling Environment (FME). The new functionality can support multi-parameter and multi-objective auto-calibration of EDCM at the both pixel and regional levels. We also developed a post-processing procedure for GEMS to provide options to save the pixel-based or aggregated county-land cover specific parameter values for subsequent simulations. In our case study, we successfully applied the updated model (EDCM-Auto) for a single crop pixel with a corn–wheat rotation and a large ecological region (Level II)—Central USA Plains. The evaluation results indicate that EDCM-Auto is applicable at multiple scales and is capable to handle land cover changes (e.g., crop rotations). The model also performs well in capturing the spatial pattern of grain yield production for crops and net primary production (NPP) for other ecosystems across the region, which is a good example for implementing calibration and validation of ecological models with readily available survey data (grain yield) and remote sensing data (NPP) at regional and national levels. The developed platform for auto-calibration can be readily expanded to incorporate other model inversion algorithms and potential R packages, and also be applied to other ecological models.

Ecological Informatics

A suggestion for computing objective function in model calibration

A parameter-optimization process (model calibration) is usually required for numerical model applications, which involves the use of an objective function to determine the model cost (model-data errors). The sum of square errors (SSR) has been widely adopted as the objective function in various optimization procedures. However, ‘square error’ calculation was found to be more sensitive to extreme or high values. Thus, we proposed that the sum of absolute errors (SAR) may be a better option than SSR for model calibration. To test this hypothesis, we used two case studies—a hydrological model calibration and a biogeochemical model calibration—to investigate the behavior of a group of potential objective functions: SSR, SAR, sum of squared relative deviation (SSRD), and sum of absolute relative deviation (SARD). Mathematical evaluation of model performance demonstrates that ‘absolute error’ (SAR and SARD) are superior to ‘square error’ (SSR and SSRD) in calculating objective function for model calibration, and SAR behaved the best (with the least error and highest efficiency). This study suggests that SSR might be overly used in real applications, and SAR may be a reasonable choice in common optimization implementations without emphasizing either high or low values (e.g., modeling for supporting resources management).

Ecological Informatics

Organization of marine phenology data in support of planning and conservation in ocean and coastal ecosystems

Among the many effects of climate change is its influence on the phenology of biota. In marine and coastal ecosystems, phenological shifts have been documented for multiple life forms; however, biological data related to marine species' phenology remain difficult to access and is under-used. We conducted an assessment of potential sources of biological data for marine species and their availability for use in phenological analyses and assessments. Our evaluations showed that data potentially related to understanding marine species' phenology are available through online resources of governmental, academic, and non-governmental organizations, but appropriate datasets are often difficult to discover and access, presenting opportunities for scientific infrastructure improvement. The developing Federal Marine Data Architecture when fully implemented will improve data flow and standardization for marine data within major federal repositories and provide an archival repository for collaborating academic and public data contributors. Another opportunity, largely untapped, is the engagement of citizen scientists in standardized collection of marine phenology data and contribution of these data to established data flows. Use of metadata with marine phenology related keywords could improve discovery and access to appropriate datasets. When data originators choose to self-publish, publication of research datasets with a digital object identifier, linked to metadata, will also improve subsequent discovery and access. Phenological changes in the marine environment will affect human economics, food systems, and recreation. No one source of data will be sufficient to understand these changes. The collective attention of marine data collectors is needed—whether with an agency, an educational institution, or a citizen scientist group—toward adopting the data management processes and standards needed to ensure availability of sufficient and useable marine data to understand marine phenology.

Ecological Informatics

Caveats for correlative species distribution modeling

Correlative species distribution models are becoming commonplace in the scientific literature and public outreach products, displaying locations, abundance, or suitable environmental conditions for harmful invasive species, threatened and endangered species, or species of special concern. Accurate species distribution models are useful for efficient and adaptive management and conservation, research, and ecological forecasting. Yet, these models are often presented without fully examining or explaining the caveats for their proper use and interpretation and are often implemented without understanding the limitations and assumptions of the model being used. We describe common pitfalls, assumptions, and caveats of correlative species distribution models to help novice users and end users better interpret these models. Four primary caveats corresponding to different phases of the modeling process, each with supporting documentation and examples, include: (1) all sampling data are incomplete and potentially biased; (2) predictor variables must capture distribution constraints; (3) no single model works best for all species, in all areas, at all spatial scales, and over time; and (4) the results of species distribution models should be treated like a hypothesis to be tested and validated with additional sampling and modeling in an iterative process.

Ecological Informatics

Simulating realistic predator signatures in quantitative fatty acid signature analysis

Diet estimation is an important field within quantitative ecology, providing critical insights into many aspects of ecology and community dynamics. Quantitative fatty acid signature analysis (QFASA) is a prominent method of diet estimation, particularly for marine mammal and bird species. Investigators using QFASA commonly use computer simulation to evaluate statistical characteristics of diet estimators for the populations they study. Similar computer simulations have been used to explore and compare the performance of different variations of the original QFASA diet estimator. In both cases, computer simulations involve bootstrap sampling prey signature data to construct pseudo-predator signatures with known properties. However, bootstrap sample sizes have been selected arbitrarily and pseudo-predator signatures therefore may not have realistic properties. I develop an algorithm to objectively establish bootstrap sample sizes that generates pseudo-predator signatures with realistic properties, thereby enhancing the utility of computer simulation for assessing QFASA estimator performance. The algorithm also appears to be computationally efficient, resulting in bootstrap sample sizes that are smaller than those commonly used. I illustrate the algorithm with an example using data from Chukchi Sea polar bears ( Ursus maritimus ) and their marine mammal prey. The concepts underlying the approach may have value in other areas of quantitative ecology in which bootstrap samples are post-processed prior to their use.

Ecological Informatics

Spatial prediction of wheat Septoria leaf blotch (Septoria tritici) disease severity in central Ethiopia

A number of studies have reported the presence of wheat septoria leaf blotch ( Septoria tritici ; SLB) disease in Ethiopia. However, the environmental factors associated with SLB disease, and areas under risk of SLB disease, have not been studied. Here, we tested the hypothesis that environmental variables can adequately explain observed SLB disease severity levels in West Shewa, Central Ethiopia. Specifically, we identified 50 environmental variables and assessed their relationships with SLB disease severity. Geographically referenced disease severity data were obtained from the field, and linear regression and Boosted Regression Trees (BRT) modeling approaches were used for developing spatial models. Moderate-resolution imaging spectroradiometer (MODIS) derived vegetation indices and land surface temperature (LST) variables highly influenced SLB model predictions. Soil and topographic variables did not sufficiently explain observed SLB disease severity variation in this study. Our results show that wheat growing areas in Central Ethiopia, including highly productive districts, are at risk of SLB disease. The study demonstrates the integration of field data with modeling approaches such as BRT for predicting the spatial patterns of severity of a pathogenic wheat disease in Central Ethiopia. Our results can aid Ethiopia's wheat disease monitoring efforts, while our methods can be replicated for testing related hypotheses elsewhere.

Ecological Informatics

AnimalFinder: A semi-automated system for animal detection in time-lapse camera trap images

Although the use of camera traps in wildlife management is well established, technologies to automate image processing have been much slower in development, despite their potential to drastically reduce personnel time and cost required to review photos. We developed AnimalFinder in MATLAB® to identify animal presence in time-lapse camera trap images by comparing individual photos to all images contained within the subset of images (i.e. photos from the same survey and site), with some manual processing required to remove false positives and collect other relevant data (species, sex, etc.). We tested AnimalFinder on a set of camera trap images and compared the presence/absence results with manual-only review with white-tailed deer ( Odocoileus virginianus ), wild pigs ( Sus scrofa ), and raccoons ( Procyon lotor ). We compared abundance estimates, model rankings, and coefficient estimates of detection and abundance for white-tailed deer using N-mixture models. AnimalFinder performance varied depending on a threshold value that affects program sensitivity to frequently occurring pixels in a series of images. Higher threshold values led to fewer false negatives (missed deer images) but increased manual processing time, but even at the highest threshold value, the program reduced the images requiring manual review by ~ 40% and correctly identified > 90% of deer, raccoon, and wild pig images. Estimates of white-tailed deer were similar between AnimalFinder and the manual-only method (~ 1–2 deer difference, depending on the model), as were model rankings and coefficient estimates. Our results show that the program significantly reduced data processing time and may increase efficiency of camera trapping surveys.

Ecological Informatics

Small values in big data: The continuing need for appropriate metadata

Compiling data from disparate sources to address pressing ecological issues is increasingly common. Many ecological datasets contain left-censored data – observations below an analytical detection limit. Studies from single and typically small datasets show that common approaches for handling censored data — e.g., deletion or substituting fixed values — result in systematic biases. However, no studies have explored the degree to which the documentation and presence of censored data influence outcomes from large, multi-sourced datasets. We describe left-censored data in a lake water quality database assembled from 74 sources and illustrate the challenges of dealing with small values in big data, including detection limits that are absent, range widely, and show trends over time. We show that substitutions of censored data can also bias analyses using ‘big data’ datasets, that censored data can be effectively handled with modern quantitative approaches, but that such approaches rely on accurate metadata that describe treatment of censored data from each source.

Ecological Informatics

Evaluation of biodiversity data portals based on requirement analysis

In recent years, concern about the misuse of natural resources has been increasing. It is essential to know in detail the biodiversity of an ecosystem to understand and analyze the impact of human activities on nature, as well as to promote the economic growth of a country. To achieve these goals, public and private institutions are aggregating and sharing biological data around the world by means of biodiversity data portals. The main purpose of those portals is to provide a set of tools that help users and institutions catalog, analyze, and publish raw data about different species in a manner that is open and freely available to any interested party. Normally the process of choosing the best software solution is not straightforward. This paper proposes a methodology to evaluate a collection of data portals to establish a clear and consistent selection process that analyzes a collection of requirements and research purposes. The proposed approach is based on three strategies: the use of software engineering techniques to identify the desired group of features to be available in the data portal; the application of the Kano Satisfaction Model to score each requirement according to a preset weight of importance; and the use of tree-maps to visualize the requirements based on their implementation priority, to establish a portal deployment road-map. The proposed methodology is broadly applicable to portal analyses for many communities of practice.

Ecological Informatics

A reporting format for leaf-level gas exchange data and metadata

Leaf-level gas exchange data support the mechanistic understanding of plant fluxes of carbon and water. These fluxes inform our understanding of ecosystem function, are an important constraint on parameterization of terrestrial biosphere models, are necessary to understand the response of plants to global environmental change, and are integral to efforts to improve crop production. Collection of these data using gas analyzers can be both technically challenging and time consuming, and individual studies generally focus on a small range of species, restricted time periods, or limited geographic regions. The high value of these data is exemplified by the many publications that reuse and synthesize gas exchange data, however the lack of metadata and data reporting conventions make full and efficient use of these data difficult. Here we propose a reporting format for leaf-level gas exchange data and metadata to provide guidance to data contributors on how to store data in repositories to maximize their discoverability, facilitate their efficient reuse, and add value to individual datasets. For data users, the reporting format will better allow data repositories to optimize data search and extraction, and more readily integrate similar data into harmonized synthesis products. The reporting format specifies data table variable naming and unit conventions, as well as metadata characterizing experimental conditions and protocols. For common data types that were the focus of this initial version of the reporting format, i.e., survey measurements, dark respiration, carbon dioxide and light response curves, and parameters derived from those measurements, we took a further step of defining required additional data and metadata that would maximize the potential reuse of those data types. To aid data contributors and the development of data ingest tools by data repositories we provided a translation table comparing the outputs of common gas exchange instruments. Extensive consultation with data collectors, data users, instrument manufacturers, and data scientists was undertaken in order to ensure that the reporting format met community needs. The reporting format presented here is intended to form a foundation for future development that will incorporate additional data types and variables as gas exchange systems and measurement approaches advance in the future. The reporting format is published in the U.S. Department of Energy's ESS-DIVE data repository, with documentation and future development efforts being maintained in a version control system.

Ecological Informatics

Accessibility of environmental data for sharing: The role of UX in large cyberinfrastructure projects

Incorporating user experience (UX) testing when creating research cyberinfrastructure is often overlooked, but if left too late, the cost of retrofitting is considerable, and the very clients the cyberinfrastructure was built to serve may be lost. Successfully integrating UX testing into the product development cycle can be difficult but rewarding. This paper describes how UX evaluations were incorporated over ten years of operation of DataONE ( www.dataone.org ), a multi-sector science research cyberinfrastructure project created to support the discovery, access, and sustainability of data about life on Earth and the environment that sustains it. The diverse stakeholders in DataONE include data creators and users such as researchers and government workers across the broad scope of the earth and environmental sciences as well as those who hold and manage data such as libraries and data repositories. Between 2009 and 2019 DataONE members designed and constructed data management tools and services to fulfill the DataONE objectives. To assist in achieving its goals, a participatory design approach was used by establishing several largely volunteer and stakeholder-representative working groups, including the Usability and Assessment Working Group. This Working Group conducted over forty UX evaluations to assess the usability of DataONE products and websites at various stages of the development process. In addition to improving the usability of DataONE products, the UX evaluations fostered community involvement by building trust and engagement with the products being developed. The DataONE UX experience yields several important lessons which will improve the success of other projects. It is our conclusion that UX testing should be a mandatory part of the design of any cyberinfrastructure project.

Ecological Informatics

PS3: The Pheno-Synthesis software suite for integration and analysis of multi-scale, multi-platform phenological data

Phenology is the study of recurring plant and animal life-cycle stages which can be observed across spatial and temporal scales that span orders of magnitude (e.g., organisms to landscapes). The variety of scales at which phenological processes operate is reflected in the range of methods for collecting phenologically relevant data, and the programs focused on these collections. Consideration of the scale at which phenological observations are made, and the platform used for observation, is critical for the interpretation of phenological data and the application of these data to both research questions and land management objectives. However, there is currently little capacity to facilitate access, integration and analysis of cross-scale, multi-platform phenological data. This paper reports on a new suite of software and analysis tools – the “Pheno-Synthesis Software Suite,” or PS3 – to facilitate integration and analysis of phenological and ancillary data, enabling investigation and interpretation of phenological processes at scales ranging from organisms to landscapes and from days to decades. We use PS3 to investigate phenological processes in a semi-aride, mixed shrub-grass ecosystem, and find that the apparent importance of seasonal precipitation to vegetation activity (i.e., “greenness”) is affected by the scale and platform of observation. We end by describing potential applications of PS3 to phenological modeling and forecasting, understanding patterns and drivers of phenological activity in real-world ecosystems, and supporting agricultural and natural resource management and decision-making.

Ecological Informatics

Characterizing mauka-to-makai connections for aquatic ecosystem conservation on Maui, Hawaiʻi

Mauka-to-makai (mountain to sea in the Hawaiian language) hydrologic connectivity – commonly referred to as ridge-to-reef – directly affects biogeochemical processes and socioecological functions across terrestrial, freshwater, and marine systems. The supply of freshwater to estuarine and nearshore environments in a ridge-to-reef system supports the food, water, and habitats utilized by marine fauna . In addition, the ecosystem services derived from this land-to-sea connectivity support social and cultural practices (hereafter referred to as socio-cultural) including fishing, aquaculture, wetland agriculture, religious ceremonies, and recreational activities. To effectively guide island resource management, a better understanding of the linkages from ridge-to-reef across natural and social usages is critical, particularly in the context of climate change, with anticipated increasing temperature and shifting precipitation patterns. The objective of this study was to identify spatial linkages that promote multiple and diverse uses, following the ridge-to-reef concept, at an island-wide scale to identify regions of high conservation importance for aquatic resources. We selected the Island of Maui as a study representative of many Pacific islands. Diverse datasets, including agricultural lands within watersheds, wetland locations, presence of stream species, indicators of freshwater input from streams, coral cover, nearshore fish biomass, socio-cultural data such as fishpond locations, wetland taro cultivation, beach recreation use, and lastly the dynamically downscaled Coupled Model Intercomparison Project Phase (CMIP5) future climate projections scenarios (Representative Concentration Pathway (RCP) 4.5 & 8.5) were used to examine the spatial linkages through hydrological connectivity from land to the sea. Zonation spatial planning software was used to prioritize areas of high management and conservation value and to help inform aquatic resources management. The resulting prioritized areas included many minimally disturbed watersheds in east Maui and western nearshore and coastal zones that are adjacent to diverse coral reefs. These results are driven by the importance of fish biomass and coral reef distribution as well as traditional wetland taro cultivation and coastal access points for recreation. These results underline the importance of examining ridge-to-reef systems for aquatic resource management and including important social and cultural values in resource management upon planning adaptation strategies for climate change. Improving our understanding of diverse natural and socio-cultural influences on habitat conditions and their values in these areas provides an opportunity to strategically plan future management and conservation actions.

Hawaii

Invaders at the doorstep: Using species distribution modeling to enhance invasive plant watch lists

Watch lists of invasive species that threaten a particular land management unit are useful tools because they can draw attention to invasive species at the very early stages of invasion when early detection and rapid response efforts are often most successful. However, watch lists typically rely on the subjective selection of invasive species by experts or on the use of spotty occurrence records. Further, incomplete records of invasive plant occurrences bias these watch lists towards the inclusion of invasive plant species that may already be present in a land management unit, because the occurrences have not been formally integrated into publicly accessible biodiversity databases. However, these problems may be overcome by an iterative approach that guides more complete detection and compilation of invasive plant species records within land management units. To address issues from unobserved or unrecorded occurrences, we combined predicted suitable habitat from species distribution models and aggregated invasive plant occurrence records to develop ranked watch lists of 146 priority invasive plant species on >4000 land management units from five different administrative types within the United States. Based on this analysis, we determined that on average 84% of priority invasive plants with suitable habitat within a given land management unit were as yet unobserved, and that 41% of those were ‘doorstep species’ – found within 50 miles of the unit boundary yet not detected within the unit. Two case studies, developed in collaboration with staff at U.S. Fish and Wildlife Service Refuges, showed that by combining both habitat suitability models and invasive plant occurrence records, we could identify additional problematic invasive plants that had been previously overlooked. Model-based watch lists of ‘doorstep species’ are useful tools because they can objectively alert land managers to threats from invasive plants with high likelihood of establishment.

contiguous United States

Evaluating a tandem human-machine approach to labelling of wildlife in remote camera monitoring

Remote cameras (“trail cameras”) are a popular tool for non-invasive, continuous wildlife monitoring, and as they become more prevalent in wildlife research, machine learning (ML) is increasingly used to automate or accelerate the labor-intensive process of labelling (i.e., tagging) photos. Human-machine hybrid tagging approaches have been shown to greatly increase tagging efficiency (i.e., time to tag a single image). However, those potential increases hinge on the extent to which an ML model makes correct vs. incorrect predictions. We performed an experiment using a ML model that produces bounding boxes around animals, people, and vehicles in remote camera imagery (MegaDetector) to consider the impact of a ML model’s performance on its ability to accelerate human labeling. Six participants tagged trail camera images collected from 12 sites in Vermont and Maine, USA (January–September 2022) using three tagging methods (one with ML bounding box assistance and two without assistance). We used a generalized linear mixed model to examine the influence of ML model performance and tagging method on tagging efficiency. We found that ML bounding boxes offer significant improvement in tagging efficiency when labelling data compared to unassisted tagging. Additionally, the time taken to label with bounding boxes was not statistically different from an unassisted tagging approach. However, we found that gains in efficiency are contingent on the ML algorithm’s performance and that incorrect ML predictions, particularly the 4.2% false positive and 3.6% false negative predictions, can slow the tagging process compared to a non-hybrid approach. These findings indicate that although practitioners usually forgo the production of bounding boxes when selecting a data labelling process due to the increased effort, ML bounding box-assisted tagging can offer an efficient method for labeling. More broadly, ML-assisted data labelling offers an opportunity to accelerate the analysis of trail camera imagery, but an assessment of the ML model’s performance can illuminate whether the hybrid-tagging approach is ultimately a help or hinderance.

Maine, Vermont

Understanding gaps in early detection of and rapid response to invasive species in the United States: A literature review and bibliometric analysis

While concepts regarding invasive species establishment patterns and eradication possibilities have long been a topic of invasion biology, the specific terminology referring to early detection of and rapid response to (EDRR) invasive species emerged in scientific literature during the early 2000s. Since then, the EDRR approach has expanded to include a suite of detection, planning, and management tools. By conducting a systematic literature review, we attempt to characterize the field of EDRR in the United States and its territories as reflected by publication records. Specifically, we assessed publication data such as the number of publications per year, the most common journals where papers were published, and the relationship between author’s keywords for studies focusing on aquatic and terrestrial habitats. For publications that used invasive species occurrence or abundance data (whether collected for the purposes of the respective publication or acquired from another data source), we manually vetted additional information such as focal taxa, data collection years and locations, sources of other data used, and whether data or code were deposited in open access formats. We also conducted network analyses for the author institutions that coauthored papers together most frequently and for the references most cited by EDRR publications. Overall, we found that silos existed in terms of which author institutions worked together, which existing literature was cited, and which topics were frequently explored. We also found evidence of substantial gaps in data access and use. For example, although a wide variety of data sources for invasive species occurrences are available, these sources were seldom cited by published literature, and newly collected data was not often deposited into invasive species databases or other open-source data repositories. Considering the continued advocation for a centralized national EDRR information system, our study suggests that facilitating access to data, decision support tools, and other informational resources represents a key opportunity for improving EDRR capabilities.

Ecological Informatics