Search USGSSearch

SEARCH · Search USGS

Results for “Data Science Journal”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The challenge of archiving and preserving remotely sensed data

Few would question the need to archive the scientific and technical (S&T) data generated by researchers. At a minimum, the data are needed for change analysis. Likewise, most people would value efforts to ensure the preservation of the archived S&T data. Future generations will use analysis techniques not even considered today. Until recently, archiving and preserving these data were usually accomplished within existing infrastructures and budgets. As the volume of archived data increases, however, organizations charged with archiving S&T data will be increasingly challenged (U.S. General Accounting Office, 2002). The U.S. Geological Survey has had experience in this area and has developed strategies to deal with the mountain of land remote sensing data currently being managed and the tidal wave of expected new data. The Agency has dealt with archiving issues, such as selection criteria, purging, advisory panels, and data access, and has met with preservation challenges involving photographic and digital media. That experience has allowed the USGS to develop management approaches, which this paper outlines.

Data Science Journal

The landsat image mosaic of the Antarctica Web Portal

People believe what they can see. The Poles exist as a frozen dream to most people. The International Polar Year wants to break the ice (so to speak), open up the Poles to the general public, support current polar research, and encourage new research projects. The IPY officially begins in March, 2007. As part of this effort, the U.S. Geological Survey (USGS) and the British Antarctic Survey (BAS), with funding from the National Science Foundation (NSF), are developing three Landsat mosaics of Antarctica and an Antarctic Web Portal with a Community site and an online map viewer. When scientists are able to view the entire scope of polar research, they will be better able to collaborate and locate the resources they need. When the general public more readily sees what is happening in the polar environments, they will understand how changes to the polar areas affect everyone.

Data Science Journal

Development of a standard database of reference sites for validating global burned area products

Over the past 2 decades, several global burned area products have been produced and released to the public. However, the accuracy assessment of such products largely depends on the availability of reliable reference data that currently do not exist on a global scale or whose production require a high level of dedication of project resources. The important lack of reference data for the validation of burned area products is addressed in this paper. We provide the Burned Area Reference Database (BARD), the first publicly available database created by compiling existing reference BA (burned area) datasets from different international projects. BARD contains a total of 2661 reference files derived from Landsat and Sentinel-2 imagery. All those files have been checked for internal quality and are freely provided by the authors. To ensure database consistency, all files were transformed to a common format and were properly documented by following metadata standards. The goal of generating this database was to give BA algorithm developers and product testers reference information that would help them to develop or validate new BA products. BARD is freely available at https://doi.org/10.21950/BBQQU7 (Franquesa et al., 2020).

Earth System Science Data journal

Post-disaster supply chain interdependent critical infrastructure system restoration: A review of data necessary and available for modeling

The majority of restoration strategies in the wake of large-scale disasters have focused on short-term emergency response solutions. Few consider medium- to long-term restoration strategies to reconnect urban areas to national supply chain interdependent critical infrastructure systems (SCICI). These SCICI promote the effective flow of goods, services, and information vital to the economic vitality of an urban environment. To re-establish the connectivity that has been broken during a disaster between the different SCICI, relationships between these systems must be identified, formulated, and added to a common framework to form a system-level restoration plan. To accomplish this goal, a considerable collection of SCICI data is necessary. The aim of this paper is to review what data are required for model construction, the accessibility of these data, and their integration with each other. While a review of publicly available data reveals a dearth of real-time data to assist modeling long-term recovery following an extreme event, a significant amount of static data does exist and these data can be used to model the complex interdependencies needed. For the sake of illustration, a particular SCICI (transportation) is used to highlight the challenges of determining the interdependencies and creating models capable of describing the complexity of an urban environment with the data publicly available. Integration of such data as is derived from public domain sources is readily achieved in a geospatial environment, after all geospatial infrastructure data are the most abundant data source and while significant quantities of data can be acquired through public sources, a significant effort is still required to gather, develop, and integrate these data from multiple sources to build a complete model. Therefore, while continued availability of high quality, public information is essential for modeling efforts in academic as well as government communities, a more streamlined approach to a real-time acquisition and integration of these data is essential.

Data Science Journal

Developing criteria to establish Trusted Digital Repositories

This paper details the drivers, methods, and outcomes of the U.S. Geological Survey’s quest to establish criteria by which to judge its own digital preservation resources as Trusted Digital Repositories. Drivers included recent U.S. legislation focused on data and asset management conducted by federal agencies spending $100M USD or more annually on research activities. The methods entailed seeking existing evaluation criteria from national and international organizations such as International Standards Organization (ISO), U.S. Library of Congress, and Data Seal of Approval upon which to model USGS repository evaluations. Certification, complexity, cost, and usability of existing evaluation models were key considerations. The selected evaluation method was derived to allow the repository evaluation process to be transparent, understandable, and defensible; factors that are critical for judging competing, internal units. Implementing the chosen evaluation criteria involved establishing a cross-agency, multi-disciplinary team that interfaced across the organization.

Data Science Journal

State of the data: Assessing the FAIRness of USGS data

In response to recent shifts towards open science that emphasize transparency, reproducibility, and access to research data, the US Geological Survey (USGS) conducted a study to assess the degree to which USGS data assets meet the FAIR data principles (Findable, Accessible, Interoperable, and Reusable). The USGS designed and applied a methodology for quantitative analysis of FAIR characteristics. A new rubric was derived from a crosswalk of existing FAIR evaluation frameworks and customized for the USGS. The rubric, consisting of 62 yes/no questions, was applied to 392 metadata records of USGS data products published between 1987 and 2022. Results were analyzed to show which FAIR characteristics were most and least present in the metadata and how these scores changed after the implementation of data policy requirements in 2016. Aggregated scores showed specific areas of strength and needed improvements. The greatest increases in FAIR scores over time were for elements that were required by new data policies, especially in the ‘Findable’ category. Based on the results, this paper presents strategies to further improve USGS alignment with FAIR. The suggested strategies are organized in four key areas: USGS data repository characteristics, training and communities of practice, data management policy considerations, and metadata standards, tools, and best practices.

Data Science Journal

Are researchers citing their data? A case study from the U.S. Geological Survey

Data citation promotes accessibility and discoverability of data through measures carried out by researchers, publishers, repositories, and the scientific community. This paper examines how a data citation workflow has been implemented by the U.S. Geological Survey (USGS) by evaluating publication and data linkages. Two different methods were used to identify data citations: examining publication structural metadata and examining the full text of the publication. A growing number of USGS researchers are complying with publisher data sharing policies aimed to capture data citation information in a standardized way within associated publications. However, inconsistencies in how data citation information is documented in publications has limited the accessibility and discoverability of the data. This paper demonstrates how organizational evaluations of publication and data linkages can be used to identify obstacles in advancing data citation efforts and improve data citation workflows.

Data Science Journal

From data to interpretable models: Machine learning for soil moisture forecasting

Soil moisture is critical to agricultural business, ecosystem health, and certain hydrologically driven natural disasters. Monitoring data, though, is prone to instrumental noise, wide ranging extrema, and nonstationary response to rainfall where ground conditions change. Furthermore, existing soil moisture models generally forecast poorly for time periods greater than a few hours. To improve such forecasts, we introduce two data-driven models, the Naive Accumulative Representation (NAR) and the Additive Exponential Accumulative Representation (AEAR). Both of these models are rooted in deterministic, physically based hydrology, and we study their capabilities in forecasting soilmoisture over time periods longer than a fewhours. Learned model parameters represent the physically based unsaturated hydrological redistribution processes of gravity and suction. We validate our models using soil moisture and rainfall time series data collected from a steep gradient, post-wildfire site in southern California. Data analysis is complicated by rapid landscape change observed in steep, burned hillslopes in response to even small to moderate rain events. The proposed NAR and AEAR models are, in forecasting experiments, shown to be competitive with several established and state-of-the-art baselines. The AEAR model fits the data well for three distinct soil textures at variable depths below the ground surface (5, 15, and 30 cm). Similar robust results are demonstrated in controlled, laboratory-based experiments. Our AEAR model includes readily interpretable hydrologic parameters and provides more accurate forecasts than existing models for time horizons of 10–24 h. Such extended periods of warning for natural disasters, such as floods and landslides, provide actionable knowledge to reduce loss of life and property.

International Journal of Data Science and Analytic

Estimating disease prevalence from preferentially sampled, pooled data

After the onset of the COVID-19 pandemic, scientific interest in coronaviruses endemic in animal populations has increased dramatically. However, investigating the prevalence of disease in animal populations across the landscape, which requires finding and capturing animals can be difficult. Spatial random sampling over a grid could be extremely inefficient because animals can be hard to locate, and the total number of samples may be small. Alternatively, preferential sampling, using existing knowledge to inform sample location, can guarantee larger numbers of samples, but estimates derived from this sampling scheme may exhibit bias if there is a relationship between higher probability sampling locations and the disease prevalence. Sample specimens are commonly grouped and tested in pools which can also be an added challenge when combined with preferential sampling. Here we present a Bayesian method for estimating disease prevalence with preferential sampling in pooled presence-absence data motivated by estimating factors related to coronavirus infection among Mexican free-tailed bats ( Tadarida brasiliensis ) in California. We demonstrate the efficacy of our approach in a simulation study, where a naive model, not accounting for preferential sampling, returns biased estimates of parameter values; however, our model returns unbiased results regardless of the degree of preferential sampling. Our model framework is then applied to data from California to estimate factors related to coronavirus prevalence. After accounting for preferential sampling impacts, our model suggests small prevalence differences between male and female bats.

California

Exploring the science and data foundation for Federal public lands decisions

Public lands provide diverse resources, values, and services worldwide. Laws and policies typically require consideration of science in public lands decisions, and resource managers are committed to science-informed decision-making. However, it can be challenging for managers to use, and document the use of, science and data in their decisions. To better understand science and data use in Federal public lands decisions in the United States, we assessed the number, type, and age of documents cited in 70 Environmental Assessments (EAs) completed by the Bureau of Land Management (BLM) in Colorado from 2015–2019. We focused on the BLM, as they manage the largest area of public lands in the United States. We selected Colorado as our study area, as actions proposed on BLM lands in Colorado are representative of those across the nation. Fifty percent of citations were categorized as science and 23% as data. EAs contained an average of 17 citations (range 0–111), with documents analyzing effects of oil and gas development and recreation actions including the highest and lowest mean number of citations (41 and 6, respectively). Of individual resource analysis sections within EAs, 24% contained ≥1 science citation and 21% contained ≥1 data citation. Journal articles were the most cited type of document (26% of citations) followed by non-BLM inventories (13%). Forty-seven percent of citations were relatively recent (2010 or later); the oldest citation was from 1927. Commonly analyzed resources with the highest mean number of citations were socioeconomics, mineral resources, and noise. Fourteen of 33 commonly analyzed resources included <1 citation on average. Actions and resources with no or few citations represent opportunities for strengthening the transparent use of science and data in public lands decision-making.

Colorado

Advances and applications of Unoccupied Aerial Systems (UAS) research in landscape ecology

Landscape ecologists have long depended on satellite and aerial remote sensing to address questions about landscape pattern and process, structure, and change (Foody 2023 ). Unoccupied aerial systems/vehicles (UAS/UAV, a.k.a. drones) technology is becoming an increasingly popular research tool in environmental sciences allowing scientists to generate low-cost, high-quality, and high-resolution imagery on demand that can be tailored to specific research questions. While satellite data are of a fixed resolution and temporal interval, UAS offer researchers control and flexibility to design studies and collect data at resolutions and scales that provide ecologically relevant information at finer spatial resolutions (e.g., < 30 cm) than what is currently available from satellite platforms (typically > 3m), thus helping capture objects such as individual plant canopies, micro-topography, and individual animals. Unlike satellites with fixed orbits, UAS can be deployed at more optimal temporal frequencies for ecological monitoring. We organized the special collection “Advances and Applications of Unoccupied Aerial Systems (UAS) Research in Landscape Ecology” to showcase the many ways that UAS tools and technologies are currently applied to advance landscape ecological research. When we announced the collection in 2023, only 11 papers published in the journal Landscape Ecology used UAS data, which was a notably small number compared to many other general ecology, environmental science and remote sensing journals. In an attempt to understand why UAS were not more widely used in landscape ecology and provide possible solutions, we published a review article (Villarreal et al. 2025) that identified the challenges, knowledge gaps, and obstacles for the adoption of UAS technologies in landscape ecology research. The main issues we identified include: (1) an abundance of UAS methods papers in the existing literature, with comparatively few studies demonstrating how UAS can be applied to address ecological questions; (2) a perceived scale mismatch between the geographic extent of UAS data collection (local) compared to larger study areas (landscapes) and a need to design robust scaling approaches to connect fine-scale UAS data with broader ecological patterns; and (3) a need for improved integration of UAS data with other commonly used remote sensing datasets including historical high resolution aerial imagery. Additionally, researchers new to UAS remote sensing may be discouraged or overwhelmed by the general lack of scientific consensus and standardized protocols for typical tasks such as data collection, vegetation classification, and change detection, as well as restrictive and/or confusing policy, regulatory, and legal issues surrounding UAS operations (Villarreal et al. 2025).

Landscape Ecology

Obtaining and applying public data for training students in technical statistical writing: Case studies with data from U.S. Geological Survey and general ecological literature

Effective undergraduate statistical education requires training using real-world data. Textbook datasets seldom match the complexities and messiness of real-world data and finding these datasets can be challenging for educators. Consulting and industrial datasets often have nondisclosure agreements. Academic datasets often require subject area expertise beyond those of a general education or lack connections to real-world applications. Many governments, including the United States, now require the release of data from projects they directly complete or fund though grants and contracts. We show how statistical educators may find datasets and incorporate them into courses. Specifically, we use two examples from the U.S. Geological Survey (USGS) and one example from the ecology literature. We demonstrate the use of these datasets in an upper-level analysis of variance (ANOVA) class. In addition to describing how we found the datasets, we describe how to include them into course work and the course’s student assessments. We have used these datasets over multiple semesters and included student feedback from these courses. Although our examples focus on an ANOVA class, the general methods for finding data shared here could be used for statistical classes ranging from high school to graduate education. Supplementary materials for this article are available online.

Journal of Statistics and Data Science Education

Studying biodiversity: is a new paradigm really needed?

Authors in this journal have recommended a new approach to the conduct of biodiversity science. This data-driven approach requires the organization of large amounts of ecological data, analysis of these data to discover complex patterns, and subsequent development of hypotheses corresponding to detected patterns. This proposed new approach has been contrasted with more-traditional knowledge-based approaches in which investigators deduce consequences of competing hypotheses to be confronted with actual data, providing a basis for discriminating among the hypotheses. We note that one approach is directed at hypothesis generation, whereas the other is also focused on discriminating among competing hypotheses. Here, we argue for the importance of using existing knowledge to the separate issues of (a) hypothesis selection and generation and (b) hypothesis discrimination and testing. In times of limited conservation funding, the relative efficiency of different approaches to learning should be an important consideration in decisions about how to study biodiversity.

BioScience

Evapotranspiration estimation using a normalized difference vegetation index transformation of satellite data

Evapotranspiration of irrigated crops on two irrigation service areas along the lower Colorado River was estimated using a normalized difference vegetation index of satellite data. A procedure was developed which equated the index to crop coefficients. Evapotranspiration estimates for fields for three dates of thematic mapper data were highly correlated with ground estimates. Service area estimates using thematic mapper and Advanced Very High Resolution Radiometer data agreed well with estimates based on US Geological Survey gauging station data.

Hydrological Sciences Journal

Advances in a distributed approach for ocean model data interoperability

An infrastructure for earth science data is emerging across the globe based on common data models and web services. As we evolve from custom file formats and web sites to standards-based web services and tools, data is becoming easier to distribute, find and retrieve, leaving more time for science. We describe recent advances that make it easier for ocean model providers to share their data, and for users to search, access, analyze and visualize ocean data using MATLAB® and Python®. These include a technique for modelers to create aggregated, Climate and Forecast (CF) metadata convention datasets from collections of non-standard Network Common Data Form (NetCDF) output files, the capability to remotely access data from CF-1.6-compliant NetCDF files using the Open Geospatial Consortium (OGC) Sensor Observation Service (SOS), a metadata standard for unstructured grid model output (UGRID), and tools that utilize both CF and UGRID standards to allow interoperable data search, browse and access. We use examples from the U.S. Integrated Ocean Observing System (IOOS®) Coastal and Ocean Modeling Testbed, a project in which modelers using both structured and unstructured grid model output needed to share their results, to compare their results with other models, and to compare models with observed data. The same techniques used here for ocean modeling output can be applied to atmospheric and climate model output, remote sensing data, digital terrain and bathymetric data.

Journal of Marine Science and Engineering

Teachers doing science: An authentic geology research experience for teachers

Fairmont State University (FSU) and the West Virginia Geological and Economic Survey (WVGES) provided a small pilot group of West Virginia science teachers with a professional development session designed to mimic experiences obtained by geology majors during a typical summer field camp. Called GEOTECH, the program served as a research capstone event complimenting the participants' multi-year association with the RockCamp professional development program. GEOTECH was funded through a Improving Teacher Quality Grant administered by West Virginia Higher Education Policy Commission. Over the course of three weeks, eight GEOTEACH participants learned field measurement and field data collection techniques which they then applied to the construction of a surficial geologic map. The program exposed participants to authentic scientific processes by emphasizing the authentic scientific application of content knowledge. As a secondary product, it also enhanced their appreciation of the true nature of science in general and geology particular. After the session, a new appreciation of the effort involved in making a geologic map emerged as tacit knowledge ready to be transferred to their students. The program was assessed using pre/post instruments, cup interviews, journals, artifacts (including geologic maps, field books, and described sections), performance assessments, and constructed response items. Evaluation of the accumulated data revealed an increase in participants demonstrated use of science content knowledge, an enhanced awareness and understanding of the processes and nature of geologic mapping, positive dispositions toward geologic research and a high satisfaction rating for the program. These findings support the efficacy of the experience and document future programmatic enhancements.

Journal of Geoscience Education

Knowledge inventory of foundational data products in planetary science

Some of the key components of any Planetary Spatial Data Infrastructure (PDSI) are the data products that end-users wish to discover, access, and interrogate. One precursor to the implementation of a PSDI is a knowledge inventory that catalogs what products are available, from which data producers, and at what initially understood data qualities. We present a knowledge inventory of foundational PSDI data products: geodetic coordinate reference frames, elevation or topography, and orthoimages or orthomosaics. Additionally, we catalog the available gravity models that serve as critical data for the assessment of spatial location, spatial accuracy, and ultimately spatial efficacy. We strengthen our previously published definitions of foundational data products to assist in solidifying a common vocabulary that will improve communication about these essential data products.

The Planetary Science Journal

Reply to: Turner, R.E., 2014. Discussion of: Olea, R.A. and Coleman, J.L., Jr., 2014. A synoptic examination of causes of land loss in southern Louisiana as related to the exploitation of subsurface geologic resources, Journal of Coastal Research, 30(5), 1025–1044; Journal of Coastal Research, 30(6), 1330–1334.

To a large extent, geology is a science of solving inverse problems based on some data and scientific principles. Solutions to these types of problems are not unique, especially when using different data, invoking different principles, or both. It is not surprising that the discussant and we have reached different conclusions on the same specific issue of land loss along the coast of Louisiana because we use different observations and view those observations in a different context. The objective of this reply is to orient the reader, who then can decide which approach is more likely to be the correct analysis.

Journal of Coastal Research