Search USGSSearch

SEARCH · Search USGS

Results for “Scientific Data - Nature”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

An underwater observation dataset for fish classification and fishery assessment

Using Dual-Frequency Identification Sonar (DIDSON), fishery acoustic observation data was collected from the Ocqueoc River, a tributary of Lake Huron in northern Michigan, USA. Data were collected March through July 2013 and 2016 and included the identification, via technology or expert analysis, of eight fish species as they passed through the DIDSON’s field of view. A set of short DIDSON clips containing identified fish was curated. Additionally, two other datasets were created that include visualizations of the acoustic data and longer DIDSON clips. These datasets could complement future research characterizing the abundance and behavior of valued fishes such as walleye ( Sander vitreus ) or white sucker ( Catostomus commersonii ) or invasive fishes such as sea lamprey ( Petromyzon marinus ) or European carp ( Cyprinus carpio ). Given the abundance of DIDSON data and the fact that a portion of it is labeled, these data could aid in the creation of machine learning tools from DIDSON data, particularly for invasive sea lamprey which are amply represented and a destructive invader of the Laurentian Great Lakes.

Scientific Data

Water quality measurements in San Francisco Bay by the U.S. Geological Survey, 1969–2015

The U.S. Geological Survey (USGS) maintains a place-based research program in San Francisco Bay (USA) that began in 1969 and continues, providing one of the longest records of water-quality measurements in a North American estuary. Constituents include salinity, temperature, light extinction coefficient, and concentrations of chlorophyll- a , dissolved oxygen, suspended particulate matter, nitrate, nitrite, ammonium, silicate, and phosphate. We describe the sampling program, analytical methods, structure of the data record, and how to access all measurements made from 1969 through 2015. We provide a summary of how these data have been used by USGS and other researchers to deepen understanding of how estuaries are structured and function differently from the river and ocean ecosystems they bridge.

California

Lilac and honeysuckle phenology data 1956–2014

The dataset is comprised of leafing and flowering data collected across the continental United States from 1956 to 2014 for purple common lilac ( Syringa vulgaris ), a cloned lilac cultivar (S. x chinensis ‘Red Rothomagensis’) and two cloned honeysuckle cultivars ( Lonicera tatarica ‘Arnold Red’ and L. korolkowii ‘Zabeli’). Applications of this observational dataset range from detecting regional weather patterns to understanding the impacts of global climate change on the onset of spring at the national scale. While minor changes in methods have occurred over time, and some documentation is lacking, outlier analyses identified fewer than 3% of records as unusually early or late. Lilac and honeysuckle phenology data have proven robust in both model development and climatic research.

Scientific Data

Meeting the challenge: U.S. Geological Survey North Atlantic and Appalachian Region fiscal year 2020 in review

The utilization, preservation, and conservation of the Nation’s resources requires well-informed management decisions. The North Atlantic and Appalachian Region (NAAR) of the U.S. Geological Survey (USGS) supports science-based decision making for Federal, State, and local policymakers to meet the challenges of today and into the future. The science centers in the NAAR have well-deserved reputations as world leaders in delivering unbiased science. We help protect the lives and property of our families, friends, neighbors, and the Nation by providing the data and scientific interpretation that decision makers need to make informed choices on a myriad of topics. Many of our jobs include inherent risk. When others are moving themselves and their families to higher ground during storms, NAAR employees can be found heading toward high water to ensure that accurate streamflow and storm-tide data continue to be collected and delivered to the public and first responders. In March 2020, the world changed, and the NAAR staff adapted to it. Despite the challenges, the NAAR has had an incredibly productive year. I am not just citing publications (with our labs and field offices closed in the spring, centers increased annual publications by 10 to 40 percent compared with 2019) or partnerships (new science initiatives and partnerships are up significantly as well). Leaders at the center level created the right environments for their teams to be safe but still meet and exceed their program goals. Our vast data collection networks were maintained and enhanced. Our laboratories met holding times and quality-control objectives. When folks asked for help, our staff provided. Some solutions were not perfect at first, but they just kept trying. What started as a short-term inconvenience may now have become the new normal, but in quickly adapting, the NAAR staff showed dedication and wisdom, made the region a little safer, and just might change the world. This general information product highlights just a few of the many accomplishments of the NAAR staff during these challenging times and offers a taste of all the great work being done by the USGS community.

Connecticut, Delaware, Kentucky, Maine, Maryland,

Status and distribution of mangrove forests of the world using earth observation satellite data

Aim Our scientific understanding of the extent and distribution of mangrove forests of the world is inadequate. The available global mangrove databases, compiled using disparate geospatial data sources and national statistics, need to be improved. Here, we mapped the status and distributions of global mangroves using recently available Global Land Survey (GLS) data and the Landsat archive. Methods We interpreted approximately 1000 Landsat scenes using hybrid supervised and unsupervised digital image classification techniques. Each image was normalized for variation in solar angle and earth–sun distance by converting the digital number values to the top-of-the-atmosphere reflectance. Ground truth data and existing maps and databases were used to select training samples and also for iterative labelling. Results were validated using existing GIS data and the published literature to map ‘true mangroves’. Results The total area of mangroves in the year 2000 was 137,760 km2 in 118 countries and territories in the tropical and subtropical regions of the world. Approximately 75% of world's mangroves are found in just 15 countries, and only 6.9% are protected under the existing protected areas network (IUCN I-IV). Our study confirms earlier findings that the biogeographic distribution of mangroves is generally confined to the tropical and subtropical regions and the largest percentage of mangroves is found between 5° N and 5° S latitude. Main conclusions We report that the remaining area of mangrove forest in the world is less than previously thought. Our estimate is 12.3% smaller than the most recent estimate by the Food and Agriculture Organization (FAO) of the United Nations. We present the most comprehensive, globally consistent and highest resolution (30 m) global mangrove database ever created. We developed and used better mapping techniques and data sources and mapped mangroves with better spatial and thematic details than previous studies.

Global Ecology and Biogeography

A geographic dataset of rocky reefs habitat areas of particular concern for the United States West Coast

The United States National Marine Fisheries Service determines “essential fish habitat (EFH)” for federally managed species in coordination with regional fishery management councils, considers adverse effects to those habitats, and provides information to further habitat conservation and enhancement. Identifying discrete subsets of EFH as “habitat areas of particular concern (HAPC)” can help focus conservation, management, and research efforts. In 2006, the Pacific Fishery Management Council designated rocky reefs along the United States (U.S.) West Coast as HAPCs for groundfishes because of their ecological significance, sensitivity to human impacts, and relative rarity. To better understand where rocky reefs occur, we (1) located rocky reef areas that were not included in the 2006 rocky reef dataset, and (2) incorporated best available data into a refined geographic dataset that enables visualization. Our update shows that rocky reefs are distributed throughout the U.S. West Coast continental margin, are patchier than previously known, and comprise 8% of the extent of all data inputs. This updated dataset will inform resource management decisions in coastal and marine environments.

California, Oregon, Washington

A continuously updated, geospatially rectified database of utility-scale wind turbines in the United States

Nearly 60,000 utility-scale wind turbines are installed in the United States as of July, 2019, representing over 97 gigawatts of electric power capacity; US wind turbine installations continue to grow at a rapid pace. Yet, until April 2018, no publicly-available, regularly updated data source existed to describe those turbines and their locations. Under a cooperative research and development agreement, analysts from three organizations collaborated to develop and release the United States Wind Turbine Database (USWTDB) - a publicly available, continuously updated, spatially rectified data source of locations and attributes of utility-scale wind turbines in the United States. Technical specifications and wind facility data, incorporated from five sources, undergo rigorous quality control. The location of each turbine is visually verified using high-resolution aerial imagery. The quarterly-updated data are available in a variety of formats, including an interactive web application, comma-separated values (CSV), shapefile, and application programming interface (API). The data are used widely by academic researchers, engineers and developers from wind energy companies, government agencies, planners, educators, and the general public.

Scientific Data

Plant macrofossil data for 48-0 ka in the USGS North American Packrat Midden Database, version 5.0

Plant macrofossils from packrat ( Neotoma spp.) middens provide direct evidence of past vegetation changes in arid regions of North America. Here we describe the newest version (version 5.0) of the U.S. Geological Survey (USGS) North American Packrat Midden Database. The database contains published and contributed data from 3,331 midden samples collected in southwest Canada, the western United States, and northern Mexico, with samples ranging in age from 48 ka to the present. The database includes original midden-sample macrofossil counts and relative-abundance data along with a standardized relative-abundance scheme that makes it easier to compare macrofossil data across midden-sample sites. In addition to the midden-sample data, this version of the midden database includes calibrated radiocarbon ( 14 C) ages for the midden samples and plant functional type (PFT) assignments for the midden taxa. We also provide World Wildlife Fund ecoregion assignments and climate and bioclimate data for each midden-sample site location. The data are provided in tabular (.xlsx), comma-separated values (.csv), and relational database (.mdb) files.

Scientific Data

Simplifying complex fault data for systems-level analysis: Earthquake geology inputs for U.S. NSHM 2023

As part of the U.S. National Seismic Hazard Model (NSHM) update planned for 2023, two databases were prepared to more completely represent Quaternary-active faulting across the western United States: the NSHM23 fault sections database (FSD) and earthquake geology database (EQGeoDB). In prior iterations of NSHM, fault sections were included only if a field-measurement-derived slip rate was estimated along a given fault. By expanding this inclusion criteria, we were able to assess a larger set of faults for use in NSHM23. The USGS Quaternary Fault and Fold Database served as a guide for assessing possible additions to the NSHM23 FSD. Reevaluating available data from published sources yielded an increase of fault sections from ~650 faults in NSHM18 to ~1,000 faults proposed for use in NSHM23. EQGeoDB, a companion dataset linked to NSHM23 FSD, contains geologic slip rate estimates for fault sections included in FSD. Together, these databases serve as common input data used in deformation modeling, earthquake rupture forecasting, and additional downstream uses in NSHM development.

Scientific Data

Development and validation of the CHIRTS-daily quasi-global high-resolution daily temperature data set

We present a high-resolution daily temperature data set, CHIRTS-daily, which is derived by merging the monthly Climate Hazards center InfraRed Temperature with Stations climate record with daily temperatures from version 5 of the European Centre for Medium-Range Weather Forecasts Re-Analysis. We demonstrate that remotely sensed temperature estimates may more closely represent true conditions than those that rely on interpolation, especially in regions with sparse in situ data. By leveraging remotely sensed infrared temperature observations, CHIRTS-daily provides estimates of 2-meter air temperature for 1983–2016 with a footprint covering 60°S-70°N. We describe this data set and perform a series of validations using station observations from two prominent climate data sources. The validations indicate high levels of accuracy, with CHIRTS-daily correlations with observations ranging from 0.7 to 0.9, and very good representation of heat wave trends.

Scientific Data

Community established best practice recommendations for tephra studies— From collection through analysis

Tephra is a unique volcanic product with an unparalleled role in understanding past eruptions, long-term behavior of volcanoes, and the effects of volcanism on climate and the environment. Tephra deposits also provide spatially widespread, high-resolution time-stratigraphic markers across a range of sedimentary settings and thus are used in numerous disciplines (e.g., volcanology, climate science, archaeology). Nonetheless, the study of tephra deposits is challenged by a lack of standardization that inhibits data integration across geographic regions and disciplines. We present comprehensive recommendations for tephra data gathering and reporting that were developed by the tephra science community to guide future investigators and to ensure that sufficient data are gathered for interoperability. Recommendations include standardized field and laboratory data collection, reporting and correlation guidance. These are organized as tabulated lists of key metadata with their definition and purpose. They are system independent and usable for template, tool, and database development. This standardized framework promotes consistent documentation and archiving, fosters interdisciplinary communication, and improves effectiveness of data sharing among diverse communities of researchers.

Scientific Data

GPS data from 2019 and 2020 campaigns in the Chesapeake Bay region towards quantifying vertical land motions

The Chesapeake Bay is a region along the eastern coast of the United States where sea-level rise is confounded with poorly resolved rates of land subsidence, thus new constraints on vertical land motions (VLM) in the region are warranted. In this paper, we provide a description of two campaign-style Global Positioning System (GPS) datasets, explain the methods used in data collection and validation, and present the experiment designed to quantify a new baseline of VLM in the Chesapeake Bay region of eastern North America. Data from GPS campaigns in 2019 and 2020 are presented as ASCII RINEX2.11 files and logsheets for each observation from the campaigns. Data were quality checked using the open-source program TEQC, resulting in average multipath 1 and 2 values of 0.68 and 0.57, respectively. All data are archived and publicly available for open access at the geodesy facility UNAVCO to abide by Findable, Accessible, Interoperable, Reusable (FAIR) data principles.

Chesapeake Bay area

Individual encounter data of six African carnivore species optimized for multi-species density estimation

The ability to estimate abundances of multiple wildlife species within an area is valuable for both conservation and ecological inquiry. Spatially explicit capture–recapture (SCR) methods are commonly used to obtain reliable population size estimates, particularly for low-density and individually identifiable carnivore species. However, estimating abundance within multi-species communities poses a methodological challenge as survey designs and analytical tools are primarily tailored for single target species. Here, we present a dataset of spatially referenced individual encounter histories of six carnivore species with varying space requirements (lion, Panthera leo ; leopard, Panthera pardus ; spotted hyena, Crocuta crocuta ; cheetah, Acinonyx jubatus ; serval, Leptailurus serval ; large-spotted genet, Genetta tigrina ). These data were collected in a South African game reserve using a camera trap array optimized for multi-species density estimation using SCR methods. This dataset will be a valuable resource for studying spatial processes among potentially interacting carnivores without the common pitfalls that come with by-catch data of non-target species, and will provide a much-needed case study for the further development of multi-species statistical method development.

Munywana Conservancy

A rasterized building footprint dataset for the United States

Microsoft released a U.S.-wide vector building dataset in 2018. Although the vector building layers provide relatively accurate geometries, their use in large-extent geospatial analysis comes at a high computational cost. We used High-Performance Computing (HPC) to develop an algorithm that calculates six summary values for each cell in a raster representation of each U.S. state, excluding Alaska and Hawaii: (1) total footprint coverage, (2) number of unique buildings intersecting each cell, (3) number of building centroids falling inside each cell, and area of the (4) average, (5) smallest, and (6) largest area of buildings that intersect each cell. These values are represented as raster layers with 30 m cell size covering the 48 conterminous states. We also identify errors in the original building dataset. We evaluate precision and recall in the data for three large U.S. urban areas. Precision is high and comparable to results reported by Microsoft while recall is high for buildings with footprints larger than 200 m2 but lower for progressively smaller buildings.

Scientific Data

A land data assimilation system for sub-Saharan Africa food and water security applications

Seasonal agricultural drought monitoring systems, which rely on satellite remote sensing and land surface models (LSMs), are important for disaster risk reduction and famine early warning. These systems require the best available weather inputs, as well as a long-term historical record to contextualize current observations. This article introduces the Famine Early Warning Systems Network (FEWS NET) Land Data Assimilation System (FLDAS), a custom instance of the NASA Land Information System (LIS) framework. The FLDAS is routinely used to produce multi-model and multi-forcing estimates of hydro-climate states and fluxes over semi-arid, food insecure regions of Africa. These modeled data and derived products, like soil moisture percentiles and water availability, were designed and are currently used to complement FEWS NET’s operational remotely sensed rainfall, evapotranspiration, and vegetation observations. The 30+ years of monthly outputs from the FLDAS simulations are publicly available from the NASA Goddard Earth Science Data and Information Services Center (GES DISC) and recommended for use in hydroclimate studies, early warning applications, and by agro-meteorological scientists in Eastern, Southern, and Western Africa.

Scientific Data

Georectified polygon database of ground-mounted large-scale solar photovoltaic sites in the United States

Over 4,400 large-scale solar photovoltaic (LSPV) facilities operate in the United States as of December 2021, representing more than 60 gigawatts of electric energy capacity. Of these, over 3,900 are ground-mounted LSPV facilities with capacities of 1 MWdc or more. Ground mounted LSPV installations continue increasing, with more than 400 projects appearing online in 2021 alone; however, a comprehensive, publicly available georectified dataset including spatial footprints of these facilities is lacking. Analysts from U.S. Over 4,400 large-scale solar photovoltaic (LSPV) facilities operate in the United States as of December 2021, representing more than 60 gigawatts of electric energy capacity. Of these, over 3,900 are ground-mounted LSPV facilities with capacities of 1 megawatt direct current (MW dc ) or more. Ground-mounted LSPV installations continue increasing, with more than 400 projects appearing online in 2021 alone; however, a comprehensive, publicly available georectified dataset including spatial footprints of these facilities is lacking. The United States Large-Scale Solar Photovoltaic Database (USPVDB) was developed to fill this gap. Using US Energy Information Administration (EIA) data, locations of 3,699 LSPV facilities were verified using high-resolution aerial imagery, polygons were digitized around panel arrays, and attributes were appended. Quality assurance and control were achieved via team peer review and comparison to other US PV datasets. Data are publicly available via an interactive web application and multiple downloadable formats, including: comma-separated value (CSV), application programming interface (API), and GIS shapefile and GeoJSON. Survey and Lawrence Berkeley National Laboratory collaborated to develop the United States Large-Scale Solar Photovoltaic Database (USPVDB). Using Energy Information Administration (EIA) data, locations of LSPV facilities were verified using high-resolution aerial imagery, polygons were digitized around panel arrays, and attributes were appended. Quality assurance and control were achieved via team peer review and comparison to other US PV datasets. Data are publicly available in an interactive web application, and a number of downloadable formats, including: comma-separated value spreadsheet (CSV), application programming interface (API), and GIS shapefile.

Scientific Data

Tidal wetland soil carbon accumulation rates for coastal California

Carbon stock and carbon accumulation rate data are vital to multiple aspects of tidal wetland conservation and restoration policy. In California, USA tidal soil data are rare outside of the San Francisco Bay and Sacramento Delta regions, despite the differing conditions experienced by the outer coastline. Here we provide carbon stocks and decadal-to-centennial-scale carbon accumulation rate calculations. This dataset presents 83 soil depth profiles from 15 sites, with 58 cores from 12 tidal wetland sites analyzed for carbon stock, mostly from the outer coastline of California. Mean organic matter content was 11%, and stocks estimated to 1 meter depth ranged from 15.4 to 44.7 kgC m −2 . Organic matter content generally declined asymptotically with depth. Carbon accumulation rates ranged from 39.2 to 130.0 gC m −2 yr −1 . Neither carbon stock nor carbon accumulation rates were notably different from global average values. Data at this level of reporting are vital for establishing restoration baselines, informing greenhouse gas mitigation planning, and projecting future ecosystem response to sea-level rise.

California

Reconstructing Great Lakes air temperature and ice dynamics data back to 1897

Ice cover on the Great Lakes plays an important role in regional climate, supports tourism and recreation, and provides ecological habitat. As the climate warms, ice cover in the Great Lakes is expected to decline, which in turn will create more lake effect precipitation, reduce ice cover for recreation, and alter habitat for aquatic species. While it is important to understand the historical ice patterns to better understand past distributions of aquatic species and improve the accuracy of forecasts for future ice cover on the lakes, Great Lakes ice cover data prior to 1973 is scarce, due to the limited routine satellite observations. We used weather station data around the Great Lakes to compile daily air temperature, calculate cumulative freezing degree-days and net melting degree-days from 1897–2023, and develop raster layers estimating ice duration and variability spatially during the historical period from 1897–1960.

Great Lakes