Search USGS⌕ Search

SEARCH · Search USGS

Results for “Data”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Water resources data, Pennsylvania, water year 2000. Volume 3. Ohio and St. Lawrence River Basins

Introduction The Water Resources Division of the U.S. Geological Survey, in cooperation with State, municipal, and Federal agencies, collects a large amount of data pertaining to the water resources of Pennsylvania each water year. These data, accumulated during many water years, constitute a valuable data base for developing an improved understanding of the water resources of the State. To make these data readily available to interested parties outside the Geological Survey, these data are published annually in this report series entitled "Water Resources Data - Pennsylvania, Volumes 1, 2, and 3." Volume 1 contains data for the Delaware River Basin; Volume 2, the Susquehanna and Potomac River Basins; and Volume 3, the Ohio and St. Lawrence River Basins. This report, Volume 3, contains: (1) discharge records for 58 continuous-record streamflow-gaging stations, 5 partial-record stations, and 12 special study and miscellaneous streamflow sites; (2) elevation and contents records for 11 lakes and reservoirs; (3) water-quality records for 1 streamflow gaging station and 8 ungaged streamsites; and (4) water-level records for 15 ground-water network observation wells and. Additional water data collected at various sites not involved in the systematic data-collection program may also be presented. Publications similar to this report are published annually by the Geological Survey for all States. For the purpose of archiving, these official reports have an identification number consisting of the two-letter State abbreviation, the last two digits of the water year, and the volume number. For example, this volume is identified as "U.S. Geological Survey Water-Data Report PA-00-3." These water-data reports, beginning with the 1971 water year, are for sale as paper copy or microfiche by the National Technical Information Service, U.S. Department of Commerce, Springfield, VA 22161. The annual series of Water Data Reports for Pennsylvania began with the 1961 water-year report and contained only data relating to quantities of surface water. With the 1964 water year, a companion report (part 2) was introduced that contained only data relating to water quality. Beginning with the 1975 water year the report was changed to three volumes (by river basin), with each volume containing data on quantities of surface water, quality of surface and ground water, and ground-water levels. Prior to the introduction of this series and for several years concurrent with it, water-resources data for Pennsylvania were published in U.S. Geological Survey Water-Supply Papers. Data on stream discharge and stage, and on lake or reservoir contents and stage, through September 1960, were published annually under the title "Surface-Water Supply of the United States," which was released in numbered parts as determined by natural drainage basins. For the 1961-70 water years, these data were published in two 5-year reports. Data prior to 1961 are included in two reports: "Compilation of Records of Surface Waters of the United States through 1950," and "Compilation of Records of Surface Waters of the United States, October 1950 to September 1960." Data for Pennsylvania are published in Parts 1, 3, and 4. Data on chemical quality, temperature, and suspended sediment for the 1941-70 water years were published annually under the title "Quality of Surface Waters of the United States," and ground-water levels for the 1935-74 water years were published annually under the title "Ground-Water Levels in the United States." The above mentioned Water-Supply Papers may be consulted in the libraries of the principal cities of the United States and may be purchased from the U.S. Geological Survey, Information Services, Box 25286, Denver, CO 80225. Information for ordering specific reports may be obtained from the Pennsylvania District Office at the address on the back of the title page or by phoning the Scientific and Technical Products Section at (717) 730-6940. Information on the availability of unpublished data or statistical analyses may be obtained from the District Information Specialist by telephone at (717) 730-6916 or by FAX at (717) 730-6997.

Water Data Report↗

Water Resources Data, Pennsylvania, Water Year 2001, Volume 2. Susquehanna and Potomac River Basins

Introduction The Water Resources Division of the U.S. Geological Survey, in cooperation with State, municipal, and Federal agencies, collects a large amount of data pertaining to the water resources of Pennsylvania each water year. These data, accumulated during many water years, constitute a valuable data base for developing an improved understanding of the water resources of the State. To make these data readily available to interested parties outside the Geological Survey, these data are published annually in this report series entitled "Water Resources Data - Pennsylvania, Volumes 1, 2, and 3." Volume 1 contains data for the Delaware River Basin; Volume 2, the Susquehanna and Potomac River Basins; and Volume 3, the Ohio and St. Lawrence River Basins. This report, Volume 2, contains: (1) discharge records for 83 continuous-record streamflow-gaging stations, 15 partial-record stations, and 24 special study and miscellaneous streamflow sites; (2) elevation and contents records for 12 lakes and reservoirs; (3) water-quality records for 9 streamflow gaging stations and 73 partial-record and project stations; and (4) water-level records for 36 ground-water network observation wells and water-quality analyses of ground water from 8 wells; (5) water-quality analyses at 123 special study ground-water wells; and, (6) miscellaneous water-level measurements at 80 special study ground-water wells. Additional water data collected at various sites not involved in the systematic data-collection program may also be presented. Publications similar to this report are published annually by the Geological Survey for all States. For the purpose of archiving, these official reports have an identification number consisting of the two-letter State abbreviation, the last two digits of the water year, and the volume number. For example, this volume is identified as "U.S. Geological Survey Water-Data Report PA-01-2." These water-data reports, beginning with the 1971 water year, are for sale as paper copy or microfiche by the National Technical Information Service, U.S. Department of Commerce, Springfield, VA 22161. The annual series of Water Data Reports for Pennsylvania began with the 1961 water-year report and contained only data relating to quantities of surface water. With the 1964 water year, a companion report (part 2) was introduced that contained only data relating to water quality. Beginning with the 1975 water year the report was changed to three volumes (by river basin), with each volume containing data on quantities of surface water, quality of surface and ground water, and ground-water levels. Prior to the introduction of this series and for several years concurrent with it, water-resources data for Pennsylvania were published in U.S. Geological Survey Water-Supply Papers. Data on stream discharge and stage, and on lake or reservoir contents and stage, through September 1960, were published annually under the title "Surface-Water Supply of the United States," which was released in numbered parts as determined by natural drainage basins. For the 1961-70 water years, these data were published in two 5-year reports. Data prior to 1961 are included in two reports: "Compilation of Records of Surface Waters of the United States through 1950," and "Compilation of Records of Surface Waters of the United States, October 1950 to September 1960." Data for Pennsylvania are published in Parts 1, 3, and 4. Data on chemical quality, temperature, and suspended sediment for the 1941-70 water years were published annually under the title "Quality of Surface Waters of the United States," and ground-water levels for the 1935-74 water years were published annually under the title "Ground-Water Levels in the United States." The above mentioned Water-Supply Papers may be consulted in the libraries of the principal cities of the United States and may be purchased from the U.S. Geological Survey, Information Services, Box 25286, Denver, CO 80225. Information for ordering specific reports may be obtained from the Pennsylvania District Office at the address on the back of the title page or by phoning the Scientific and Technical Products Section at (717) 730-6940. Information on the availability of unpublished data or statistical analyses may be obtained from the District Information Specialist by telephone at (717) 730-6916 or by FAX at (717) 730-6997.

Water Data Report↗

Description of Existing Data for Integrated Landscape Monitoring in the Puget Sound Basin, Washington

This report summarizes existing geospatial data and monitoring programs for the Puget Sound Basin in northwestern Washington. This information was assembled as a preliminary data-development task for the U.S. Geological Survey (USGS) Puget Sound Integrated Landscape Monitoring (PSILM) pilot project. The PSILM project seeks to support natural resource decision-making by developing a 'whole system' approach that links ecological processes at the landscape level to the local level (Benjamin and others, 2008). Part of this effort will include building the capacity to provide cumulative information about impacts that cross jurisdictional and regulatory boundaries, such as cumulative effects of land-cover change and shoreline modification, or region-wide responses to climate change. The PSILM project study area is defined as the 23 HUC-8 (hydrologic unit code) catchments that comprise the watersheds that drain into Puget Sound and their near-shore environments. The study area includes 13 counties and more than four million people. One goal of the PSILM geospatial database is to integrate spatial data collected at multiple scales across the Puget Sound Basin marine and terrestrial landscape. The PSILM work plan specifies an iterative process that alternates between tasks associated with data development and tasks associated with research or strategy development. For example, an initial work-plan goal was to delineate the study area boundary. Geospatial data required to address this task included data from ecological regions, watersheds, jurisdictions, and other boundaries. This assemblage of data provided the basis for identifying larger research issues and delineating the study-area boundary based on these research needs. Once the study-area boundary was agreed upon, the next iteration between data development and research activities was guided by questions about data availability, data extent, data abundance, and data types. This report is not intended as an exhaustive compilation of all available geospatial data, rather, it is a collection of information about geospatial data that can be used to help answer the suite of questions posed after the study-area boundary was defined. This information will also be useful to the PSILM team for future project tasks, such as assessing monitoring gaps, exploring monitoring-design strategies, identifying and deriving landscape indicators and metrics, and visual geographic communication. The two main geospatial data types referenced in this report - base-reference layers and monitoring data - originated from numerous and varied sources. In addition to collecting information and metadata about the base-reference layers, the data also were collected for project needs, such as developing maps for visual communication among team members and with outside groups. In contrast, only information about the data was typically required for the monitoring data. The information on base-reference layers and monitoring data included in this report is only as detailed as what was readily available from the sources themselves. Although this report may appear to lack consistency between data records, the varying degree of details contained in this report are merely a reflection of varying source detail. This compilation is just a beginning. All data listed also are being catalogued in spreadsheets and knowledge-management systems. Our efforts are continual as we develop a geospatial catalog for the PSILM pilot project.

Open-File Report↗

NBII-SAIN Data Management Toolkit

The Strategic Plan for the U.S. Geological Survey Biological Informatics Program (2005-2009) recognizes the need for effective data management: Though the Federal government invests more than $600 million per year in biological data collection, it is difficult to address these issues because of limited accessibility and lack of standards for data and information...variable quality, sources, methods, and formats (for example observations in the field, museum specimens, and satellite images) present additional challenges. This is further complicated by the fast-moving target of emerging and changing technologies such as GPS and GIS. Even though these technologies offer new solutions, they also create new informatics challenges (Ruggiero and others, 2005). The USGS National Biological Information Infrastructure program, hereafter referred to as NBII, is charged with the mission to improve the way data and information are gathered, documented, stored, and accessed. The central objective of this project is a direct reflection of the purpose of NBII as described by John Mosesso, Program Manager of the U.S. Geological Survey-Biological Informatics Program-GAP Analysis: At the outset, the reason for bringing about NBII was that there were significant amounts of data and information scattered all over the U.S., not accessible, in incompatible formats, and that NBII was tasked with addressing this problem...NBII's focus is to pull data together that truly matters to someone or communities. Essentially, the core questions are: 1) what are the issues, 2) where is the data, and 3) how can we make it usable and accessible (John Mosesso, U.S. Geological Survey, oral commun., 2006). Redundancy in data collection can be a major issue when multiple stakeholders are involved with a common effort. In 2001 the U.S. General Accounting Office (USGAO) estimated that about 50 percent of the Federal government's geospatial data at the time was redundant. In addition, approximately 80 percent of the cost of a spatial information system is associated with spatial data collection and management (U.S. General Accounting Office, 2003). These figures indicate that the resources (time, personnel, money) of many agencies and organizations could be used more efficiently and effectively. Dedicated and conscientious data management coordination and documentation is critical for reducing such redundancy. Substantial cost savings and increased efficiency are direct results of a pro-active data management approach. In addition, details of projects as well as data and information are frequently lost as a result of real-world occurrences such as the passing of time, job turnover, and equipment changes and failure. A standardized, well documented database allows resource managers to identify issues, analyze options, and ultimately make better decisions in the context of adaptive management (National Land and Water Resources Audit and the Australia New Zealand Land Information Council on behalf of the Australian National Government, 2003). Many environmentally focused, scientific, or natural resource management organizations collect and create both spatial and non-spatial data in some form. Data management appropriate for those data will be contingent upon the project goal(s) and objectives and thus will vary on a case-by-case basis. This project and the resulting Data Management Toolkit, hereafter referred to as the Toolkit, is therefore not intended to be comprehensive in terms of addressing all of the data management needs of all projects that contain biological, geospatial, and other types of data. The Toolkit emphasizes the idea of connecting a project's data and the related management needs to the defined project goals and objectives from the outset. In that context, the Toolkit presents and describes the fundamental components of sound data and information management that are common to projects involving biological, geospatial, and other related data

Open-File Report↗

Water-quality, streamflow, and meteorological data for the Tualatin River Basin, Oregon, 1991-93

Surface-water-quality data, ground-water-quality data, streamflow data, field measurements, aquatic-biology data, meteorological data, and quality-assurance data were collected in the Tualatin River Basin from 1991 to 1993 by the U.S. Geological Survey (USGS) and the Unified Sewerage Agency of Washington County, Oregon (USA). The data from that study, which are part of this report, are presented in American Standard Code for Information Interchange (ASCII) format in subject-specific data files on a Compact Disk-Read Only Memory (CD-ROM). The text of this report describes the objectives of the study, the location of sampling sites, sample-collection and processing techniques, equipment used, laboratory analytical methods, and quality-assurance procedures. The data files on CD-ROM contain the analytical results of water samples collected in the Tualatin River Basin, streamflow measurements of the main-stem Tualatin River and its major tributaries, flow data from the USA wastewater-treatment plants, flow data from stations that divert water from the main-stem Tualatin River, aquatic-biology data, and meteorological data from the Tualatin Valley Irrigation District (TVID) Agrimet Weather Station located in Verboort, Oregon. Specific information regarding the contents of each data file is given in the text. The data files use a series of letter codes that distinguish each line of data. These codes are defined in data tables accompanying the text. Presenting data on CD-ROM offers several advantages: (1) the data can be accessed easily and manipulated by computers, (2) the data can be distributed readily over computer networks, and (3) the data may be more easily transported and stored than a large printed report. These data have been used by the USGS to (1) identify the sources, transport, and fate of nutrients in the Tualatin River Basin, (2) quantify relations among nutrient loads, algal growth, low dissolved-oxygen concentrations, and high pH, and (3) develop and calibrate a water- quality model that allows managers to test options for alleviating water-quality problems.

Oregon↗

Comparison of TOPMODEL streamflow simulations using NEXRAD-based and measured rainfall data, McTier Creek watershed, South Carolina

Rainfall is an important forcing function in most watershed models. As part of a previous investigation to assess interactions among hydrologic, geochemical, and ecological processes that affect fish-tissue mercury concentrations in the Edisto River Basin, the topography-based hydrological model (TOPMODEL) was applied in the McTier Creek watershed in Aiken County, South Carolina. Measured rainfall data from six National Weather Service (NWS) Cooperative (COOP) stations surrounding the McTier Creek watershed were used to calibrate the McTier Creek TOPMODEL. Since the 1990s, the next generation weather radar (NEXRAD) has provided rainfall estimates at a finer spatial and temporal resolution than the NWS COOP network. For this investigation, NEXRAD-based rainfall data were generated at the NWS COOP stations and compared with measured rainfall data for the period June 13, 2007, to September 30, 2009. Likewise, these NEXRAD-based rainfall data were used with TOPMODEL to simulate streamflow in the McTier Creek watershed and then compared with the simulations made using measured rainfall data. NEXRAD-based rainfall data for non-zero rainfall days were lower than measured rainfall data at all six NWS COOP locations. The total number of concurrent days for which both measured and NEXRAD-based data were available at the COOP stations ranged from 501 to 833, the number of non-zero days ranged from 139 to 209, and the total difference in rainfall ranged from -1.3 to -21.6 inches. With the calibrated TOPMODEL, simulations using NEXRAD-based rainfall data and those using measured rainfall data produce similar results with respect to matching the timing and shape of the hydrographs. Comparison of the bias, which is the mean of the residuals between observed and simulated streamflow, however, reveals that simulations using NEXRAD-based rainfall tended to underpredict streamflow overall. Given that the total NEXRAD-based rainfall data for the simulation period is lower than the total measured rainfall at the NWS COOP locations, this bias would be expected. Therefore, to better assess the use of NEXRAD-based rainfall estimates as compared to NWS COOP rainfall data on the hydrologic simulations, TOPMODEL was recalibrated and updated simulations were made using the NEXRAD-based rainfall data. Comparisons of observed and simulated streamflow show that the TOPMODEL results using measured rainfall data and NEXRAD-based rainfall are comparable. Nonetheless, TOPMODEL simulations using NEXRAD-based rainfall still tended to underpredict total streamflow volume, although the magnitude of differences were similar to the simulations using measured rainfall. The McTier Creek watershed was subdivided into 12 subwatersheds and NEXRAD-based rainfall data were generated for each subwatershed. Simulations of streamflow were generated for each subwatershed using NEXRAD-based rainfall and compared with subwatershed simulations using measured rainfall data, which unlike the NEXRAD-based rainfall were the same data for all subwatersheds (derived from a weighted average of the six NWS COOP stations surrounding the basin). For the two simulations, subwatershed streamflow were summed and compared to streamflow simulations at two U.S. Geological Survey streamgages. The percentage differences at the gage near Monetta, South Carolina, were the same for simulations using measured rainfall data and NEXRAD-based rainfall. At the gage near New Holland, South Carolina, the percentage differences using the NEXRAD-based rainfall were twice as much as those using the measured rainfall. Single-mass curve comparisons showed an increase in the total volume of rainfall from north to south. Similar comparisons of the measured rainfall at the NWS COOP stations showed similar percentage differences, but the NEXRAD-based rainfall variations occurred over a much smaller distance than the measured rainfall. Nonetheless, it was concluded that in some cases, using NEXRAD-based rainfall data in TOPMODEL streamflow simulations may provide an effective alternative to using measured rainfall data. For this investigation, however, TOPMODEL streamflow simulations using NEXRAD-based rainfall data for both calibration and simulations did not show significant improvements with respect to matching observed streamflow over simulations generated using measured rainfall data.

South Carolina↗

Summary of West Virginia Water-Resource Data through September 2008

The West Virginia Water Science Center of the U.S. Geological Survey, in cooperation with State and Federal agencies, obtains a large amount of data pertaining to the water resources of West Virginia each water year. A water year is the 12-month period beginning October 1 and ending September 30. These data, accumulated during many years, constitute a valuable database for developing an improved understanding of the water resources of the State. These data are maintained in the National Water Information System (NWIS) and are available through its World-Wide Web interface, NWISWeb, at http://waterdata.usgs.gov/wv/nwis. Data can be retrieved in a variety of common formats, and a tutorial is available at http://nwis.waterdata.usgs.gov/tutorial. Location information for all continuous-record gaging stations operated in West Virginia through September 2008 is provided in this report, as well as statistical summaries of the available daily records. This report can serve as an index to the daily records data available on the World-Wide Web. Hydrologic data for nearly all of the gaging stations identified in this report are also available in the annual publication series titled Water-Resources Data - West Virginia. This series of annual reports for West Virginia began with the 1961 water year with a report that contained only data relating to quantities of surface water. For the 1964 water year, a similar report was introduced that contained only data relating to water quality. Beginning with the 1975 water year, the report format was changed to include data on quantities of surface water, quality of surface water and groundwater, and groundwater levels. Prior to the introduction of the Water-Resources Data - West Virginia series and for several water years concurrent with it, water-resources data for West Virginia were published in U.S. Geological Survey Water-Supply Papers. Data on stream discharge and stage and on lake or reservoir contents and stage through September 1960 were published annually under the title Surface-Water Supply of the United States, Parts 6A and 6B. For the 1961 through 1970 water years, the data were published in two 5-year reports. Data on chemical quality, temperature, and suspended sediment for the 1941 through 1970 water years were published annually under the title Quality of Surface Water of the United States, and water levels for the 1935 through 1974 water years were published under the title Ground-Water Levels in the United States. Many of the above mentioned Water-Supply Papers are available at the USGS Publications Warehouse (http://pubs.er.usgs.gov), and most of the others may be found in the collections of large libraries or may be purchased from the U.S. Geological Survey, Books and Open-File Reports, Federal Center, Box 25425, Denver, Colorado 80225. Annual reports on hydrologic data are published by the Geological Survey for all states, and each has an identification number consisting of the two-letter state abbreviation, the last two digits of the water year, and the volume number. For example, the 2005 water year report for West Virginia is identified as U.S. Geological Survey Water-Data Report WV-05-01. Water-Data Reports for West Virginia for 2001-2005 are available online at http://pubs.usgs.gov/wdr/#WV. Water-Data Reports for water years prior to 2006 are for sale in paper copy or microfiche by the National Technical Information Service, U.S. Department of Commerce, Springfield, Virginia 22161. Since the 2006 water year, the report is published online only and is available at http://wdr.water.usgs.gov/. When substantial errors in published records are discovered, the records are revised. Such revisions are routine and are made to records regardless of the age of the original records. Revisions have been made for many stations for which data are published in this report. The USGS National Water Information System always contains the most recent data revisions. For critical a

Open-File Report↗

Promoting synergy in the innovative use of environmental data—Workshop summary

From December 2 to 4, 2015, NatureServe and the U.S. Geological Survey organized and hosted a biodiversity and ecological informatics workshop at the U.S. Department of the Interior in Washington, D.C. The workshop objective was to identify user-driven future directions and areas of collaboration in advanced applications of environmental data applied to forecasting and decision making for the sustainability of biodiversity and ecosystem services. Substantial effort to recruit attendees from diverse Federal, State, and private sector organizations successfully attracted participants from 20 Federal agencies and 48 different institutions in the academic, nonprofit, State government, and commercial sectors; the total number of attendees ranged from 100 to 144 during the 3-day workshop. The first one-half of the workshop was divided into 7 plenary sessions and 3 sets of lightning talk sessions organized by sector, providing 48 oral and visual plenary presentations that shared diverse perspectives on biodiversity and ecological informatics, including original biospatial analyses from 6 graduate student map contest winners. The second one-half of the workshop focused on 10 breakout sessions with participant-driven themes from the environmental data sphere and concluded with an address by the Director of the U.S. Fish and Wildlife Service. The workshop was structured to encourage interactivity. About 80–90 percent of attendees provided direct feedback using clicker devices for specific questions related to biodiversity and ecological data uses and needs, and 10 breakout session leaders shared the highlights of their group discussions during the final workshop plenary sessions. Participants were encouraged to use the Twitter hashtag #ShareUrData. Over lunch on day 2 there were 20 simultaneous presentations of tools and apps during a special “Tools Café” session. The 10 participant-defined breakout session topics are listed below: Ecosystem services and ecological indicators Inventory and monitoring Biogeographic map of the Nation Pollinators Invasive species Remote sensing Drivers of agricultural change Citizen science Climate Hydrology and watersheds Numerous common themes that emerged from the workshop include the following: The vital importance of completing foundational environmental datasets that are nationally consistent and are essential to multiple sectors, such as the Soil Survey Geographic database high-resolution soils data, a minimum 5-meter resolution digital elevation model, national hydrographic data, high-resolution land cover data, time series high-resolution spatial climate data from historical to future time steps, and a national wetland inventory. Improved, nationally consistent environmental datasets (integrated with targeted observations) will dramatically advance forecasting capacity and support early warning systems (that is, drought, forest disease); however, multiagency coordination should focus on decision support tools that convey appropriate actions and responses to adapt to, and mitigate, potential negative consequences. Digitizing and providing access to the vast stores of underused historical data that can be leveraged for this purpose is of national importance. Modern computational techniques and the ever-increasing flow of environmental data from ground and remote observations can support improved understanding of environmental change. Success of understanding patterns of change for decision making requires establishing baselines from which change can be measured. The value of digitized historical data is greater than ever before. There is a need to recognize the multifaceted potential of citizen science to engage the public in resource stewardship, to create the next generation of science, technology, engineering, math, and environmental leaders, and to have sufficient field personnel to monitor environmental trends, including early detection of alien invasive species, phenological shifts, shifting distribution and abundance of indicator species, and species inventories. The Federal government has an essential role in creating the infrastructure to dramatically improve mobilization of citizen science (and other) data by fostering the following: creation of data standards, creation of nationally consistent framework datasets, vertical integration of observation data, visualization and dissemination of aggregated datasets, and calculation and communication of derived trends. Current and near future trends in the availability of remotely sensed data (rapid expansion of satellite fleets and drones) is revolutionizing access to near-real-time ecological data. Targeted integration with ground-based observations and instrumentation has an extremely valuable role in validating remotely sensed data, filling data gaps, improving data quality, and fully realizing the potential of the near-real-time monitoring of environmental indicator trends. Integrated management of environmental data at the landscape scale is required even as specific actions on the ground are largely local in nature. The workshop highlighted numerous success stories; however, almost every breakout group pointed out the still-too-fragmented nature of the current data landscape. Management and delivery of the necessary data, tools, and analyses to sustain our Nation’s environmental capital must be a collaborative effort between Federal, State, and local governments, academia, nonprofits, and the commercial sector, even though the responsibilities of each sector are different.

Open-File Report↗

Department of the Interior metadata implementation guide—Framework for developing the metadata component for data resource management

The Department of the Interior (DOI) is a Federal agency with over 90,000 employees across 10 bureaus and 8 agency offices. Its primary mission is to protect and manage the Nation’s natural resources and cultural heritage; provide scientific and other information about those resources; and honor its trust responsibilities or special commitments to American Indians, Alaska Natives, and affiliated island communities. Data and information are critical in day-to-day operational decision making and scientific research. DOI is committed to creating, documenting, managing, and sharing high-quality data and metadata in and across its various programs that support its mission. Documenting data through metadata is essential in realizing the value of data as an enterprise asset. The completeness, consistency, and timeliness of metadata affect users’ ability to search for and discover the most relevant data for the intended purpose; and facilitates the interoperability and usability of these data among DOI bureaus and offices. Fully documented metadata describe data usability, quality, accuracy, provenance, and meaning. Across DOI, there are different maturity levels and phases of information and metadata management implementations. The Department has organized a committee consisting of bureau-level points-of-contacts to collaborate on the development of more consistent, standardized, and more effective metadata management practices and guidance to support this shared mission and the information needs of the Department. DOI’s metadata implementation plans establish key roles and responsibilities associated with metadata management processes, procedures, and a series of actions defined in three major metadata implementation phases including: (1) Getting started—Planning Phase, (2) Implementing and Maintaining Operational Metadata Management Phase, and (3) the Next Steps towards Improving Metadata Management Phase. DOI’s phased approach for metadata management addresses some of the major data and metadata management challenges that exist across the diverse missions of the bureaus and offices. All employees who create, modify, or use data are involved with data and metadata management. Identifying, establishing, and formalizing the roles and responsibilities associated with metadata management are key to institutionalizing a framework of best practices, methodologies, processes, and common approaches throughout all levels of the organization; these are the foundation for effective data resource management. For executives and managers, metadata management strengthens their overarching views of data assets, holdings, and data interoperability; and clarifies how metadata management can help accelerate the compliance of multiple policy mandates. For employees, data stewards, and data professionals, formalized metadata management will help with the consistency of definitions, and approaches addressing data discoverability, data quality, and data lineage. In addition to data professionals and others associated with information technology; data stewards and program subject matter experts take on important metadata management roles and responsibilities as data flow through their respective business and science-related workflows. The responsibilities of establishing, practicing, and governing the actions associated with their specific metadata management roles are critical to successful metadata implementation.

Techniques and Methods↗

Alaska Geochemical Database Version 3.0 (AGDB3)—Including “Best Value” Data Compilations for Rock, Sediment, Soil, Mineral, and Concentrate Sample Media

The Alaska Geochemical Database Version 3.0 (AGDB3) contains new geochemical data compilations in which each geologic material sample has one “best value” determination for each analyzed species, greatly improving speed and efficiency of use. Like the Alaska Geochemical Database Version 2.0 before it, the AGDB3 was created and designed to compile and integrate geochemical data from Alaska to facilitate geologic mapping, petrologic studies, mineral resource assessments, definition of geochemical baseline values and statistics, element concentrations and associations, environmental impact assessments, and studies in public health associated with geology. This relational database, created from data-bases and published datasets of the U.S. Geological Survey (USGS), Atomic Energy Commission National Uranium Resource Evaluation (NURE), Alaska Division of Geological & Geophysical Surveys (DGGS), U.S. Bureau of Mines, and U.S. Bureau of Land Management serves as a data archive in support of Alaskan geologic and geochemical projects and contains data tables in several different formats describing historical and new quantitative and qualitative geochemical analyses. The analytical results were determined by 112 laboratory and field analytical methods on 396,343 rock, sediment, soil, mineral, heavy-mineral concentrate, and oxalic acid leachate samples. Most samples were collected by personnel of these agencies and analyzed in agency laboratories or, under contracts, in commercial analytical laboratories. These data represent analyses of samples collected as part of various agency programs and projects from 1938 through 2017. In addition, mineralogical data from 18,138 nonmagnetic heavy-mineral concentrate samples are included in this database. The AGDB3 includes historical geochemical data archived in the USGS National Geochemical Database (NGDB) and NURE National Uranium Resource Evaluation-Hydrogeochemical and Stream Sediment Reconnaissance databases, and in the DGGS Geochemistry database. Retrievals from these data-bases were used to generate most of the AGDB data set. These data were checked for accuracy regarding sample location, sample media type, and analytical methods used. In other words, the data of AGDB3 supersedes data in the AGDB and the AGDB2, but the background about the data in these two earlier versions are needed by users of the current AGDB3 to understand what has been done to amend, clean up, correct and format this data. Corrections were entered, resulting in a significantly improved Alaska geochemical dataset, the AGDB3. Data that were not previously in these databases because the data predate the earliest agency geochemical data-bases, or were once excluded for programmatic reasons, are included here in the AGDB3 and will be added to the NGDB and Alaska Geochemistry. The AGDB3 data provided here are the most accurate and complete to date and should be useful for a wide variety of geochemical studies. The AGDB3 data provided in the online version of the database may be updated or changed periodically.

Alaska↗

Data-resolution matrix and model-resolution matrix for Rayleigh-wave inversion using a damped least-squares method

Inversion of multimode surface-wave data is of increasing interest in the near-surface geophysics community. For a given near-surface geophysical problem, it is essential to understand how well the data, calculated according to a layered-earth model, might match the observed data. A data-resolution matrix is a function of the data kernel (determined by a geophysical model and a priori information applied to the problem), not the data. A data-resolution matrix of high-frequency (>2 Hz) Rayleigh-wave phase velocities, therefore, offers a quantitative tool for designing field surveys and predicting the match between calculated and observed data. We employed a data-resolution matrix to select data that would be well predicted and we find that there are advantages of incorporating higher modes in inversion. The resulting discussion using the data-resolution matrix provides insight into the process of inverting Rayleigh-wave phase velocities with higher-mode data to estimate S-wave velocity structure. Discussion also suggested that each near-surface geophysical target can only be resolved using Rayleigh-wave phase velocities within specific frequency ranges, and higher-mode data are normally more accurately predicted than fundamental-mode data because of restrictions on the data kernel for the inversion system. We used synthetic and real-world examples to demonstrate that selected data with the data-resolution matrix can provide better inversion results and to explain with the data-resolution matrix why incorporating higher-mode data in inversion can provide better results. We also calculated model-resolution matrices in these examples to show the potential of increasing model resolution with selected surface-wave data. ?? Birkhaueser 2008.

Pure and Applied Geophysics↗

Ten practical questions to improve data quality

High-quality rangeland data are critical to supporting adaptive management. However, concrete, cost-saving steps to ensure data quality are often poorly defined and understood. Data quality is more than data management. Ensuring data quality requires 1) clear communication among team members; 2) appropriate sample design; 3) training of data collectors, data managers, and data users; 4) observer and sensor calibration; and 5) active data management. Quality assurance and quality control are ongoing processes to help rangeland managers and scientists identify, prevent, and correct errors in past, current, and future monitoring data. We present 10 guiding data quality questions to help managers and scientists identify appropriate workflows to improve data quality by 1) describing the data ecosystem, 2) creating a data quality plan, 3) identifying roles and responsibilities, 4) building data collection and data management workflows, 5) training and calibrating data collectors, 6) detecting and correcting errors, and 7) describing sources of variability. Iteratively improving rangeland data quality is a key part of adaptive monitoring and rangeland data collection. All members of the rangeland community are invited to participate in ensuring rangeland data quality.

Rangelands↗

A trade-off between model resolution and variance with selected Rayleigh-wave data

Inversion of multimode surface-wave data is of increasing interest in the near-surface geophysics community. For a given near-surface geophysical problem, it is essential to understand how well the data, calculated according to a layered-earth model, might match the observed data. A data-resolution matrix is a function of the data kernel (determined by a geophysical model and a priori information applied to the problem), not the data. A data-resolution matrix of high-frequency (??? 2 Hz) Rayleigh-wave phase velocities, therefore, offers a quantitative tool for designing field surveys and predicting the match between calculated and observed data. First, we employed a data-resolution matrix to select data that would be well predicted and to explain advantages of incorporating higher modes in inversion. The resulting discussion using the data-resolution matrix provides insight into the process of inverting Rayleigh-wave phase velocities with higher mode data to estimate S-wave velocity structure. Discussion also suggested that each near-surface geophysical target can only be resolved using Rayleigh-wave phase velocities within specific frequency ranges, and higher mode data are normally more accurately predicted than fundamental mode data because of restrictions on the data kernel for the inversion system. Second, we obtained an optimal damping vector in a vicinity of an inverted model by the singular value decomposition of a trade-off function of model resolution and variance. In the end of the paper, we used a real-world example to demonstrate that selected data with the data-resolution matrix can provide better inversion results and to explain with the data-resolution matrix why incorporating higher mode data in inversion can provide better results. We also calculated model-resolution matrices of these examples to show the potential of increasing model resolution with selected surface-wave data. With the optimal damping vector, we can improve and assess an inverted model obtained by a damped least-square method.

Conference Paper↗

Geodatabase compilation of hydrogeologic, remote sensing, and water-budget-component data for the High Plains aquifer, 2011

The High Plains aquifer underlies almost 112 million acres in the central United States. It is one of the largest aquifers in the Nation in terms of annual groundwater withdrawals and provides drinking water for 2.3 million people. The High Plains aquifer has gained national and international attention as a highly stressed groundwater supply primarily because it has been appreciably depleted in some areas. The U.S. Geological Survey has an active program to monitor the changes in groundwater levels for the High Plains aquifer and has documented substantial water-level changes since predevelopment: the High Plains Groundwater Availability Study is part of a series of regional groundwater availability studies conducted to evaluate the availability and sustainability of major aquifers across the Nation. The goals of the regional groundwater studies are to quantify current groundwater resources in an aquifer system, evaluate how these resources have changed over time, and provide tools to better understand a systems response to future demands and environmental stresses. The purpose of this report is to present selected data developed and synthesized for the High Plains aquifer as part of the High Plains Groundwater Availability Study. The High Plains Groundwater Availability Study includes the development of a water-budget-component analysis for the High Plains completed in 2011 and development of a groundwater-flow model for the northern High Plains aquifer. Both of these tasks require large amounts of data about the High Plains aquifer. Data pertaining to the High Plains aquifer were collected, synthesized, and then organized into digital data containers called geodatabases. There are 8 geodatabases, 1 file geodatabase and 7 personal geodatabases, that have been grouped in three categories: hydrogeologic data, remote sensing data, and water-budget-component data. The hydrogeologic data pertaining to the northern High Plains aquifer is included in three separate geodatabases: (1) base data from a groundwater-flow model; (2) hydrogeology and hydraulic properties data; and (3) groundwater-flow model data to be used as calibration targets. The remote sensing data for this study were developed by the U. S. Geological Survey Earth Resources Observation and Science Center and include historical and predicted land-use/land-cover data and actual evapotranspiration data by using remotely sensed temperature data. The water-budget-component data contains selected raster data from maps in the “Selected Approaches to Estimate Water-Budget Components of the High Plains, 1940 Through 1949 and 2000 Through 2009” report completed in 2011 ( http://pubs.usgs.gov/sir/2011/5183/ ). Federal Geographic Data Committee compliant metadata were created for each spatial and tabular data layer in the geodatabases.

Colorado, Kansas, Nebraska, New Mexico, Oklahoma, ↗

Water resources data, Iowa, water year 2001, Volume 2. surface water--Missouri River basin, and ground water

The Water Resources Division of the U.S. Geological Survey, in cooperation with State, county, municipal, and other Federal agencies, obtains a large amount of data pertaining to the water resources of Iowa each water year. These data, accumulated during many water years, constitute a valuable data base for developing an improved understanding of the water resources of the State. To make this data readily available to interested parties outside of the Geological Survey, the data is published annually in this report series entitled “Water Resources Data - Iowa” as part of the National Water Data System. Water resources data for water year 2001 for Iowa consists of records of stage, discharge, and water quality of streams; stage and contents of lakes and reservoirs; and water levels and water quality of ground water. This report, in two volumes, contains stage or discharge records for 132 gaging stations; stage records for 9 lakes and reservoirs; water-quality records for 4 gaging stations; sediment records for 13 gaging stations; and water levels for 163 ground-water observation wells. Also included are peak-flow data for 92 crest-stage partial-record stations, water-quality data from 86 municipal wells, and precipitation data collected at 6 gaging stations and 2 precipitation sites. Additional water data were collected at various sites not included in the systematic data-collection program, and are published here as miscellaneous measurements and analyses. These data represent that part of the National Water Data System operated by the U.S. Geological Survey and cooperating local, State, and Federal agencies in Iowa. Records of discharge or stage of streams, and contents or stage of lakes and reservoirs were first published in a series of U.S. Geological Survey water-supply papers entitled “Surface Water Supply of the United States.” Through September 30, 1960, these water-supply papers were published in an annual series; during 1961-65 and 1966-70, they were published in 5- year series. Records of chemical quality, water temperatures, and suspended sediment were published from 1941 to 1970 in an annual series of water-supply papers entitled “Quality of Surface Waters of the United States.” Records of ground-water levels were published from 1935 to 1974 in a series of water-supply papers entitled “Ground-Water Levels in the United States.” Water-supply papers may be consulted in the libraries of the principal cities in the United States, or they may be purchased from Books and Open-File Reports Section, Federal Center, Box 25425, Denver, Colorado 80225. For water years 1961 through 1970, streamflow data were released by the Geological Survey in annual reports on a State-boundary basis. Water-quality records for water years 1964 through 1970 were similarly released either in separate reports or in conjunction with streamflow records. Beginning with the 1971 water year, water data for streamflow, water quality, and ground water is published in official U.S. Geological Survey reports on a State-boundary basis. These official reports carry an identification number consisting of the two-letter State postal abbreviation, the last two digits of the water year, and the volume number. For example, this report is identified as “U.S. Geological Survey Water-Data Report IA-01-1.” These water-data reports are for sale by the National Technical Information Service, U.S. Department of Commerce, Springfield, Virginia 22161.

Iowa↗

A multi-decade record of high-quality fCO2 data in version 3 of the Surface Ocean CO2 Atlas (SOCAT)

The Surface Ocean CO 2 Atlas (SOCAT) is a synthesis of quality-controlled f CO 2 (fugacity of carbon dioxide) values for the global surface oceans and coastal seas with regular updates. Version 3 of SOCAT has 14.7 million f CO 2 values from 3646 data sets covering the years 1957 to 2014. This latest version has an additional 4.6 million f CO 2 values relative to version 2 and extends the record from 2011 to 2014. Version 3 also significantly increases the data availability for 2005 to 2013. SOCAT has an average of approximately 1.2 million surface water f CO 2 values per year for the years 2006 to 2012. Quality and documentation of the data has improved. A new feature is the data set quality control (QC) flag of E for data from alternative sensors and platforms. The accuracy of surface water f CO 2 has been defined for all data set QC flags. Automated range checking has been carried out for all data sets during their upload into SOCAT. The upgrade of the interactive Data Set Viewer (previously known as the Cruise Data Viewer) allows better interrogation of the SOCAT data collection and rapid creation of high-quality figures for scientific presentations. Automated data upload has been launched for version 4 and will enable more frequent SOCAT releases in the future. High-profile scientific applications of SOCAT include quantification of the ocean sink for atmospheric carbon dioxide and its long-term variation, detection of ocean acidification, as well as evaluation of coupled-climate and ocean-only biogeochemical models. Users of SOCAT data products are urged to acknowledge the contribution of data providers, as stated in the SOCAT Fair Data Use Statement. This ESSD (Earth System Science Data) “living data” publication documents the methods and data sets used for the assembly of this new version of the SOCAT data collection and compares these with those used for earlier versions of the data collection (Pfeil et al., 2013; Sabine et al., 2013; Bakker et al., 2014).

Earth System Science Data↗

Spectral assessment of new ASTER SWIR surface reflectance data products for spectroscopic mapping of rocks and minerals

ASTER reflectance spectra from Cuprite, Nevada, and Mountain Pass, California, were compared to spectra of field samples and to ASTER-resampled AVIRIS reflectance data to determine spectral accuracy and spectroscopic mapping potential of two new ASTER SWIR reflectance datasets: RefL1b and AST_07XT. RefL1b is a new reflectance dataset produced for this study using ASTER Level 1B data, crosstalk correction, radiance correction factors, and concurrently acquired level 2 MODIS water vapor data. The AST_07XT data product, available from EDC and ERSDAC, incorporates crosstalk correction and non-concurrently acquired MODIS water vapor data for atmospheric correction. Spectral accuracy was determined using difference values which were compiled from ASTER band 5/6 and 9/8 ratios of AST_07XT or RefL1b data subtracted from similar ratios calculated for field sample and AVIRIS reflectance data. In addition, Spectral Analyst, a statistical program that utilizes a Spectral Feature Fitting algorithm, was used to quantitatively assess spectral accuracy of AST_07XT and RefL1b data.Spectral Analyst matched more minerals correctly and had higher scores for the RefL1b data than for AST_07XT data. The radiance correction factors used in the RefL1b data corrected a low band 5 reflectance anomaly observed in the AST_07XT and AST_07 data but also produced anomalously high band 5 reflectance in RefL1b spectra with strong band 5 absorption for minerals, such as alunite. Thus, the band 5 anomaly seen in the RefL1b data cannot be corrected using additional gain adjustments. In addition, the use of concurrent MODIS water vapor data in the atmospheric correction of the RefL1b data produced datasets that had lower band 9 reflectance anomalies than the AST_07XT data. Although assessment of spectral data suggests that RefL1b data are more consistent and spectrally more correct than AST_07XT data, the Spectral Analyst results indicate that spectral discrimination between some minerals, such as alunite and kaolinite, are still not possible unless additional spectral calibration using site specific spectral data are performed. ?? 2010.

Remote Sensing of Environment↗

Data sharing by scientists: Practices and perceptions

Background Scientific research in the 21st century is more data intensive and collaborative than in the past. It is important to study the data practices of researchers – data accessibility, discovery, re-use, preservation and, particularly, data sharing. Data sharing is a valuable part of the scientific method allowing for verification of results and extending research from prior results. Methodology/Principal Findings A total of 1329 scientists participated in this survey exploring current data sharing practices and perceptions of the barriers and enablers of data sharing. Scientists do not make their data electronically available to others for various reasons, including insufficient time and lack of funding. Most respondents are satisfied with their current processes for the initial and short-term parts of the data or research lifecycle (collecting their research data; searching for, describing or cataloging, analyzing, and short-term storage of their data) but are not satisfied with long-term data preservation. Many organizations do not provide support to their researchers for data management both in the short- and long-term. If certain conditions are met (such as formal citation and sharing reprints) respondents agree they are willing to share their data. There are also significant differences and approaches in data management practices based on primary funding agency, subject discipline, age, work focus, and world region. Conclusions/Significance Barriers to effective data sharing and preservation are deeply rooted in the practices and culture of the research process as well as the researchers themselves. New mandates for data management plans from NSF and other federal agencies and world-wide attention to the need to share and preserve data could lead to changes. Large scale programs, such as the NSF-sponsored DataNET (including projects like DataONE) will both bring attention and resources to the issue and make it easier for scientists to apply sound data management principles.

PLoS ONE↗