A computer-based library reference system
Explore the source record for details and available documents.
SEARCH · Search USGS
Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Since computers became widely available in the early 1960s, seismologists have been using them for data acquisition, processing, and analysis, as well as theoretical computation and modeling. For example, the book by Doornbos (1988) contains a collection of seismological algorithms with the corresponding computer programs available on tape or disk from the World Data Center A for Solid Earth Geophysics. The introduction of personal computers in the early 1980s further revolutionized the use of computers for scientific research. Instead of expensive mainframe computers that required a large staff to operate, inexpensive personal computers allowed creative applications to be implemented by individuals with a shoestring budget.
Trace contaminants in water, including metals and organics, often are measured at sufficiently low concentrations to be reported only as values below the instrument detection limit. Interpretation of these "less thans" is complicated when multiple detection limits occur. Statistical methods for multiply censored, or multiple-detection limit, datasets have been developed for medical and industrial statistics, and can be employed to estimate summary statistics or model the distributions of trace-level environmental data. We describe S-language-based software tools that perform robust linear regression on order statistics (ROS). The ROS method has been evaluated as one of the most reliable procedures for developing summary statistics of multiply censored data. It is applicable to any dataset that has 0 to 80% of its values censored. These tools are a part of a software library, or add-on package, for the R environment for statistical computing. This library can be used to generate ROS models and associated summary statistics, plot modeled distributions, and predict exceedance probabilities of water-quality standards. ?? 2005 Elsevier Ltd. All rights reserved.
The U.S. Geological Survey (USGS) Library, where the author serves as the digital services librarian, is increasingly challenged to make it easier for users to find information from many heterogeneous information sources. Information is scattered throughout different software applications (i.e., library catalog, federated search engine, link resolver, and vendor websites), and each specializes in one thing. How could the library integrate the functionalities of one application with another and provide a single point of entry for users to search across? To improve the user experience, the library launched an effort to integrate the federated search engine into the library's intranet website. The result is a simple search box that leverages the federated search engine's built-in application programming interfaces (APIs). In this article, the author describes how this project demonstrated the power of APIs and their potential to be used by other enterprise search portals inside or outside of the library.
This report presents precision and accuracy data for volatile organic compounds (VOCs) in the nanogram-per-liter range, including aromatic hydrocarbons, reformulated fuel components, and halogenated hydrocarbons using purge and trap capillary-column gas chromatography/mass spectrometry. One-hundred-four VOCs were initially tested. Of these, 86 are suitable for determination by this method. Selected data are provided for the 18 VOCs that were not included. This method also allows for the reporting of semiquantitative results for tentatively identified VOCs not included in the list of method compounds. Method detection limits, method performance data, preservation study results, and blank results are presented. The authors describe a procedure for reporting low-concentration detections at less than the reporting limit. The nondetection value (NDV) is introduced as a statistically defined reporting limit designed to limit false positives and false negatives to less than 1 percent. Nondetections of method compounds are reported as ?less than NDV.? Positive detections measured at less than NDV are reported as estimated concentrations to alert the data user to decreased confidence in accurate quantitation. Instructions are provided for analysts to report data at less than the reporting limits. This method can support the use of either method reporting limits that censor detections at lower concentrations or the use of NDVs as reporting limits. The data-reporting strategy for providing analytical results at less than the reporting limit is a result of the increased need to identify the presence or absence of environmental contaminants in water samples at increasingly lower concentrations. Long-term method detection limits (LTMDLs) for 86 selected compounds range from 0.013 to 2.452 micrograms per liter (?g/L) and differ from standard method detection limits (MDLs) in that the LTMDLs include the long-term variance of multiple instruments, multiple operators, and multiple calibrations over a longer time. For these reasons, LTMDLs are expected to be slightly higher than standard MDLs. Recoveries for all of the VOCs tested ranged from 36 (tert-butyl formate) to 155 percent (pentachlorobenzene). The majority of the compounds ranged from 85 to 115 percent recovery and had less than 5 percent relative standard deviation for concentrations spiked between 1 to 500 ?g/L in volatile blank-, surface-, and ground-water samples. Recoveries of 60 set spikes at low concentrations ranged from 70 to 114 percent (1,2,3- trimethylbenzene and acetone). Recovery data were collected over 6 months with multiple instruments, operators, and calibrations. In this method, volatile organic compounds are extracted from a water sample by actively purging with helium. The VOCs are collected onto a sorbent trap, thermally desorbed, separated by a Megabore gas chromatographic capillary column, and finally determined by a full-scan quadrupole mass spectrometer. Compound identification is confirmed by the gas chromatographic retention time and by the resultant mass spectrum, typically identified by three unique ions. An unknown compound detected in a sample can be tentatively identified by comparing the unknown mass spectrum to reference spectra in the mass-spectra computer-data system library compiled by the National Institute of Standards and Technology.
Knowledge of predator diets provides essential insights into their ecology, yet diet estimation is challenging and remains an active area of research. Quantitative fatty acid signature analysis (QFASA) is a popular method of estimating diet composition that continues to be investigated and extended. However, software to implement QFASA has only recently become publicly available. I summarize a new R package, qfasar , for diet estimation using QFASA methods. The package also provides functionality to evaluate and potentially improve the performance of a library of prey signature data, compute goodness-of-fit diagnostics, and support simulation-based research. Several procedures in the package have not previously been published. qfasar makes traditional and recently published QFASA diet estimation methods accessible to ecologists for the first time. Use of the package is illustrated with signature data from Chukchi Sea polar bears and potential prey species.
Although many of the algorithms now considered to be machine learning algorithms (MLAs) have existed for nearly a century (e.g., Rosenblatt 1958 ), interest in MLAs has recently increased exponentially for solving data-driven problems across a variety of fields due to the expanded availability of large, complex datasets that may be difficult to interrogate using other methods, increases in computing power, and a growing library of easily implemented machine learning tools. While MLAs are often similar to statistical methods, there are key differences in the approach to problem solving. Namely, statistical methods are more concerned with generating informative models from “long” data (i.e., many more observations than explanatory variables), whereas MLAs are typically concerned with generating accurate predictions from “wide” data (i.e., a large number of variables with relatively fewer observations, Bzdok et al. 2018 ). In hydrogeologic studies, such wide datasets may be available from boreholes, where various types of geophysical, geochemical, and lithological information may exist. Borehole datasets are therefore a tempting target for MLAs to reveal hidden relations among gathered data and parameters of interest (e.g., contaminant concentration), and as a method of parameter reduction (e.g., reduce costs by collecting fewer datasets).
We have assembled a library of spectra measured with laboratory, field, and airborne spectrometers. The instruments used cover wavelengths from the ultraviolet to the far infrared (0.2 to 200 microns [μm]). Laboratory samples of specific minerals, plants, chemical compounds, and manmade materials were measured. In many cases, samples were purified, so that unique spectral features of a material can be related to its chemical structure. These spectro-chemical links are important for interpreting remotely sensed data collected in the field or from an aircraft or spacecraft. This library also contains physically constructed as well as mathematically computed mixtures. Four different spectrometer types were used to measure spectra in the library: (1) Beckman™ 5270 covering the spectral range 0.2 to 3 µm, (2) standard, high resolution (hi-res), and high-resolution Next Generation (hi-resNG) models of Analytical Spectral Devices (ASD) field portable spectrometers covering the range from 0.35 to 2.5 µm, (3) Nicolet™ Fourier Transform Infra-Red (FTIR) interferometer spectrometers covering the range from about 1.12 to 216 µm, and (4) the NASA Airborne Visible/Infra-Red Imaging Spectrometer AVIRIS, covering the range 0.37 to 2.5 µm. Measurements of rocks, soils, and natural mixtures of minerals were made in laboratory and field settings. Spectra of plant components and vegetation plots, comprising many plant types and species with varying backgrounds, are also in this library. Measurements by airborne spectrometers are included for forested vegetation plots, in which the trees are too tall for measurement by a field spectrometer. This report describes the instruments used, the organization of materials into chapters, metadata descriptions of spectra and samples, and possible artifacts in the spectral measurements. To facilitate greater application of the spectra, the library has also been convolved to selected spectrometer and imaging spectrometers sampling and bandpasses, and resampled to selected broadband multispectral sensors. The native file format of the library is the SPECtrum Processing Routines (SPECPR) data format. This report describes how to access freely available software to read the SPECPR format. To facilitate broader access to the library, we produced generic formats of the spectra and metadata in text files. The library is provided on digital media and online at https://speclab.cr.usgs.gov/spectral-lib.html . A long-term archive of these data are stored on the USGS ScienceBase data server ( https://dx.doi.org/10.5066/F7RR1WDJ ).
In this work we describe the AutoCNet library, written in Python, to support the application of Computer Vision techniques for n-image correspondence identication in remotely sensed planetary images and subsequent bundle adjustment. The library is designed to support exploratory data analysis, algorithm and processing pipeline development, and application at scale in High Performance Computing (HPC) environments for processing large data sets and generating foundational data products. We also present a brief case study illustrating high level usage for the Apollo 15 Metric camera.
Introduction We have assembled a digital reflectance spectral library that covers the wavelength range from the ultraviolet to far infrared along with sample documentation. The library includes samples of minerals, rocks, soils, physically constructed as well as mathematically computed mixtures, plants, vegetation communities, microorganisms, and man-made materials. The samples and spectra collected were assembled for the purpose of using spectral features for the remote detection of these and similar materials. Analysis of spectroscopic data from laboratory, aircraft, and spacecraft instrumentation requires a knowledge base. The spectral library discussed here forms a knowledge base for the spectroscopy of minerals and related materials of importance to a variety of research programs being conducted at the U.S. Geological Survey. Much of this library grew out of the need for spectra to support imaging spectroscopy studies of the Earth and planets. Imaging spectrometers, such as the National Aeronautics and Space Administration (NASA) Airborne Visible/Infra Red Imaging Spectrometer (AVIRIS) or the NASA Cassini Visual and Infrared Mapping Spectrometer (VIMS) which is currently orbiting Saturn, have narrow bandwidths in many contiguous spectral channels that permit accurate definition of absorption features in spectra from a variety of materials. Identification of materials from such data requires a comprehensive spectral library of minerals, vegetation, man-made materials, and other subjects in the scene. Our research involves the use of the spectral library to identify the components in a spectrum of an unknown. Therefore, the quality of the library must be very good. However, the quality required in a spectral library to successfully perform an investigation depends on the scientific questions to be answered and the type of algorithms to be used. For example, to map a mineral using imaging spectroscopy and the mapping algorithm of Clark and others (1990a, 2003b), one simply needs a diagnostic absorption band. The mapping system uses continuum-removed reference spectral features fitted to features in observed spectra. Spectral features for such algorithms can be obtained from a spectrum of a sample containing large amounts of contaminants, including those that add other spectral features, as long as the shape of the diagnostic feature of interest is not modified. If, however, the data are needed for radiative transfer models to derive mineral abundances from reflectance spectra, then completely uncontaminated spectra are required. This library contains spectra that span a range of quality, with purity indicators to flag spectra for (or against) particular uses. Acquiring spectral measurements and performing sample characterizations for this library has taken about 15 person-years of effort. Software to manage the library and provide scientific analysis capability is provided (Clark, 1980, 1993). A personal computer (PC) reader for the library is also available (Livo and others, 1993). The program reads specpr binary files (Clark, 1980, 1993) and plots spectra. Another program that reads the specpr format is written in IDL (Kokaly, 2005). In our view, an ideal spectral library consists of samples covering a very wide range of materials, has large wavelength range with very high precision, and has enough sample analyses and documentation to establish the quality of the spectra. Time and available resources limit what can be achieved. Ideally, for each mineral, the sample analysis would include X-ray diffraction (XRD), electron microprobe (EM) or X-ray fluorescence (XRF), and petrographic microscopic analyses. For some minerals, such as iron oxides, additional analyses such as Mossbauer would be helpful. We have found that to make the basic spectral measurements, provide XRD, EM or XRF analyses, and microscopic analyses, document the results, and complete an entry of one spectral library sample, all takes about
GKS-PC is a Graphical Kernel System for the family of microcomputers compatible with the IBM-PC/MS-DOS standard. This library of functions allows an application program to control a computer graphics display conforming to a variety of graphics display standards. GKS-PC contains a subset of the GKS system of graphics primitive functions. GKS is an ANSI standardized computer graphics software development environment. General descriptions of GKS are in the references (American National Standards Committee X3, 1984) and (Hopgood et.al., 1983) listed on page 52. GKS-PC is not a comprehensive package for producing commercial grade graphics software, and it is not a general animation environment even though certain kinds of animation are possible with this software. The 5.25" diskette containing the GKS-PC object library files and installation program is available as Open-File Report 93-241-B.
This report documents a graphical display program for the U. S. Geological Survey finite-element groundwater flow and solute transport model. Graphic features of the program, SUTRA-PLOT (SUTRA-PLOT = saturated/unsaturated transport), include: (1) plots of the finite-element mesh, (2) velocity vector plots, (3) contour plots of pressure, solute concentration, temperature, or saturation, and (4) a finite-element interpolator for gridding data prior to contouring. SUTRA-PLOT is written in FORTRAN 77 on a PRIME 750 computer system, and requires Version 9.0 or higher of the DISSPLA graphics library. The program requires two input files: the SUTRA input data list and the SUTRA simulation output listing. The program is menu driven and specifications for individual types of plots are entered and may be edited interactively. Installation instruction, a source code listing, and a description of the computer code are given. Six examples of plotting applications are used to demonstrate various features of the plotting program. (Author 's abstract)
The genomics revolution has initiated a new era of population genetics where genome-wide data are frequently used to understand complex patterns of population structure and selection. However, the application of genomic tools to inform management and conservation has been somewhat rare outside a few well studied species. Fortunately, two recently developed approaches, amplicon sequencing and sequence capture, have the potential to significantly advance the field of conservation genomics. Here, amplicon sequencing refers to highly multiplexed PCR followed by high-throughput sequencing (e.g., GTseq), and sequence capture refers to using capture probes to isolate loci from reduced-representation libraries (e.g., Rapture). Both approaches allow sequencing of thousands of individuals at relatively low costs, do not require any specialized equipment for library preparation, and generate data that can be analyzed without sophisticated computational infrastructure. Here, we discuss the advantages and disadvantages of each method and provide a decision framework for geneticists who are looking to integrate these methods into their research programme. While it will always be important to consider the specifics of the biological question and system, we believe that amplicon sequencing is best suited for projects aiming to genotype <500 loci on many individuals (>1,500) or for species where continued monitoring is anticipated (e.g., long-term pedigrees). Sequence capture, on the other hand, is best applied to projects including fewer individuals or where >500 loci are required. Both of these techniques should smooth the transition from traditional genetic techniques to genomics, helping to usher in the conservation genomics era.
In the spring of 2004, the U.S. Geological Survey (USGS) Menlo Park Center Council commissioned an interdisciplinary working group to develop a forward-looking science strategy for the USGS Menlo Park Science Center in California (hereafter also referred to as "the Center"). The Center has been the flagship research center for the USGS in the western United States for more than 50 years, and the Council recognizes that science priorities must be the primary consideration guiding critical decisions made about the future evolution of the Center. In developing this strategy, the working group consulted widely within the USGS and with external clients and collaborators, so that most stakeholders had an opportunity to influence the science goals and operational objectives. The Science Goals are to: Natural Hazards: Conduct natural-hazard research and assessments critical to effective mitigation planning, short-term forecasting, and event response. Ecosystem Change: Develop a predictive understanding of ecosystem change that advances ecosystem restoration and adaptive management. Natural Resources: Advance the understanding of natural resources in a geologic, hydrologic, economic, environmental, and global context. Modeling Earth System Processes: Increase and improve capabilities for quantitative simulation, prediction, and assessment of Earth system processes. The strategy presents seven key Operational Objectives with specific actions to achieve the scientific goals. These Operational Objectives are to: Provide a hub for technology, laboratories, and library services to support science in the Western Region. Increase advanced computing capabilities and promote sharing of these resources. Enhance the intellectual diversity, vibrancy, and capacity of the work force through improved recruitment and retention. Strengthen client and collaborative relationships in the community at an institutional level. Expand monitoring capability by increasing density, sensitivity, and efficiency and reducing costs of instruments and networks. Encourage a breadth of scientific capabilities in Menlo Park to foster interdisciplinary science. Communicate USGS science to a diverse audience.
During the past several years, the sediment transport group in the Coastal and Marine Geology Program (CMGP) of the U. S. Geological Survey has made major revisions to its methodology of processing, analyzing, and maintaining the variety of oceanographic time-series data. First, CMGP completed the transition of the its oceanographic time-series database to a self-documenting NetCDF (Rew et al., 1997) data format. Second, CMGP’s oceanographic data variety and complexity have been greatly expanded from traditional 2-dimensional, single-point time-series measurements (e.g., Electro-magnetic current meters, transmissometers) to more advanced 3-dimensional and profiling time-series measurements due to many new acquisitions of modern instruments such as Acoustic Doppler Current Profiler (RDI, 1996), Acoustic Doppler Velocitimeter, Pulse-Coherence Acoustic Doppler Profiler (SonTek, 2001), Acoustic Bacscatter Sensor (Aquatec, 1001001001001001001). In order to accommodate the NetCDF format of data from the new instruments, a software package of processing, analyzing, and visualizing time-series oceanographic data was developed. It is named CMGTooL. The CMGTooL package contains two basic components: a user-friendly GUI for NetCDF file analysis, processing and manipulation; and a data analyzing program library. Most of the routines in the library are stand-alone programs suitable for batch processing. CMGTooL is written in MATLAB computing language (The Mathworks, 1997), therefore users must have MATLAB installed on their computer in order to use this software package. In addition, MATLAB’s Signal Processing Toolbox is also required by some CMGTooL’s routines. Like most MATLAB programs, all CMGTooL codes are compatible with different computing platforms including PC, MAC, and UNIX machines (Note: CMGTooL has been tested on different platforms that run MATLAB 5.2 (Release 10) or lower versions. Some of the commands related to MAC may not be compatible with later releases of MATLAB). The GUI and some of the library routines call low-level NetCDF file I/O, variable and attribute functions. These NetCDF exclusive functions are supported by a MATLAB toolbox named NetCDF, created by Dr. Charles Denham . This toolbox has to be installed in order to use the CMGTooL GUI. The CMGTooL GUI calls several routines that were initially developed by others. The authors would like to acknowledge the following scientists for their ideas and codes: Dr. Rich Signell (USGS), Dr. Chris Sherwood (USGS), and Dr. Bob Beardsley (WHOI). Many special terms that carry special meanings in either MATLAB or the NetCDF Toolbox are used in this manual. Users are encouraged to read the documents of MATLAB and NetCDF for references.
We have assembled a digital reflectance spectral library of spectra that covers wavelengths from the ultraviolet to near-infrared along with sample documentation. The library includes samples of minerals, rocks, soils, physically constructed as well as mathematically computed mixtures, vegetation, microorganisms, and man-made materials. The samples and spectra collected were assembled for the purpose of using spectral features for the remote detection of these and similar materials.
SSR_pipeline is a flexible set of programs designed to efficiently identify simple sequence repeats (e.g., microsatellites) from paired-end high-throughput Illumina DNA sequencing data. The program suite contains 3 analysis modules along with a fourth control module that can automate analyses of large volumes of data. The modules are used to 1) identify the subset of paired-end sequences that pass Illumina quality standards, 2) align paired-end reads into a single composite DNA sequence, and 3) identify sequences that possess microsatellites (both simple and compound) conforming to user-specified parameters. The microsatellite search algorithm is extremely efficient, and we have used it to identify repeats with motifs from 2 to 25bp in length. Each of the 3 analysis modules can also be used independently to provide greater flexibility or to work with FASTQ or FASTA files generated from other sequencing platforms (Roche 454, Ion Torrent, etc.). We demonstrate use of the program with data from the brine fly Ephydra packardi (Diptera: Ephydridae) and provide empirical timing benchmarks to illustrate program performance on a common desktop computer environment. We further show that the Illumina platform is capable of identifying large numbers of microsatellites, even when using unenriched sample libraries and a very small percentage of the sequencing capacity from a single DNA sequencing run. All modules from SSR_pipeline are implemented in the Python programming language and can therefore be used from nearly any computer operating system (Linux, Macintosh, and Windows).