Search USGS⌕ Search

SEARCH · Search USGS

Results for “Algorithms”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Coupling validation effort with in situ bioacoustic data improves estimating relative activity and occupancy for multiple species with cross-species misclassifications

The increasing complexity and pace of ecological change requires natural resource managers to consider entire species assemblages. Acoustic recording units (ARUs) require minimal cost and effort to deploy and inform relative activity, or encounter rates, for multiple species simultaneously. ARU-based surveys require post-processing of the recordings via software algorithms that assign a species label to each recording. The automated classification process can result in cross-species misidentifications that should be accounted for when employing statistical modelling for conservation decision-making. Using simulation and ARU-based detection counts from 17 bat species in British Columbia, Canada, we investigate three strategies for adjusting statistical inference for species misclassification: (a) ‘coupling’ ambiguous and unambiguous detections by validating a subset of survey events post-hoc, (b) using a calibration dataset on the software algorithm's (in)accuracy for species identification or (c) specifying informative Bayesian priors on classification probabilities. We explore the impact of different Bayesian prior specifications for the classification probabilities on posterior estimation. We then consider how the quantity of data validated post-hoc impacts model convergence and resulting inferences for bat species relative activity as related to nightly conditions and yearly site occupancy after accounting for site-level environmental variables. Coupled methods resulted in less bias and uncertainty when estimating relative activity and species classification probabilities relative to calibration approaches. We found that species that were difficult-to-detect and those that were often inaccurately identified by the software required more validation effort than more easily detected and/or identified species. Our results suggest that, when possible, acoustic surveys should rely on coupled validated detection information to account for false-positive detections, rather than uncoupled calibration datasets. However, if the assemblage of interest contains a large number of rarely detected or less prevalent species, an intractable amount of effort may be required, suggesting there are benefits to curating a calibration dataset that is representative of the observation process. Our findings provide insights into the practical challenges associated with statistical analyses of ARU data and possible analytical solutions to support reliable and cost-effective decision-making for wildlife conservation/management in the face of known sources of observation errors.

British Columbia↗

Simulated soundscapes and transfer learning boost the performance of acoustic classifiers under data scarcity

1. The biodiversity crisis necessitates spatially extensive methods to monitor multiple taxonomic groups for evidence of change in response to evolving environmental conditions. Programs that combine passive acoustic monitoring and machine learning are increasingly used to meet this need. These methods require large, annotated datasets, which are time-consuming and expensive to produce, creating potential barriers to adoption in data- and funding-poor regions. Recently released pre-trained avian acoustic classification models provide opportunities to reduce the need for manual labelling and accelerate the development of new acoustic classification algorithms through transfer learning. Transfer learning is a strategy for developing algorithms under data scarcity that uses pre-trained models from related tasks to adapt to new tasks. 2. Our primary objective was to develop a transfer learning strategy using the feature embeddings of a pre-trained avian classification model to train custom acoustic classification models in data-scarce contexts. We used three annotated avian acoustic datasets to test whether transfer learning and soundscape simulation-based data augmentation could substantially reduce the annotated training data necessary to develop performant custom acoustic classifiers. We also conducted a sensitivity analysis for hyperparameter choice and model architecture. We then assessed the generalizability of our strategy to increasingly novel non-avian classification tasks. 3. With as few as two training examples per class, our soundscape simulation data augmentation approach consistently yielded new classifiers with improved performance relative to the pre-trained classification model and transfer learning classifiers trained with other augmentation approaches. Performance increases were evident for three avian test datasets, including single-class and multi-label contexts. We observed that the relative performance among our data augmentation approaches varied for the avian datasets and nearly converged for one dataset when we included more training examples. 4. We demonstrate an efficient approach to developing new acoustic classifiers leveraging open-source sound repositories and pre-trained networks to reduce manual labelling. With very few examples, our soundscape simulation approach to data augmentation yielded classifiers with performance equivalent to those trained with many more examples, showing it is possible to reduce manual label-ling while still achieving high-performance classifiers and, in turn, expanding the potential for passive acoustic monitoring to address rising biodiversity monitoring needs.

Methods in Ecology and Evolution↗

Discrete-storm water-table fluctuation method to estimate episodic recharge.

We have developed a method to identify and quantify recharge episodes, along with their associated infiltration-related inputs, by a consistent, systematic procedure. Our algorithm partitions a time series of water levels into discrete recharge episodes and intervals of no episodic recharge. It correlates each recharge episode with a specific interval of rainfall, so storm characteristics such as intensity and duration can be associated with the amount of recharge that results. To be useful in humid climates, the algorithm evaluates the separability of events, so that those whose recharge cannot be associated with a single storm can be appropriately lumped together. Elements of this method that are subject to subjectivity in the application of hydrologic judgment are values of lag time, fluctuation tolerance, and master recession parameters. Because these are determined once for a given site, they do not contribute subjective influences affecting episode-to-episode comparisons. By centralizing the elements requiring scientific judgment, our method facilitates such comparisons by keeping the most subjective elements openly apparent, making it easy to maintain consistency. If applied to a period of data long enough to include recharge episodes with broadly diverse characteristics, the method has value for predicting how climatic alterations in the distribution of storm intensities and seasonal duration may affect recharge.

Groundwater↗

Multi‐constrained catchment scale optimization of groundwater abstraction using linear programming

Due to increasing water demands globally, freshwater ecosystems are under constant pressure. Groundwater resources, as the main source of accessible freshwater, are crucially important for irrigation worldwide. Over‐abstraction of groundwater leads to declines in groundwater levels; consequently, the groundwater inflow to streams decreases. The reduction in base flow and alteration of the stream flow regime can potentially have an adverse impact on groundwater‐dependent ecosystems. A spatially distributed, coupled groundwater‐surface water model can simulate the impacts of groundwater abstraction on aquatic ecosystems. A constrained optimization algorithm and a simulation model in combination can provide an objective tool for the water practitioner to evaluate the interplay between economic benefits of groundwater abstractions and requirements to environmental flow. In this study, a holistic catchment‐scale groundwater abstraction optimization framework has been developed that allows for a spatially explicit optimization of groundwater abstraction, while fulfilling a pre‐defined maximum allowed reduction of stream flow (base flow (Q95) or median flow (Q50)) as constraint criteria for 1484 stream locations across the catchment. A balanced K‐Means clustering method was implemented to reduce the computational burden of the optimization. The model parameters and observation uncertainties calculated based on Bayesian linear theory allow for a risk assessment on the optimized groundwater abstraction values. The results from different optimization scenarios indicated that using the linear programming optimization algorithm in conjunction with integrated models provides valuable information for guiding the water practitioners in designing an effective groundwater abstraction plan with the consideration of environmental flow criteria important for the ecological status of the entire system.

Groundwater↗

Evaluating lower computational burden approaches for calibration of large environmental models

Realistic environmental models used for decision making typically require a highly parameterized approach. Calibration of such models is computationally intensive because widely used parameter estimation approaches require individual forward runs for each parameter adjusted. These runs construct a parameter-to-observation sensitivity, or Jacobian, matrix used to develop candidate parameter upgrades. Parameter estimation algorithms are also commonly adversely affected by numerical noise in the calculated sensitivities within the Jacobian matrix, which can result in unnecessary parameter estimation iterations and less model-to-measurement fit. Ideally, approaches to reduce the computational burden of parameter estimation will also increase the signal-to-noise ratio related to observations influential to the parameter estimation even as the number of forward runs decrease. In this work a simultaneous increments, an iterative ensemble smoother (IES), and a randomized Jacobian approach were compared to a traditional approach that uses a full Jacobian matrix. All approaches were applied to the same model developed for decision making in the Mississippi Alluvial Plain, USA. Both the IES and randomized Jacobian approach achieved a desirable fit and similar parameter fields in many fewer forward runs than the traditional approach; in both cases the fit was obtained in fewer runs than the number of adjustable parameters. The simultaneous increments approach did not perform as well as the other methods due to inability to overcome suboptimal dropping of parameter sensitivities. This work indicates that use of highly efficient algorithms can greatly speed parameter estimation, which in turn increases calibration vetting and utility of realistic models used for decision making.

Mississippi Embayment regional aquifer system↗

Effects of auto-adaptive localization on a model calibration using ensemble methods

Simulations of the natural systems for environmental decision-making typically benefit from a highly parameterized approach (Hunt et al. 2007; Doherty and Hunt 2010), which enhances the flow of information contained in state observations to the parameters and improves application to decision support. However, parameter estimation (PE) with highly parameterized environmental models using traditional approaches (e.g., Doherty and Hunt 2010) is computationally intensive. Attempts at addressing the computational burden include improved computing approaches (e.g., Schreüder 2009; Hunt et al. 2010) and advances in algorithmic approaches (e.g., Tonkin and Doherty 2005; Welter et al. 2012, 2015). Recently, the iterative ensemble smoother (IES) approach (Chen and Oliver 2013; White 2018; White et al. 2020a) has greatly improved the efficiency of the PE calibration process compared to previous algorithms while concurrently providing nonlinear estimates of uncertainty (Hunt et al. 2021).

Groundwater↗

Waveform modelling using locked-mode synthetic and differential seismograms: application to determination of the structure of Mexico

We have developed algorithms for modelling seismic waveforms to constrain regional Earth structure. The seismogram is represented as a sum of locked-mode travelling waves in a layered medium. This representation is convenient as it allows us to model structures with slowly varying heterogeneity and to construct differential seismograms. Describes the techniques we have implemented that enable us to compute synthetic and differential seismograms in an efficient and stable manner. The computational methods are sufficiently rapid that many modes can be included and in some cases the entire seismogram may be modified. These algorithms are applied to model a set of seismograms of southern Mexican earthquakes recorded in northern Mexico. The frequency bandwidth of these data is centred at 0.067 Hz and we demonstrate that even at these relatively high frequencies, many features of the seismogram can be successfully modelled. Our results suggest that the structure within the recording array in northern Mexico is resolvably different from that to the south. We find that the average shear velocity of the lower lithosphere of southern Mexico is very low, approximately 4.3 km s-1. If the low-velocity region is confined to the Trans Mexican Volcanic Belt, the shear velocities between 20-80 km depth are approximately 3.3 km s-1. This may be correlated with partial melt and is consistent with the active volcanism and high heat flow found in the region. -Authors

Geophysical Journal International↗

Time-lapse three-dimensional inversion of complex conductivity data using an active time constrained (ATC) approach

Induced polarization (more precisely the magnitude and phase of impedance of the subsurface) is measured using a network of electrodes located at the ground surface or in boreholes. This method yields important information related to the distribution of permeability and contaminants in the shallow subsurface. We propose a new time-lapse 3-D modelling and inversion algorithm to image the evolution of complex conductivity over time. We discretize the subsurface using hexahedron cells. Each cell is assigned a complex resistivity or conductivity value. Using the finite-element approach, we model the in-phase and out-of-phase (quadrature) electrical potentials on the 3-D grid, which are then transformed into apparent complex resistivity. Inhomogeneous Dirichlet boundary conditions are used at the boundary of the domain. The calculation of the Jacobian matrix is based on the principles of reciprocity. The goal of time-lapse inversion is to determine the change in the complex resistivity of each cell of the spatial grid as a function of time. Each model along the time axis is called a ‘reference space model’. This approach can be simplified into an inverse problem looking for the optimum of several reference space models using the approximation that the material properties vary linearly in time between two subsequent reference models. Regularizations in both space domain and time domain reduce inversion artefacts and improve the stability of the inversion problem. In addition, the use of the time-lapse equations allows the simultaneous inversion of data obtained at different times in just one inversion step (4-D inversion). The advantages of this new inversion algorithm are demonstrated on synthetic time-lapse data resulting from the simulation of a salt tracer test in a heterogeneous random material described by an anisotropic semi-variogram.

Geophysical Journal International↗

W phase source inversion for moderate to large earthquakes (1990-2010)

Rapid characterization of the earthquake source and of its effects is a growing field of interest. Until recently, it still took several hours to determine the first-order attributes of a great earthquake (e.g. M w ≥ 7.5), even in a well-instrumented region. The main limiting factors were data saturation, the interference of different phases and the time duration and spatial extent of the source rupture. To accelerate centroid moment tensor (CMT) determinations, we have developed a source inversion algorithm based on modelling of the W phase, a very long period phase (100–1000 s) arriving at the same time as the P wave. The purpose of this work is to finely tune and validate the algorithm for large-to-moderate-sized earthquakes using three components of W phase ground motion at teleseismic distances. To that end, the point source parameters of all M w ≥ 6.5 earthquakes that occurred between 1990 and 2010 (815 events) are determined using Federation of Digital Seismograph Networks, Global Seismographic Network broad-band stations and STS1 global virtual networks of the Incorporated Research Institutions for Seismology Data Management Center. For each event, a preliminary magnitude obtained from W phase amplitudes is used to estimate the initial moment rate function half duration and to define the corner frequencies of the passband filter that will be applied to the waveforms. Starting from these initial parameters, the seismic moment tensor is calculated using a preliminary location as a first approximation of the centroid. A full CMT inversion is then conducted for centroid timing and location determination. Comparisons with Harvard and Global CMT solutions highlight the robustness of W phase CMT solutions at teleseismic distances. The differences in M w rarely exceed 0.2 and the source mechanisms are very similar to one another. Difficulties arise when a target earthquake is shortly (e.g. within 10 hr) preceded by another large earthquake, which disturbs the waveforms of the target event. To deal with such difficult situations, we remove the perturbation caused by earlier disturbing events by subtracting the corresponding synthetics from the data. The CMT parameters for the disturbed event can then be retrieved using the residual seismograms. We also explore the feasibility of obtaining source parameters of smaller earthquakes in the range 6.0 ≤M w < 6.5. Results suggest that the W phase inversion can be implemented reliably for the majority of earthquakes of M w = 6 or larger.

Geophysical Journal International↗

Code to generate random identifiers and select QA/QC samples

SAMPLID is a PC-based, FORTH AN-77 code which generates unique numbers for identification of samples, selection of QA/QC samples, and generation of labels. These procedures are tedious, but using a computer code such as SAMPLID can increase efficiency and reduce or eliminate errors and bias. The algorithm, used in SAMPLID, for generation of pseudorandom numbers is free of statistical flaws present in commonly available algorithms.

Groundwater↗

Mapping river bathymetry with a small footprint green LiDAR: Applications and challenges

Airborne bathymetric Light Detection And Ranging (LiDAR) systems designed for coastal and marine surveys are increasingly sought after for high-resolution mapping of fluvial systems. To evaluate the potential utility of bathymetric LiDAR for applications of this kind, we compared detailed surveys collected using wading and sonar techniques with measurements from the United States Geological Survey’s hybrid topographic⁄ bathymetric Experimental Advanced Airborne Research LiDAR (EAARL). These comparisons, based upon data collected from the Trinity and Klamath Rivers, California, and the Colorado River, Colorado, demonstrated that environmental conditions and postprocessing algorithms can influence the accuracy and utility of these surveys and must be given consideration. These factors can lead to mapping errors that can have a direct bearing on derivative analyses such as hydraulic modeling and habitat assessment. We discuss the water and substrate characteristics of the sites, compare the conventional and remotely sensed river-bed topographies, and investigate the laser waveforms reflected from submerged targets to provide an evaluation as to the suitability and accuracy of the EAARL system and associated processing algorithms for riverine mapping applications.

Journal of the American Water Resources Associatio↗

Modeling global Hammond landform regions from 250-m elevation data

In 1964, E.H. Hammond proposed criteria for classifying and mapping physiographic regions of the United States. Hammond produced a map entitled “Classes of Land Surface Form in the Forty-Eight States, USA”, which is regarded as a pioneering and rigorous treatment of regional physiography. Several researchers automated Hammond?s model in GIS. However, these were local or regional in application, and resulted in inadequate characterization of tablelands. We used a global 250 m DEM to produce a new characterization of global Hammond landform regions. The improved algorithm we developed for the regional landform modeling: (1) incorporated a profile parameter for the delineation of tablelands; (2) accommodated negative elevation data values; (3) allowed neighborhood analysis window (NAW) size to vary between parameters; (4) more accurately bounded plains regions; and (5) mapped landform regions as opposed to discrete landform features. The new global Hammond landform regions product builds on an existing global Hammond landform features product developed by the U.S. Geological Survey, which, while globally comprehensive, did not include tablelands, used a fixed NAW size, and essentially classified pixels rather than regions. Our algorithm also permits the disaggregation of “mixed” Hammond types (e.g. plains with high mountains) into their component parts.

Transactions in GIS↗

Satellite remote sensing of river discharge: A framework for assessing the accuracy of discharge estimates made from satellite remote sensing observations

This research presents an evaluation of the accuracy and uncertainty of estimates of river discharge made using satellite observed data sources as input to a modified form of Manning’s equation. Conventional U.S. Geological Survey (USGS) streamflow gaging station data and in-situ measurements of width, depth, height, slope, discharge, and velocity from 30 USGS gage sites were used as ground-truth to assess accuracy. This study explores accuracy in relation to the amount of ground truth information available, the number of calibration points available, and the accuracy of the input data. This research indicates that remotely sensed discharge estimates associated with the modified Manning equation may be expected to have an uncertainty in range of 10% overall given a sufficient number of calibration points. The uncertainty associated with the modified Manning algorithm increased markedly for depths <3 meters (m) and for discharges <1000 cubic meters per second (m 3 / s) for many rivers after calibration. Rivers that exhibit (1) a wide range of flow conditions, (2) a significant number of dams in the watershed and along the channel, and (3) a high baseflow index are more likely to have relatively large errors overall and particularly at the low end of the streamflow range. Uncertainty in remotely sensed measurements of water-surface elevation (WSE) and width in the expected range (WSE, + / − 10 cm; Width, + / − 15 m) introduces uncertainty in the discharge estimates on the order of 10% and is greatest at the low end of discharge as rivers get shallower and narrower. As WSE and width measurement uncertainty increases, discharge uncertainty increases accordingly. In general, the observation errors are greater than the errors associated with the algorithm for a well-calibrated model (e.g., 20 calibration points).

Journal of Applied Remote Sensing↗

Lava lake thermal pattern classification using self organizing maps and relationships to eruption processes at Kilauea Volcano, Hawaii

Kīlauea Volcano’s active summit lava lake poses hazards to downwind residents and over 1.6 million Hawai‘i Volcanoes National Park visitors each year. The lava lake surface is dynamic; crustal plates separated by incandescent cracks move across the lake as magma circulates below. We hypothesize that these dynamic thermal patterns are related to changes in other volcanic processes, such that sequences of thermal images may provide information about eruption parameters that are sometimes difficult to measure. The ability to learn about current gas emissions and seismic activity from a remote thermal time-lapse camera would be beneficial when conditions are too hazardous for field measurements. We apply a machine learning algorithm called self-organizing maps (SOM) to thermal infrared time-lapse images of the lava lake collected hourly over 23 April – 21 October 2013 (n=4354). The SOM algorithm can take thousands of seemingly different images, each representing the spatial distribution of relative temperature across the lava lake surface, and group them into clusters based on their similarities. We then relate the resulting clusters to sulfur dioxide emissions and seismic tremor to characterize ties between the SOM classification and different emplacement conditions. The SOM classification results are highly sensitive to the normalization method applied to the input images. The standard pixel-by-pixel normalization method yields a cluster of images defined by the highest observed SO2 emission levels, elevated surface temperatures, and a high proportion of cracks between crustal plates. When lava lake surface patterns are isolated by minimizing the effect of temperature variation between images, relationships with seismic tremor activity emerge, revealing an “intense spatter” cluster, characterized by unstable, broken-up crustal plate patterns on the lava lake surface. This proof of concept study provides a basis for extending the SOM classification method to hazard forecasting and real-time volcanic monitoring applications, as well as comparative studies at other lava lakes.

Hawaii↗

Harvest–release decisions in recreational fisheries

Most fishery regulations aim to control angler harvest. Yet, we lack a basic understanding of what actually determines the angler’s decision to harvest or release fish caught. We used XGBoost, a machine learning algorithm, to develop a predictive angler harvest–release model by taking advantage of an extensive recreational fishery data set (24 water bodies, 9 years, and 193 523 fish). We were able to successfully predict the harvest–release outcome for 99% of fish caught in the training data set and 96% of fish caught in the test data set. Unsuccessful predictions were mostly attributed to predicting harvest of fish that were released. Fish length was the most essential feature examined for predicting angler harvest. Other important predictive harvest–release features included the number of individuals of the same species caught, geographic location of an angler’s residence, distance traveled, and time spent fishing. The XGBoost algorithm was able to effectively predict the harvest–release decision and revealed hidden and intricate relationships that are often unaccounted for with classical analysis techniques. Exposing and accounting for these angler–fish intricacies is critical for fisheries conservation and management.

Nebraska↗

Efficiently optimizing for dendritic connectivity on tree-structured networks in a multi-objective framework

We provide an exact and approximation algorithm based on Dynamic Programming and an approximation algorithm based on Mixed Integer Programming for optimizing for the so-called dendritic connectivity on tree-structured networks in a multi-objective setting. Dendritic connectivity describes the degree of connectedness of a network. We consider different variants of dendritic connectivity to capture both network connectivity with respect to long and short-to-middle distances. Our work is motivated by a problem in computational sustainability concerning the evaluation of trade-offs in ecosystem services due to the proliferation of hydropower dams throughout the Amazon basin. In particular, we consider trade-offs between energy production and river connectivity. River fragmentation can dramatically affect fish migrations and other ecosystem services, such as navigation and transportation. In the context of river networks, different variants of dendritic connectivity are important to characterize the movements of different fish species and human populations. Our approaches are general and can be applied to optimizing for dendritic connectivity for a variety of multi-objective problems on tree-structured networks.

Conference Paper↗

Risk implications of Poisson assumptions and declustering inferred from a fully time-dependent earthquake forecast

We use the Third Uniform California Earthquake Rupture Forecast Epidemic Type Aftershock Sequence model, which is fully time-dependent in terms of including spatiotemporal clustering, to evaluate the effects of the Poisson assumption and declustering algorithms on statewide loss exceedance curves. The model is simulation based, meaning it produces synthetic catalogs that exhibit realistic behavior with respect to aftershocks and multi-fault earthquakes. A Poisson version of the model was constructed by randomizing event times, and the influence of two declustering algorithms was examined as well. We demonstrate that the probability of one-or-more loss exceedances (occurrence exceedance probability) is greater for the Poisson model because it has fewer seismically quiet time windows. The discrepancy between dollar loss estimates with a given exceedance probability is up to a factor of 32% but varies depending on the loss threshold (the x-axis value) and the forecast duration (we examined a range between 24 h and 50 years, with the discrepancy for the latter being negligible). We discuss how the one-or-more loss exceedance metric is questionable because it ignores all but the maximum loss experienced in each timeframe. An alternative metric based on total aggregate loss in each time window (aggregate exceedance probability) was therefore also examined, for which the Poisson model again implies higher risk at intermediate losses but lower risk at higher losses (because large, triggered events now contribute to total aggregate losses for the fully time-dependent model). We also argue that declustering is not a scientifically justifiable way to deal with full time dependence, in agreement with a chorus from other recent studies. It is difficult to draw generally applicable conclusions from our study, in part because application specific details will likely be important, but our results highlight how full time dependence can be reckoned with once authoritative forecast models are made available.

California↗

Can we estimate total magnetization directions from aeromagnetic data using Helbig's integrals?

An algorithm that implements Helbig’s (1963) integrals for estimating the vector components ( m x , m y , m z ) of the magnetic dipole moment from the first order moments of the vector magnetic field components (Δ X , Δ Y , Δ Z ) is tested on real and synthetic data. After a grid of total field aeromagnetic data is converted to vector component grids using Fourier filtering, Helbig’s infinite integrals are evaluated as finite integrals in small moving windows using a quadrature algorithm based on the 2-D trapezoidal rule. Prior to integration, best-fit planar surfaces must be removed from the component data within the data windows in order to make the results independent of the coordinate system origin. Two different approaches are described for interpreting the results of the integration. In the “direct” method, results from pairs of different window sizes are compared to identify grid nodes where the angular difference between solutions is small. These solutions provide valid estimates of total magnetization directions for compact sources such as spheres or dipoles, but not for horizontally elongated or 2-D sources. In the “indirect” method, which is more forgiving of source geometry, results of the quadrature analysis are scanned for solutions that are parallel to a specified total magnetization direction.

Earth, Planets and Space↗