Search USGSSearch

SEARCH · Search USGS

Results for “Algorithms”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Two-stage approach to automatic detection with machine learning for improved surveillance of the invasive Cuban treefrog

The Cuban treefrog ( Osteopilus septentrionalis ), as an invasive species in the southern United States, presents a need for effective surveillance. Automated detection expedites processing of audio data for large-scale surveillance and monitoring programs. However, current available methods commonly used for anuran species have not been sufficient to detect Cuban treefrogs. Here, we present results from a two-stage method for automated detection that employs both cross-correlation template matching and secondary supervised learning classifiers. In the first stage, audio data are screened for initial detections using template matching, in which the detections contain both true and false positives. In the second stage, the false positives are screened out using classifier algorithms. We used this method to process 139,985 audio recordings, consisting of 596,046 total minutes, collected at 13 locations in Louisiana and Florida from 2014 to 2022. From the stage 1 template matching, we detected 83,191 Cuban treefrog signals across recordings. The stage 2 machine learning model was able to identify stage 1 false positive detections with a testing accuracy of 98.46% and a testing false positive rate of 1.116%. After pruning false positive detections, a total of 20,271 individual Cuban treefrog detections remained, distributed mainly across 3 sites in an area with known presence. Locations with presumed absence had an easily verifiable number of false positive detections ( n = 109 across all other sites). The two-stage methodology utilizing both template matching and machine learning algorithms can be integrated into wildlife surveillance or monitoring programs for species with distinctive, conserved calls as an effective way to achieve sensitive species detection with a low incidence of false positives.

Florida, Louisiana

Markov decision processes in natural resources management: observability and uncertainty

The breadth and complexity of stochastic decision processes in natural resources presents a challenge to analysts who need to understand and use these approaches. The objective of this paper is to describe a class of decision processes that are germane to natural resources conservation and management, namely Markov decision processes, and to discuss applications and computing algorithms under different conditions of observability and uncertainty. A number of important similarities are developed in the framing and evaluation of different decision processes, which can be useful in their applications in natural resources management. The challenges attendant to partial observability are highlighted, and possible approaches for dealing with it are discussed.

Ecological Modelling

Minimizing effects of methodological decisions on interpretation and prediction in species distribution studies: An example with background selection

Evaluating the conditions where a species can persist is an important question in ecology both to understand tolerances of organisms and to predict distributions across landscapes. Presence data combined with background or pseudo-absence locations are commonly used with species distribution modeling to develop these relationships. However, there is not a standard method to generate background or pseudo-absence locations, and method choice affects model outcomes. We evaluated combinations of both model algorithms (simple and complex generalized linear models, multivariate adaptive regression splines, Maxent, boosted regression trees, and random forest) and background methods (random, minimum convex polygon, and continuous and binary kernel density estimator (KDE)) to assess the sensitivity of model outcomes to choices made. We evaluated six questions related to model results, including five beyond the common comparison of model accuracy assessment metrics (biological interpretability of response curves, cross-validation robustness, independent data accuracy and robustness, and prediction consistency). For our case study with cheatgrass in the western US, random forest was least sensitive to background choice and the binary KDE method was least sensitive to model algorithm choice. While this outcome may not hold for other locations or species, the methods we used can be implemented to help determine appropriate methodologies for particular research questions.

Ecological Modelling

Augmented normalized difference water index for improved monitoring of surface water

We present a comprehensive critical review of well-established satellite remote sensing water indices and offer a novel, robust Augmented Normalized Difference Water Index (ANDWI). ANDWI employs an expanded set of spectral bands, RGB, NIR, and SWIR 1-2 , to maximize the contrast between water and non-water pixels. Further, we implement a dynamic thresholding method, the Otsu algorithm, to enhance ANDWI's performance. Applied to a variety of environmental conditions, ANDWI with Otsu-thresholding offered the highest overall accuracy (accuracy = 0.98, F1 = 0.98, and Kappa = 0.96) compared to other indices (NDWI, MNDWI, AWEI, WI). We also propose a novel cloud filtering algorithm that substantially increases the number of useable images compared to the conventional cloud-free composites (124% increased observations in the studied area) and resolves inappropriate masking of water bodies and hot sands as clouds by conventional methods. Finally, we develop a Google Earth Engine App to readily delineate 16-day surface water bodies across the globe.

Environmental Modeling and Software

NWTOPT — A hyperparameter optimization approach for selection of environmental model solver settings

Hyperparameter optimization approaches were applied to improve performance and accuracy of groundwater flow models. Freely available new software, NWTOPT, is described that uses Tree of Parzen Estimators (TPE) and Random Search algorithms to optimize MODFLOW-NWTs solver settings. We ran 3500 trials on a steady-state and transient model. To quantify the performance of candidate solver settings, we defined a loss function based on time elapsed and mass balance error of the MODFLOW-NWT forward run. Before optimization the steady- state model ran in ~12 min and the transient model ran in ~5 h with acceptable mass balance error (<1%). After optimization runtimes were reduced to ~2.7 min (steady state) and ~48 min (transient) with errors below 0.1%. In both cases TPE found hyperparameters that resulted in faster running and lower error models than those found by Random Search. The time to complete the optimization trials was also shorter with the TPE algorithm.

Environmental Modelling and Software

Combining process-based and data-driven approaches to forecast beach and dune change

Producing accurate hindcasts and forecasts with coupled models is challenging due to complex parameterizations that are difficult to ground in observational data. We present a calibration workflow that utilizes a series of machine learning algorithms paired with Windsurf, a coupled beach-dune model (Aeolis, the Coastal Dune Model, and XBeach), to produce hindcasts and forecasts of morphologic change along Bogue Banks, North Carolina. Neural networks paired with genetic algorithms allow us to fine tune calibration parameters for the hindcast, and then a long short-term memory neural network, trained on the hindcast, produces a 4-year forecast. We compare our hindcasts to observations from 2016 to 2017 and find they successfully reproduce observed modes of dune and beach change except for seaward growth of the dune face. We compare our forecasts to observations from 2016 to 2020 and find that they produce reasonably accurate predictions of dune change except when there are significant instances of erosion during the forecast period.

North Carolina

A generalized framework for inferring river bathymetry from image-derived velocity fields

Although established techniques for remote sensing of river bathymetry perform poorly in turbid water, image velocimetry can be effective under these conditions. This study describes a framework for mapping both of these attributes: Depths Inferred from Velocities Estimated by Remote Sensing, or DIVERS. The workflow involves linking image-derived velocities to depth via a flow resistance equation and invoking an optimization algorithm. We generalized an earlier formulation of DIVERS by: (1) using moving aircraft river velocimetry (MARV) to obtain a continuous, spatially extensive velocity field; (2) working within a channel-centered coordinate system; (3) allowing for local optimization of multiple parameters on a per-cross section basis; and (4) introducing a second objective function that can be used when discharge is not known. We also quantified the sensitivity of depth estimates to each parameter and input variable. MARV-based velocity estimates agreed closely with field measurements ( R 2 = 0.81 "> R 2 =0.81 ) and the use of DIVERS led to cross-sectional mean depths that were correlated with in situ observations ( R 2 = 0.75 "> R 2 =0.75 ). Errors in the input velocity field had the greatest impact on depth estimates, but the algorithm was not highly sensitive to initial parameter estimates when a known discharge was available to constrain the optimization. The DIVERS framework is predicated upon a number of simplifying assumptions — steady, uniform, one-dimensional flow and a strict, purely local proportionality between depth and velocity — that impose important limitations, but our results suggest that the approach can provide plausible, first-order estimates of river depths.

Alaska

Stochastic inversion of gravity, magnetic, tracer, lithology, and fault data for geologically realistic structural models: Patua Geothermal Field case study

Financial risk due to geological uncertainty is a major barrier for geothermal development. Production from a geothermal well depends on the unknown location of subsurface geological structures, such as faults that contain hydrothermal fluids. Traditionally, geoscientists collect many different datasets, interpret the datasets manually, and create a single model estimating faults' locations. This method, however, does not provide information about the uncertainty regarding the location of faults and often does not fully respect all observed datasets. Previous researchers investigated the use of stochastic inversion schemes for addressing geological uncertainty, but often at the expense of geologic realism. In this paper, we present algorithms and open-source code to stochastically invert five typical datasets for creating geologically realistic structural models. Using a case study with real data from the Patua Geothermal Field, we show that these inversion algorithms are successful in finding an ensemble of structural models that are geologically realistic and match the observed data sufficiently. Geoscientists can use this ensemble of models to optimize reservoir management decisions given structural uncertainty.

Nevada

When less is more: How increasing the complexity of machine learning strategies for geothermal energy assessments may not lead toward better estimates

Previous moderate- and high-temperature geothermal resource assessments of the western United States utilized data-driven methods and expert decisions to estimate resource favorability. Although expert decisions can add confidence to the modeling process by ensuring reasonable models are employed, expert decisions also introduce human and, thereby, model bias. This bias can present a source of error that reduces the predictive performance of the models and confidence in the resulting resource estimates. Our study aims to develop robust data-driven methods with the goals of reducing bias and improving predictive ability. We present and compare nine favorability maps for geothermal resources in the western United States using data from the U.S. Geological Survey's 2008 geothermal resource assessment. Two favorability maps are created using the expert decision-dependent methods from the 2008 assessment ( i.e., weight-of-evidence and logistic regression). With the same data, we then create six different favorability maps using logistic regression (without underlying expert decisions), XGBoost, and support-vector machines paired with two training strategies. The training strategies are customized to address the inherent challenges of applying machine learning to the geothermal training data, which have no negative examples and severe class imbalance. We also create another favorability map using an artificial neural network. We demonstrate that modern machine learning approaches can improve upon systems built with expert decisions. We also find that XGBoost, a non-linear algorithm, produces greater agreement with the 2008 results than linear logistic regression without expert decisions, because the expert decisions in the 2008 assessment rendered the otherwise linear approaches non-linear despite the fact that the 2008 assessment used only linear methods. The F1 scores for all approaches appear low (F1 score < 0.10), do not improve with increasing model complexity, and, therefore, indicate the fundamental limitations of the input features ( i.e., training data). Until improved feature data are incorporated into the assessment process, simple non-linear algorithms ( e.g., XGBoost) perform equally well or better than more complex methods ( e.g., artificial neural networks) and remain easier to interpret.

Geothermics

Optimizing selection of training and auxiliary data for operational land cover classification for the LCMAP initiative

The U.S. Geological Survey’s Land Change Monitoring, Assessment, and Projection (LCMAP) initiative is a new end-to-end capability to continuously track and characterize changes in land cover, use, and condition to better support research and applications relevant to resource management and environmental change. Among the LCMAP product suite are annual land cover maps that will be available to the public. This paper describes an approach to optimize the selection of training and auxiliary data for deriving the thematic land cover maps based on all available clear observations from Landsats 4–8. Training data were selected from map products of the U.S. Geological Survey’s Land Cover Trends project. The Random Forest classifier was applied for different classification scenarios based on the Continuous Change Detection and Classification (CCDC) algorithm. We found that extracting training data proportionally to the occurrence of land cover classes was superior to an equal distribution of training data per class, and suggest using a total of 20,000 training pixels to classify an area about the size of a Landsat scene. The problem of unbalanced training data was alleviated by extracting a minimum of 600 training pixels and a maximum of 8000 training pixels per class. We additionally explored removing outliers contained within the training data based on their spectral and spatial criteria, but observed no significant improvement in classification results. We also tested the importance of different types of auxiliary data that were available for the conterminous United States, including: (a) five variables used by the National Land Cover Database, (b) three variables from the cloud screening ‘‘Function of mask” (Fmask) statistics, and (c) two variables from the change detection results of CCDC. We found that auxiliary variables such as a Digital Elevation Model and its derivatives (aspect, position index, and slope), potential wetland index, water probability, snow probability, and cloud probability improved the accuracy of land cover classification. Compared to the original strategy of the CCDC algorithm (500 pixels per class), the use of the optimal strategy improved the classification accuracies substantially (15-percentage point increase in overall accuracy and 4-percentage point increase in minimum accuracy).

ISPRS Journal of Photogrammetry and Remote Sensing

Developing bare-earth digital elevation models from structure-from-motion data on barrier islands

Unoccupied aerial systems can collect aerial imagery that can be used to develop structure-from-motion products with a temporal resolution well-suited to monitoring dynamic barrier island environments. However, topographic data created using photogrammetric techniques such as structure-from-motion represent the surface elevation including the vegetation canopy . Additional processing is required for estimating bare-earth elevation, which is critical for understanding the underlying geomorphology of these islands. In this study, we used a vegetation and elevation survey to produce bare-earth digital elevation models from structure-from-motion-derived elevation products for two sites on Dauphin Island, Alabama (USA). One site was exposed to high wave energy and included a mix of beach, dune, and barrier flat habitats that were dominated by supratidal/upland herbaceous vegetation. The second site was exposed to low wave energy and was dominated by intertidal marsh. Aerial imagery was collected in late fall of 2018 and spring of 2019. We tested several machine learning algorithms for predicting and removing elevation bias for vegetated areas using predictors that included spectral indices from unoccupied aerial systems-based multispectral imagery and landscape position information (e.g., relative topography and distance from shore). Models were developed for each site and season. We also explored how well the model from one season generalized to data from a different season for the same site. For developing initial digital surface models, we found that utilizing a minimum bin algorithm, as opposed to interpolation, led to lower elevation bias. For bias removal, Gaussian process regression performed the best and led to a root mean square error for the bare-earth digital elevation models of around 0.10 m for the high energy site and 0.15 m for the low energy site. Compared to the digital surface models, the root mean square error for the bare-earth digital elevation models was reduced by at least 29 percent for the high energy site and 69 percent for the low energy site. For all models, common predictors included surface elevation, vegetation greenness, and distance from the shoreline. The models produced comparable results when trained using data from a different season. The error estimates for all analyses were within published elevation standards for lidar data for vegetated areas. With calibration, this approach could be portable to other areas or data, such as aerial lidar (conventional or unoccupied), to provide an efficient and repeatable framework for monitoring geomorphology or provide baseline elevations for predicting changes to these environments under future conditions.

Alabama

An automated compositing method for producing annual clear images from Landsat Collection 2 for annual NLCD production

Quality image input is fundamental to the quality of derived land cover products. Substantial time and effort are usually required to prepare images. Here, we present a novel and streamlined compositing algorithm that ingests Landsat Collection 2 Analysis Ready Data (ARD) and outputs cloud-free and gap-free composite imagery, which can be directly used for classification. This method leverages and improves the previous National Land Cover Database (NLCD) Virtual Median Value Point (VMVP) compositing method, the first part of the image preparation for NLCD 2019 operational production. The NLCD 2019 image preparation approach includes a second part, a residual cloud and cloud shadow detection and gap-filling method, to produce final cloud-free and gap-free composite imagery. The second part requires one clear reference image for each target year. Additional reference images are needed for producing reasonable observations for perennial ice/snow areas because Pixel QA (Quality Assessment) from ARD has difficulties differentiating ice/snow areas from clouds. Unlike the NLCD 2019 image preparation approach, our new compositing method, which is referred to as Automated VMVP (AVMVP), uses Landsat ARD as the only input and does not require reference images and extra steps. In this method, we developed new spectral filter criteria coupled with counts of clear observations using Pixel QA to identify potential cloud and cloud shadow observations on initially selected observations from the NLCD VMVP compositing algorithm. We also automate “gap-filling” using clear observations retrieved from a maximum of ±2 years around the target year when needed. Finally, a percentile-filtered compositing method was developed for the perennial ice/snow areas. All these steps are streamlined, pixel-based, and directly run on Landsat Collection 2 ARD. We have run successful tests on the conterminous United States (CONUS). Composite images derived from our innovative method were used to produce the CONUS Annual NLCD Collection 1 product suite that covers the period from 1985 to 2023.

conterminous United States

Ocean forecasting in terrain-following coordinates: Formulation and skill assessment of the Regional Ocean Modeling System

Systematic improvements in algorithmic design of regional ocean circulation models have led to significant enhancement in simulation ability across a wide range of space/time scales and marine system types. As an example, we briefly review the Regional Ocean Modeling System, a member of a general class of three-dimensional, free-surface, terrain-following numerical models. Noteworthy characteristics of the ROMS computational kernel include: consistent temporal averaging of the barotropic mode to guarantee both exact conservation and constancy preservation properties for tracers; redefined barotropic pressure-gradient terms to account for local variations in the density field; vertical interpolation performed using conservative parabolic splines; and higher-order, quasi-monotone advection algorithms. Examples of quantitative skill assessment are shown for a tidally driven estuary, an ice-covered high-latitude sea, a wind- and buoyancy-forced continental shelf, and a mid-latitude ocean basin. The combination of moderate-order spatial approximations, enhanced conservation properties, and quasi-monotone advection produces both more robust and accurate, and less diffusive, solutions than those produced in earlier terrain-following ocean models. Together with advanced methods of data assimilation and novel observing system technologies, these capabilities constitute the necessary ingredients for multi-purpose regional ocean prediction systems.

Journal of Computational Physics

A machine learning approach to identify barriers in stream networks demonstrates high prevalence of unmapped riverine dams

Restoring stream ecosystem integrity by removing unused or derelict dams has become a priority for watershed conservation globally. However, efforts to restore connectivity are constrained by the availability of accurate dam inventories which often overlook smaller unmapped riverine dams. Here we develop and test a machine learning approach to identify unmapped dams using a combination of publicly available topographic and geospatial habitat data. Specifically, we trained a random forest classification algorithm to identify unmapped dams using digitally engineered predictor variables and known dam sites for validation. We applied our algorithm to two subbasins in the Hudson River watershed, USA , and quantified connectivity impacts, as well as evaluated a range of predictor sets to examine tradeoffs between classification accuracy and model parameterization effort. The random forest classifier achieved high accuracy in predicting dam sites (true positive rate = 89%, false positive rate = 1.2%) using a subset of variables related to stream slope and presence of upstream lentic habitats. Unmapped dams were prevalent throughout the two test watersheds. In fact, existing dam inventories underestimated the true number of dams by ∼80–94%. Accounting for previously unmapped dams resulted in a 62–90% decrease in dendritic connectivity indices for migratory fishes. Unmapped dams may be pervasive and can dramatically bias stream connectivity information. However, we find that machine learning approaches can provide an accurate and scalable means of identifying unmapped dams that can guide efforts to develop accurate dam inventories, thereby informing and empowering efforts to better manage them.

New York

Evaluating the applicability of the generalized power-law rating curve model: With applications to paired discharge-stage data from Iceland, Sweden, and the United States

Hydrologic research and operations make extensive use of streamflow time series. In most applications, these time series are estimated from rating curves, which relate flow to some easy-to-measure surrogate, typically stage. The conventional stage-discharge rating takes the form of a segmented power law, with one segment for each hydrologic control at the stream gauge. However, these ratings are notoriously difficult to estimate with numerical methods, so that most are still developed manually. A few automated algorithms have emerged, but their use is sporadic, and their relative merits have not been rigorously assessed. One recently developed approach, the generalized power-law, avoids the segmenting problem by representing the power-law exponent as a Gaussian process. On the one hand, this representation is more flexible and easier to fit, but its flexibility might allow unrealistic solutions, so it needs to be tested under a range of conditions to assess its operational viability. This study evaluates the generalized power-law rating curve model by applying it to observations from 180 streams in Iceland, Sweden, and the United States. Overall, the model proved flexible and computationally robust, generating convincing rating curves across a range of geographic settings and was comparable to curves generated by a segmented rating model. Lastly, we propose a model-selection algorithm based on information theory to help identify the best rating curve model for a particular stream gauge.

Journal of Hydrology

Ultraviolet and visible remote sensing of volcanic gases

As magma rises in volcanic systems, volatile species exsolve from the silicate melt and are emitted as gases into the atmosphere. Measuring the magnitude and composition of gas emissions from volcanoes provides insights into processes occurring deep within the Earth and helps constrain the impact of volcanic degassing on atmospheric chemistry. Optical remote sensing techniques allow volcanic gas emissions to be characterized without the need to access hazardous areas near active volcanic vents. This paper reviews the state of the art in ultraviolet and visible volcanic gas remote sensing from the ground, air, and space. Special attention is given to discussing the physics of atmospheric radiative transfer on which these techniques are based. The functionality and limitations of different remote sensing instruments are examined, making clear that the ideal choice of instrumentation will depend on the volcanic system to which it is applied and the sought measurement parameters. Common algorithms for determining trace gas column densities, gas burdens, and volcanic emission rates from measurements of spectral radiance are outlined and compared, showing how some algorithms attempt to model the physics of the measurement while others maximize sensitivity. Several examples demonstrate how remote sensing measurements continue to advance our understanding of volcanic systems and their impact on the atmosphere. Finally, a few promising directions of inquiry are suggested that could lead to improvements in remote sensing instrumentation and analysis techniques. By combining spectroscopic and imaging techniques, improving our understanding of atmospheric radiative transfer, expanding the suite of target gases, and increasing the coverage and frequency of observations, we stand to significantly improve our ability to detect and quantify volcanic gas emissions and gain new insights into important Earth-system processes.

Journal of Volcanology and Geothermal Research

Monitoring cyanobacteria temporal trends in a hypereutrophic lake using remote sensing: From multispectral to hyperspectral

Cyanobacterial harmful algal blooms (cyanoHABs) and associated cyanotoxins are a concern for inland waters. Due to the extensive spatial coverage and frequent availability of satellite images, multispectral remote sensing tools demonstrate utility for monitoring these blooms. The next frontier for remote sensing of cyanoHABs in inland waters is hyperspectral data. Recent and upcoming hyperspectral satellite missions using narrow wavelength imaging spectrometers could have a major impact on advancing our ability to detect, quantify, and characterize cyanobacterial blooms. This study compares multispectral and hyperspectral remote sensing capabilities and processing tools for monitoring cyanoHAB dynamics. We evaluated the temporal trends of cyanoHABs in Clear Lake, California, a hypereutrophic lake with diverse cyanobacteria genera based on 38 sampling events over a five-year monitoring period (2019–2023). We validated the Sentinel-3 Ocean and Land Color Instrument (multispectral) Cyanobacteria Index algorithm for Clear Lake using in situ cyanobacteria measurements, which complemented our field-based evaluation of cyanobacteria trends in Clear Lake. We then demonstrate the advantages of hyperspectral data from both in situ spectroradiometer measurements and full-lake hyperspectral satellite images. We apply the Spectral Mixture Analysis for Surveillance of HABs (SMASH) workflow, a Multiple Endmember Spectral Mixture Analysis (MESMA) algorithm, to the hyperspectral images to assess the potential of satellite imaging spectrometer data to identify cyanobacteria genera – the first study to test this tool outside its original study sites. We developed a Clear Lake-specific cyanobacteria spectral library using our field spectroradiometer measurements to improve SMASH performance in Clear Lake, which supports the continued development of this tool.

California

Assessment of estuarine water-quality indicators using MODIS medium-resolution bands: initial results from Tampa Bay, FL

Using Tampa Bay, FL as an example, we explored the potential for using MODIS medium-resolution bands (250- and 500-m data at 469-, 555-, and 645-nm) for estuarine monitoring. Field surveys during 21–22 October 2003 showed that Tampa Bay has Case-II waters, in that for the salinity range of 24–32 psu, (a) chlorophyll concentration (11 to 23 mg m −3 ), (b) colored dissolved organic matter (CDOM) absorption coefficient at 400 nm (0.9 to 2.5 m −1 ), and (c) total suspended sediment concentration (TSS: 2 to 11 mg L −1 ) often do not co-vary. CDOM is the only constituent that showed a linear, inverse relationship with surface salinity, although the slope of the relationship changed with location within the bay. The MODIS medium-resolution bands, although designed for land use, are 4–5 times more sensitive than Landsat-7/ETM+ data and are comparable to or higher than those of CZCS. Several approaches were used to derive synoptic maps of water constituents from concurrent MODIS medium-resolution data. We found that application of various atmospheric-correction algorithms yielded no significant differences, due primarily to uncertainties in the sensor radiometric calibration and other sensor artifacts. However, where each scene could be groundtruthed, simple regressions between in situ observations of constituents and at-sensor radiances provided reasonable synoptic maps. We address the need for improvements of sensor calibration/characterization, atmospheric correction, and bio-optical algorithms to make operational and quantitative use of these medium-resolution bands.

Florida