Search USGSSearch

SEARCH · Search USGS

Results for “JGR Machine Learning and Computation”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,224 records · Page 2Linked to original sources

Surface variable‐based machine learning for scalable arsenic prediction in undersampled areas

In the United States, private wells are not federally regulated, and many households do not test for Arsenic (As). Chronic exposure is linked with multiple health outcomes, and risk can change sharply over short distances and with well depth. Coarse maps or sparse sampling often miss exceedances. Most existing models operate at ∼1 km resolution and use groundwater chemistry or detailed geologic logs, which limits their use in undersampled areas where improved guidance is most needed. We overcome these limitations by developing a machine learning model for Minnesota, USA, that predicts As exposure risk using only surficial variables from remote sensing and global data sets. Variables related to surface water hydrology and geomorphology are selected based on mechanistic links that control redox conditions and As mobilization. Local training was essential, and surficial geology variables that are more sensitive to local conditions were needed to maximize model accuracy. The resulting complete model was sufficiently sensitive to generate accurate and detailed risk maps and depth profiles of As concentrations above the 10 μg/L maximum contaminant level. Accuracy depended on local training data density. We identified a training data density of 0.07 wells/km 2 as a practical target for stable county-level performance. Maps of exceedance probabilities highlight priority areas for testing that are particularly important in rural communities that have received less sampling. These results support public health action by guiding where to install wells and where to test them, how much new sampling is needed, and where treatment outreach is most urgent.

Minnesota

HyFlood: A surrogate-model-based framework for compound coastal flooding

Compound coastal flooding is a major threat to low-lying coastal regions and is expected to intensify under future climate change projections. However, modeling the joint interaction of waves, storm surge, tides, and rainfall remains computationally demanding, limiting the development of fast and reliable forecast tools. Here we present HyFlood, a hybrid statistical-numerical downscaling framework capable of computing and mapping high-resolution compound flood hazards while substantially reducing the computational cost compared with fully process-based hydrodynamic modeling. HyFlood combines statistical sampling and selection algorithms with a cascade of reduced-complexity surrogate models that emulate nearshore wave transformation, surf-zone hydrodynamics, and coastal, fluvial, and pluvial flooding. The surrogate models employ machine-learning and regression algorithms applied to a low-dimensional representation of the flooding outputs, obtained through statistical dimensionality reduction. The framework is demonstrated in southern O'ahu, Hawai'i, a region exposed to elevated sea levels driven by tides, waves, and storm surge along with frequent precipitation-driven flash flooding. Validation of the surrogates against the physics-based model outputs demonstrates that HyFlood accurately reproduces daily maxima of spatially distributed flooding depths. This hybrid approach offers a scalable and efficient tool to better quantify how changes in flooding drivers translate into hazard and impact assessments, and to support compound-flood risk assessments and climate-change adaptation planning.

Hawaii

Forecasting water levels using the ConvLSTM algorithm in the Everglades, USA

Forecasting water levels in complex ecosystems like wetlands can support effective water resource management, ecological conservation, and understanding surface and groundwater hydrology. Predictive models can be used to simulate the complex interactions among natural processes, hydrometeorological factors, and human activities. The Greater Everglades in the USA is a well-known example of an ecosystem where complexity has motivated adoption of machine learning algorithms in water level prediction studies. This paper aims to contribute to extending existing machine learning algorithms by integrating spatiotemporal data with deep-learning algorithms in the forecasting process. In this study, a deep-learning model is developed to predict water levels on a regional scale, covering a large area of approximately 9,138 square kilometers in the Everglades ecosystem. This model has the architecture of Convolutional Long Short-Term Memory which can deal with spatiotemporal data by capturing both spatial and temporal dependencies in the training data. The forecasting capabilities of this model (referred to as the global model) are assessed by comparing the global model to two Artificial Neural Networks developed at two different gaging stations, referred to here as local models. One local model is developed at a gaging station directly influenced by nearby water control structures, whereas the other is developed at a gaging station located farther away from these structures. By leveraging data from the Everglades Depth Estimation Network spanning from January 2002 to May 2023, the global and local models were trained to forecast water levels with a two-day lead time. Our findings suggest that both the global and local models perform with approximately the same level of accuracy, with Mean Absolute Relative Error values ranging from 0.38% to 1.4% at the selected stations. The developed global model has demonstrated strong potential as a standalone forecasting tool for the entire study area in the Everglades and could eliminate the need for developing multiple local models. This finding also highlights how machine learning can capture complex spatial and temporal relationships to generate accurate water level predictions on a regional scale.

Florida

Efficient physics‐informed ground‐motion simulations with reduced‐order models: CyberShake implications and high‐resolution site terms for southern San Andreas fault earthquakes

Recent advances in Probabilistic Seismic Hazard Analysis (PSHA) leverage physics‐based ground‐motion simulations to estimate seismic hazard, such as the CyberShake project. However, computational costs quickly escalate when performing PSHA for numerous faults or sites and can become prohibitively expensive. To reduce computational demands, CyberShake uses reciprocity and interpolates physics‐informed corrections from simulations conducted at fewer locations, but the accuracy of these interpolations remains poorly quantified. To quantify the interpolation accuracy, we derive high‐resolution, frequency‐dependent site terms for southern California and compare them with interpolated site terms using the CyberShake approach. We accomplish this by performing a set of earthquake point‐source simulations distributed along the nonplanar fault geometry for the southern San Andreas fault (SSAF) extending from Bombay Beach to Lake Hughes. Using SeisSol, we simulate three minutes of viscoelastic seismic wave propagation for these sources and store the horizontal‐component Green’s functions for 480,000 sites. We then use a scientific machine learning approach based on interpolated proper orthogonal decomposition to construct an accurate reduced‐order model of the Green’s functions to efficiently predict effective amplitude spectra (EAS) for finite‐source rupture models of SSAF earthquakes. Using minimum curvature interpolation with tension, as used in CyberShake, we compare the interpolated site terms against our high‐resolution site terms. We identify local discrepancies with EAS differing by up to a factor of approximately three. Furthermore, we identify locations where unexpectedly high or low ground motions are missed when using the interpolated dataset for these earthquakes. We estimate that our approach may be used within CyberShake to reduce the time‐to‐solution by a factor of 336 for the entire earthquake rupture forecast. Our analysis of physics‐based site terms provides more insight into the seismic hazard due to SSAF ruptures and guides future developments by combining high‐performance computing and reduced‐order modeling techniques for PSHA.

California

False positives in the identification of dynamic earthquake triggering

Dynamic earthquake triggering is commonly identified through the temporal correlation between increased seismicity rates and global earthquakes that are possible triggering events. However, correlation does not imply causation. False positives may occur when unrelated seismicity rate changes coincidently occur at around the time of candidate triggers. We investigate the expected false positive rate in Southern California with global M ≥ 6 earthquakes as candidate triggers. We compute the false positive rate by applying the statistical tests used by DeSalvio and Fan (2023), https://doi.org/10.1029/2023jb026487 to synthetic earthquake catalogs with no real dynamic triggering. We find a false positive rate of ∼3.5%–8.5% when realistic earthquake clustering is present, consistent with the 95% confidence typically used in seismology. However, when this false positive rate is applied to the tens of thousands of spatial-temporal windows in Southern California tested in DeSalvio and Fan (2023), https://doi.org/10.1029/2023jb026487 , thousands of false positives are expected. The expected false positive occurrence is large enough to explain the observed apparent triggering following 70% of large global earthquakes (DeSalvio & Fan, 2023, https://doi.org/10.1029/2023jb026487 ), without requiring any true dynamic triggering. Aside from the known triggering from the nearby El Mayor-Cucapah, Mexico, earthquake, the spatial and temporal characteristics of the reported triggering are indistinguishable from random false positives. This implies that best practice for dynamic triggering studies that depend on temporal correlation is to estimate the false positive rate and investigate whether the observed apparent triggering is distinguishable from the correlations that may occur by chance.

JGR Solid Earth

Cajon Pass and the southern San Andreas Fault System: Earthquake cycle stress accumulation and present-day loading

With over a century since the last major rupture affecting the wider Los Angeles region, tectonic stress has steadily built along the southern San Andreas and San Jacinto fault systems, raising concerns of an imminent large earthquake. Cajon Pass, located at the junction of these faults, represents a critical site for potential through-going ruptures in Southern California. We constructed new 4D earthquake cycle simulations using a 1000-year paleoseismic rupture history of the San Andreas Fault System (SAFS) to assess spatial and temporal variations in stress. A semi-analytic Fourier transform model was used to compute stress from 3D dislocations in an elastic plate overlying a Maxwell viscoelastic half-space, assuming a complete coseismic reset of resolved shear stress on ruptured elements. Results show highest stress accumulation north of Cajon Pass (∼1.8 MPa/100 years) due to greater slip rates, and lower rates south of Cajon Pass (∼1.0–1.5 MPa/100 years). By 2025, Coulomb stress is estimated at 2.8 MPa on the Mojave South (MOS) segment, 1.8 MPa on the North San Bernardino (NSB1) segment and 3.6 MPa on the San Jacinto Bernardino (SJB) segment. Segments accumulate stresses with characteristic ranges of pre-event stress interpreted as failure thresholds: 1.2–2.7 MPa for MOS, 0.4–1.6 MPa for NSB1, and 1.2–2.9 MPa for SJB. When the stress disparity between segments SJB and MOS narrows, the faults appear to rupture jointly, suggesting that stress levels may control how Cajon Pass acts as an earthquake gate. These results may inform seismic hazard assessments by linking stress evolution to fault interactions.

California

A generalized deep learning model to detect and classify volcano seismicity

Volcano seismicity is often detected and classified based on its spectral properties. However, the wide variety of volcano seismic signals and increasing amounts of data make accurate, consistent, and efficient detection and classification challenging. Machine learning (ML) has proven very effective at detecting and classifying tectonic seismicity, particularly using Convolutional Neural Networks (CNNs) and leveraging labeled datasets from regional seismic networks. Progress has been made applying ML to volcano seismicity, but efforts have typically been focused on a single volcano and are often hampered by the limited availability of training data. We build on the method of Tan et al. [2024] ( 10.1029/2024JB029194 ) to generalize a spectrogram-based CNN termed the VOlcano Infrasound and Seismic Spectrogram Neural Network ( VOISS-Net ) to detect and classify volcano seismicity at any volcano. We use a diverse training dataset of over 270,000 spectrograms from multiple volcanoes: Pavlof, Semisopochnoi, Tanaga, Takawangha, and Redoubt volcanoes\replaced (Alaska, USA); Mt. Etna (Italy); and Kīlauea, Hawai`i (USA). These volcanoes present a wide range of volcano seismic signals, source-receiver distances, and eruption styles. Our generalized VOISS-Net model achieves an accuracy of 87 % on the test set. We apply this model to continuous data from several volcanoes and eruptions included within and outside our training set, and find that multiple types of tremor, explosions, earthquakes, long-period events, and noise are successfully detected and classified. The model occasionally confuses transient signals such as earthquakes and explosions and misclassifies seismicity not included in the training dataset (e.g. teleseismic earthquakes). We envision the generalized VOISS-Net model to be applicable in both research and operational volcano monitoring settings.

Volcanica

Preventing overfitting when using tree-based methods for mapping hydrothermal favorability

Ensemble tree-based algorithms are robust tools for estimating sparsely distributed resources with non-linear dependencies (e.g., hydrothermal systems). These algorithms naturally accommodate the threshold conditions necessary to enable and support hydrothermal systems (e.g., having sufficient heat and permeability) and are simpler than many other non-linear machine learning strategies (e.g., artificial neural networks), which is an advantage when working with few labeled examples from which to learn. In previous work, we used eXtreme Gradient Boosting (XGBoost) to produce regional prediction and uncertainty maps of hydrothermal favorability; however, recent studies suggest that, even when properly applied, XGBoost has some risk of overfitting when there are few labeled examples from which to learn. To evaluate overfitting when constructing hydrothermal favorability maps with tree-based methods, we compare XGBoost with Extremely Randomized Trees (ExtraTrees), another ensemble tree-based algorithm that has the potential to underfit when using few labeled examples. We hold all other modeling parameters constant, resulting in two contrasting favorability maps of conventional geothermal resources for the Great Basin. Our results indicate that ExtraTrees demonstrably reduces overfitting compared with XGBoost. After considering overall performance, we conclude that ExtraTrees provides a more suitable modeling approach than XGBoost for the purposes of conventional hydrothermal resource assessments.

Conference Paper

(Re)discovering the seismicity of Antarctica: A new seismic catalog for the southernmost continent

We apply a machine learning (ML) earthquake detection technique on over 21 yr of seismic data from on‐continent temporary and long‐term networks to obtain the most complete catalog of seismicity in Antarctica to date. The new catalog contains 60,006 seismic events within the Antarctic continent for 1 January 2000–1 January 2021, with estimated moment magnitudes (⁠Mw ⁠) between −1.0 and 4.5. Most detected seismicity occurs near Ross Island, large ice shelves, ice streams, ice‐covered volcanoes, or in distinct and isolated areas within the continental interior. The event locations and waveform characteristics indicate volcanic, tectonic, and cryospheric sources. The catalog shows that Antarctica is more seismically active than prior catalogs would indicate, examples include new tectonic events in East Antarctica, seismic events near and around the vicinity of David Glacier, and many thousands of events in the Mount Erebus region. This catalog provides a resource for more specific studies using other detection and analysis methods such as template matching or transfer learning to further discriminate source types and investigate diverse seismogenic processes across the continent.

Seismological Research Letters

Distinguishing natural sources from anthropogenic events in seismic data

As seismic data are increasingly used to investigate a diverse range of subsurface phenomena beyond regular fast-rupturing earthquakes (Peng and Gomberg, 2010; Beroza and Ide, 2011), it is important to acknowledge that human-generated ground vibrations may be mistaken for naturally generated subsurface processes (Larose et al., 2015; Li et al., 2018). Correct discrimination of natural processes from anthropogenic noise is especially pressing given the trend in seismic detection research toward automated algorithms and machine learning methods (Yoon et al., 2015; Kong et al., 2019;Mousavi and Beroza, 2022) and the growth in seismic data collection in new environments such as urban and industry settings (e.g., Díaz et al.,2017).

Seismological Research Letters

Estimating the probability of export restrictions to inform mineral criticality

To assess risks associated with advanced technologies’ supply chain disruptions, governmental agencies and others have developed mineral “criticality” assessments, with criticality described using the economic impact and probability of supply chain disruptions. Previous work developed subjective supply risk indicators to approximate this probability, typically combining several factors such as supply diversity and trading partners’ political stability, where indicator weightings can substantially impact results. This work explicitly quantifies export barrier probability using an ensemble of machine learning classifiers, with probability estimates informed by exogenous variables, including prior barrier implementation and global export dominance. Major differences in high-probability countries and commodities are observed across models, but the ensemble method highlights Indonesia, China, Tanzania, and the United States as particularly high risk. The Supplementary Data File provides export barrier probability estimates for each analyzed country-commodity pair, enabling a direct, quantitative, objective contribution to assessing mineral criticality, enhancing risk identification and prioritization for policymakers.

Resources, Conservation, and Recycling

Estimating the probability of export restrictions to inform mineral criticality

As demand for advanced technologies rises, mineral commodities will increase in geopolitical importance. To assess risks associated with mineral commodity supply chain disruptions, governmental agencies and others have developed "criticality" assessments, with criticality described using the economic impact and probability of supply chain disruptions. In previous work, subjective supply risk indicators were developed to approximate this probability, typically combining several factors such as supply diversity and political stability of trading partners, where indicator weightings can substantially impact results. This work explicitly quantifies trade barrier probability using an ensemble of several machine learning classifiers, with probability estimates informed by exogenous variables such as prior trade barrier implementation and global export dominance. Major differences in the high-probability countries and commodities are observed across models, but the ensemble method highlights Indonesia, China, Tanzania, and the United States as particularly high risk. This approach enables a direct, quantitative, objective approach to assessing trade barrier probability, enhancing risk identification and prioritization for policymakers.

SSRN

Zircon as a pathfinder to REE mineralization

Carbonatites and alkaline silicate rocks are major primary sources of the rare earth elements (REE) and other critical metals, such as Nb. Despite the economic significance of these rocks, their formation and the processes of REE enrichment are poorly understood. Here, statistical analysis of a global dataset demonstrates that zircon geochemistry is a powerful recorder of REE metallogenesis and a potential pathfinder for REE deposits. Zircons from REE and Nb fertile intrusions lack Eu anomalies and have elevated Gd/Yb and Th/Yb, indicating they crystallised from magmas that originated from deep, oxidised and enriched mantle sources. Complexes with Nb enrichment have low U/Nb, reflecting an enriched mantle source, whereas high U/Nb in REE-only fertile intrusions suggest a subduction-metasomatised mantle source. Machine learning models demonstrate high accuracy in classifying zircon from barren and fertile deposits. Classification of detrital zircons shows that REE-enriched deposits correlate with supercontinent assembly, whereas Nb fertile complexes are associated with supercontinent breakup. This approach offers a new, mineral to global scale, petrologic and exploration tool that enhances understanding of REE metallogenesis.

Geochemical Perspectives Letters

A benchmark dataset and workflow for landslide susceptibility zonation

Landslide susceptibility shows the spatial likelihood of landslide occurrence in a specific geographical area and is a relevant tool for mitigating the impact of landslides worldwide. As such, it is the subject of countless scientific studies. Many methods exist for generating a susceptibility map, mostly falling under the definition of statistical or machine learning. These models try to solve a classification problem: given a collection of spatial variables, and their combination associated with landslide presence or absence, a model should be trained, tested to reproduce the target outcome, and eventually applied to unseen data. Contrary to many fields of science that use machine learning for specific tasks, no reference data exist to assess the performance of a given method for landslide susceptibility. Here, we propose a benchmark dataset consisting of 7360 slope units encompassing an area of about 4,100 km 2 "> 4,100 km 2 in Central Italy. Using the dataset, we tried to answer two open questions in landslide research: (1) what effect does the human variability have in creating susceptibility models; (2) how can we develop a reproducible workflow for allowing meaningful model comparisons within the landslide susceptibility research community. With these questions in mind, we released a preliminary version of the dataset, along with a “call for collaboration,” aimed at collecting different calculations using the proposed data, and leaving the freedom of implementation to the respondents. Contributions were different in many respects, including classification methods, use of predictors, implementation of training/validation, and performance assessment. That feedback suggested refining the initial dataset, and constraining the implementation workflow. This resulted in a final benchmark dataset and landslide susceptibility maps obtained with many classification methods. Values of area under the receiver operating characteristic curve obtained with the final benchmark dataset were rather similar, as an effect of constraints on training, cross–validation, and use of data. Brier score results show larger variability, instead, ascribed to different model predictive abilities. Correlation plots show similarities between results of different methods applied by the same group, ascribed to a residual implementation dependence. We stress that the experiment did not intend to select the “best” method but only to establish a first benchmark dataset and workflow, that may be useful as a standard reference for calculations by other scholars. The experiment, to our knowledge, is the first of its kind for landslide susceptibility modeling. The data and workflow presented here comparatively assess the performance of independent methods for landslide susceptibility and we suggest the benchmark approach as a best practice for quantitative research in geosciences.

Earth-Science Reviews

Evidence for fluid pressurization of fault zones and persistent sensitivity to injection rate beneath the Raton Basin

Subsurface wastewater injection has increased the seismicity rate within the Raton Basin over more than two decades, with the basin-wide injection rate peaked between 2009-2015. To understand the evolution of injection-induced earthquakes, we systematically analyzed 2016-2024 broadband recordings with a machine-learning-based phase picker and constructed a catalog with 95,993 earthquakes (-1≤ M L ≤4.3). We then inverted for full centroid moment tensors (CMT) for 90 M L ≥ 2 events, with a special interest in constraining the non-double-couple components via probabilistic metrics. Both relocations and CMT solutions support basement-rooted normal faults, including graben and half-graben structures. Furthermore, we observe the non-double-couple components that imply elevated pore pressure in the fault zones. An earthquake cluster emerged in the north-central basin in 2023, preceded by ~1-yr of increased injection volume from wells within 15km. Despite a basin-wide decrease in the injection volume, we highlights the persistence of seismicity that remains to sensitive to injection rates within the Raton Basin.

Colorado, New Mexico

Rapid earthquake magnitude classification via P-wave strains from borehole strainmeters and Distributed Acoustic Sensing

Distributed Acoustic Sensing (DAS) offers a promising approach for earthquake early warning (EEW) in settings where seismic networks are costly to maintain. By repurposing fiber-optic cables as dense strainmeter arrays, DAS enables real-time earthquake detection wherever those fibers are accessible. However, poor azimuthal coverage and challenges in estimating magnitude from strain measurements remain key hurdles in applying for earthquake monitoring. Here, we develop a machine learning method to distinguish large (M≥5.4) earthquakes from smaller ones within the first 4 seconds of a strain waveform after a P-wave arrival without determining location. Using ensemble decision tree models trained on borehole strainmeter data (3.5≤M≤7.1) and tested on onshore DAS waveforms (including the 2024 M7 Offshore Cape Mendocino earthquake), we find that low-frequency (0.2–0.5 Hz) continuous wavelet transform coefficients are the strongest predictors of magnitude, in addition to strain amplitude. Both DAS and borehole strainmeters effectively capture long-period strain signals, making these findings valuable for EEW systems. Our method shows high precision compared to the real-time EEW system, ShakeAlert®, supporting the position that DAS is a viable technology for earthquake monitoring and magnitude classification.

California

A method to obtain remotely sensed grain size distributions from nonplanar granular deposits

Constraining the grain size distribution of granular deposits with complex surfaces is difficult with existing approaches. Field and laboratory techniques are time consuming and limited by the maximum grain size that laboratories can accommodate. In this study, we present a new method to identify the coarse fraction of the grain size distribution at a debris-flow fan deposit surveyed with terrestrial laser scanning (TLS) in Glenwood Canyon, Colorado, USA. This method is a novel grain segmentation algorithm developed for application to point cloud data of deposits with complex surfaces and angular grains ranging in size from centimeters to a meter. This approach combines an existing random forest machine learning method with a novel iterative clustering algorithm. We compared the grain size distribution from our algorithm with a Wolman pebble count conducted in the field, and found a root mean squared error of less than 2 cm from the 5th to 95th percentile of the grain size distribution of grains ranging from cobble to boulder sized (6.3–78 cm in our application). Finally, we compared our new algorithm with an existing open-source grain segregation algorithm, and our method outperformed the selected alternative when applied to the debris-flow deposit point cloud.

Colorado