Search USGSSearch

SEARCH · Search USGS

Results for “Statistical Methods & Applications”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Decomposing leftovers: Event, path, and site residuals for a small magnitude ANZA region GMPE

Ground‐motion prediction equations (GMPEs) are critical elements of probabilistic seismic hazard analysis (PSHA), as well as for other applications of ground motions. To isolate the path component for the purpose of building nonergodic GMPEs, we compute a regional GMPE using a large dataset of peak ground accelerations (PGAs) from small‐magnitude earthquakes ( ⁠0.5 ≤ M ≤ 4.5 with > 10,000 events, yielding ∼120,000 recordings) that occurred in 2013 centered around the ANZA seismic network (hypocentral distances ≤180 km⁠ ) in southern California. We examine two separate methods of obtaining residuals from the observed and predicted ground motions: a pooled ordinary least‐squares model and a mixed‐effects maximum‐likelihood model. Whereas the former is often used by the broader seismological community, the latter is widely used by the ground‐motion and engineering seismology community. We confirm that mixed‐effects models are the preferred and most statistically robust method to obtain event, path, and site residuals and discuss the reasoning behind this. Our results show that these methods yield different consequences for the uncertainty of the residuals, particularly for the event residuals. Finally, our results show no correlation (correlation coefficient [CC] <0.03⁠ ) between site residuals and the classic site‐characterization term V S30 ⁠ , the time‐averaged shear‐wave velocity in the top 30 m at a site. We propose that this is due to the relative homogeneity of the site response in the region and perhaps due to shortcomings in the formulation of V S30 ⁠ and suggest applying the provided PGA site correction terms to future ground‐motion studies for increased accuracy.

California

An empirical determination of the minimum number of measurements needed to estimate the mean random vitrinite reflectance of disseminated organic matter

In coal samples, published recommendations based on statistical methods suggest 100 measurements are needed to estimate the mean random vitrinite reflectance ( R v−r ) to within ±2%. Our survey of published thermal maturation studies indicates that those using dispersed organic matter (DOM) mostly have an objective of acquiring 50 reflectance measurements. This smaller objective size in DOM versus that for coal samples poses a statistical contradiction because the standard deviations of DOM reflectance distributions are typically larger indicating a greater sample size is needed to accurately estimate R v−r in DOM. However, in studies of thermal maturation using DOM, even 50 measurements can be an unrealistic requirement given the small amount of vitrinite often found in such samples. Furthermore, there is generally a reduced need for assuring precision like that needed for coal applications. Therefore, a key question in thermal maturation studies using DOM is how many measurements of R v−r are needed to adequately estimate the mean. Our empirical approach to this problem is to compute the reflectance distribution statistics: mean, standard deviation, skewness, and kurtosis in increments of 10 measurements. This study compares these intermediate computations of R v−r statistics with a final one computed using all measurements for that sample. Vitrinite reflectance was measured on mudstone and sandstone samples taken from borehole M-25 in the Cerro Prieto, Mexico geothermal system which was selected because the rocks have a wide range of thermal maturation and a comparable humic DOM with depth. The results of this study suggest that after only 20–30 measurements the mean R v−r is generally known to within 5% and always to within 12% of the mean R v−r calculated using all of the measured particles. Thus, even in the worst case, the precision after measuring only 20–30 particles is in good agreement with the general precision of one decimal place recommended for mean R v−r measurements on DOM. The coefficient of variation ( V = standard deviation/mean) is proposed as a statistic to indicate the reliability of the mean R v−r estimates made at n ⪡ 20. This preliminary study suggests a V < 0.1 indicates a reliable mean and a V > 0.2 suggests an unreliable mean in such small samples.

Organic Geochemistry

Assessing potential effects of highway runoff on receiving-water quality at selected sites in Oregon with the Stochastic Empirical Loading and Dilution Model (SELDM)

In 2012, the U.S. Geological Survey and the Oregon Department of Transportation began a cooperative study to demonstrate use of the Stochastic Empirical Loading and Dilution Model (SELDM) for runoff-quality analyses in Oregon. SELDM can be used to estimate stormflows, constituent concentrations, and loads from the area upstream of a stormflow discharge site, from the site of interest and in the receiving waters downstream of the discharge. SELDM also can be used to assess the potential effectiveness of best management practices (BMP) for mitigating potential effects of runoff in receiving waters. Nominally, SELDM is a highway-runoff model, but it is well suited for analysis of runoff from other land uses as well. This report provides case studies and examples to demonstrate stochastic-runoff modeling concepts and to demonstrate application of the model. Basin characteristics from six Oregon highway study sites were used to demonstrate various applications of the model. The highway catchment and upstream basin drainage areas of these study sites ranged from 3.85 to 11.83 acres and from 0.16 to 6.56 square miles, respectively. The upstream basins of two sites are urbanized, and the remaining four sites are less than 5 percent impervious. SELDM facilitates analysis by providing precipitation, pre-storm streamflow, and other variables by region or from hydrologically similar sites. In Oregon, there can be large variations in precipitation and streamflow among nearby sites. Therefore, spatially interpolated geographic information system data layers containing storm-event precipitation and pre-storm streamflow statistics specific to Oregon were created for the study using Kriging techniques. Concentrations and loads of cadmium, chloride, chromium, copper, iron, lead, nickel, phosphorus, and zinc were simulated at the six Oregon highway study sites by using statistics from sites in other areas of the country. Water‑quality datasets measured at hydrologically similar basins in the vicinity of the study sites in Oregon were selected and compiled to estimate stormflow-quality statistics for the upstream basins. The quality of highway runoff and some upstream stormflow constituents were simulated by using statistical moments (average, standard deviation, and skew) of the logarithms of data. Some upstream stormflow constituents were simulated by using transport curves, which are relations between stormflow and constituent concentrations. Stochastic analyses were done by using SELDM to demonstrate use of the model and to illustrate the types of information that stochastic analyses may provide: 1. An analysis was done to demonstrate use of dilution factors as an initial reconnaissance tool for comparing relative risk among sites. 2. An analysis of hardness-dependent, water-quality criteria was done to illustrate the effects of variations in hardness and flow on the application and interpretation of such criteria. This analysis shows that hardness-dependent criteria can vary by an order of magnitude among storm events because hardness is diluted by stormflows. 3. An analysis of uncertainties in input and output values was done to demonstrate that properly selected robust datasets are needed to represent conditions at a site of interest. This analysis shows that the rate of water-quality exceedances that are measured or simulated may depend on sample size and the luck of the draw. 4. An analysis was done to demonstrate that SELDM and other Monte Carlo models may generate extreme values from input statistics, which may or may not be feasible based on physicochemical or hydrological limits. 5. An analysis of BMP modeling methods was done to demonstrate use of the model for estimating treatment requirements for meeting water-quality objectives. 6. An analysis of the use of grab sampling and nonstochastic upstream modeling methods was done to evaluate the potential effects on modeling outcomes. Additional analyses using surrogate water-quality datasets for the upstream basin and highway catchment were provided for six Oregon study sites to illustrate the risk-based information that SELDM will produce. These analyses show that the potential effects of highway runoff on receiving-water quality downstream of the outfall depends on the ratio of drainage areas (dilution), the quality of the receiving water upstream of the highway, and the concentration of the criteria of the constituent of interest. These analyses also show that the probability of exceeding a water-quality criterion may depend on the input statistics used, thus careful selection of representative values is important.

Oregon

Geographic information system as country-level development and monitoring tool, Senegal example

Geographic information systems (GIS) allow an investigator the capability to merge and analyze numerous types of country-level resource data. Hypothetical resource analysis applications in Senegal were conducted to illustrate the utility of a GIS for development planning and resource monitoring. Map and attribute data for soils, vegetation, population, infrastructure, and administrative units were merged to form a database within a GIS. Several models were implemented using a GIS to: analyze development potential for sustainable dryland agriculture; prioritize where agricultural development should occur based upon a regional food budget; and monitor dynamic events with remote sensing. The steps for implementing a GIS analysis are described and illustrated, and the use of a GIS for conducting an economic analysis is outlined. Using a GIS for analysis and display of results opens new methods of communication between resource scientists and decision makers. Analyses yielding country-wide map output and detailed statistical data for each level of administration provide the advantage of a single system that can serve a variety of users.

Conference Paper

Visitor spending effects: assessing and showcasing America's investment in national parks

This paper provides an overview of the evolution, future, and global applicability of the U.S. National Park Service's (NPS) visitor spending effects framework and discusses the methods used to effectively communicate the economic return on investment in America's national parks. The 417 parks represent many of America's most iconic destinations: in 2016, they received a record 331 million visits. Competing federal budgetary demands necessitate that, in addition to meeting their mission to preserve unimpaired natural and cultural resources for the enjoyment of the people, parks also assess and showcase their contributions to the economic vitality of their regions and the nation. Key approaches explained include the original Money Generation Model (MGM) from 1990, MGM2 used from 2001, and the visitor spending effects model which replaced MGM2 in 2012. Detailed discussion explains the NPS's visitor use statistics system, the formal program for collecting, compiling, and reporting visitor use data. The NPS is now establishing a formal socioeconomic monitoring (SEM) program to provide a standard visitor survey instrument and a long-term, systematic sampling design for in-park visitor surveys. The pilot SEM survey is discussed, along with the need for international standardization of research methods.

Journal of Sustainable Tourism

The Amphibian Research and Monitoring Initiative (ARMI): 5-year report

The Amphibian Research and Monitoring Initiative (ARMI) is an innovative, multidisciplinary program that began in 2000 in response to a congressional directive for the Department of the Interior to address the issue of amphibian declines in the United States. ARMI&rsquo;s formulation was cross-disciplinary, integrating U.S. Geological Survey scientists from Biology, Water, and Geography to develop a course of action (Corn and others, 2005a). The result has been an effective program with diverse, yet complementary, expertise. ARMI&rsquo;s approach to research and monitoring is multiscale. Detailed investigations focus on a few species at selected local sites throughout the country; monitoring addresses a larger number of species over broader areas (typically, National Parks and National Wildlife Refuges); and inventories to document species occurrence are conducted more extensively across the landscape. Where monitoring is conducted, the emphasis is on an ability to draw statistically defensible conclusions about the status of amphibians. To achieve this objective, ARMI has instituted a monitoring response variable that has nationwide applicability. At research sites, ARMI focuses on studying species/environment interactions, determining causes of observed declines, and developing new techniques to sample populations and analyze data. Results from activities at all scales are provided to scientists, land managers, and policymakers, as appropriate. The ARMI program and the scientists involved contribute significantly to understanding amphibian declines at local, regional, national, and international levels. Within National Parks and National Wildlife Refuges, findings help land managers make decisions applicable to amphibian conservation. For example, the National Park Service (NPS) selected amphibians as a vital sign for several of their monitoring networks, and ARMI scientists provide information and assistance in developing monitoring methods for this NPS effort. At the national level, ARMI has had major exposure at a variety of meetings, including a dedicated symposium at the 2004 joint meetings of the Herpetologists&rsquo; League, the American Society of Ichthyologists and Herpetologists, and the Society for the Study of Amphibians and Reptiles. Several principal investigators have brought international exposure to ARMI through venues such as the World Congress of Herpetology in South Africa in 2005 (invited presentation by Dr. Gary Fellers), the Global Amphibian Summit, sponsored by the International Union for Conservation of Nature (IUCN) and Wildlife Conservation International, in Washington, D.C., 2005 (invited participation by Dr. P.S. Corn), and a special issue of the international herpetological journal Alytes focused on ARMI in 2004 (edited by Dr. C.K. Dodd, Jr.). ARMI research and monitoring efforts have addressed at least 7 of the 21 Threatened and Endangered Species listed by the U.S. Fish and Wildlife Service (California red-legged frog [Rana draytonii], Chiricahua leopard frog [R. chiricahuensis], arroyo toad [Bufo californicus], dusky gopher frog [Rana sevosa], mountain yellow-legged frog [R. muscosa], flatwoods salamander [Ambystoma cingulatum], and the golden coqui [Eleutherodactylus jasperi]), and 9 additional species of concern recognized by the IUCN. ARMI investigations have addressed time-sensitive research, such as emerging infectious diseases and effects on amphibians related to natural disasters like wildfire, hurricanes, and debris flows, and the effects of more constant, environmental change, like urban expansion, road development, and the use of pesticides. Over the last 5 years, ARMI has partnered with an extensive list of government, academic, and private entities. These partnerships have been fruitful and have assisted ARMI in developing new field protocols and analytic tools, in using and refining emerging technologies to improve accuracy and efficiency of data handling, in conducting amphibian disease, malformation, and environmental effects research, and in implementing a network of monitoring and research sites. Accomplishments from these endeavors include more than 40 publications on amphibian status and trends, nearly 100 publications on amphibian ecology and causes of declines, and over 30 methodological publications. Several databases have emerged as a result of ARMI and its partnerships; one, a digital atlas of ranges for all U.S. amphibian species, was used by the IUCN to display amphibian distribution maps in the Global Amphibian Assessment Project. Given the scope of ARMI and the panoply of projects, findings have had implications for policy. Investigations that demonstrate amphibian declines or illuminate causes of declines provide valuable information about habitat management, environmental effects, mechanisms for the spread of disease, and human/amphibian interfaces. This information has been made available to land managers, scientists, educators, Congress and other policymakers, and the public. The support afforded ARMI by Congress has been influential in the program&rsquo;s development and success. The value of ARMI&rsquo;s efforts will continue to increase as we are able to extend our studies spatially and temporally to answer critical questions with more confidence. We are using ARMI&rsquo;s resources efficiently and continuing to develop innovative mechanisms for leveraging resources for maximum effectiveness during challenging financial times. This report is a 5-year retrospective of the structure, methodology, progress, and contributions to the broader scientific community that have resulted from this national USGS program. We evaluate ARMI&rsquo;s success to date, with regard to the challenges faced by the program and the strengths that have emerged. We chart objectives for the next 5 years that build on current accomplishments, highlight areas meriting further research, and direct efforts to overcome existing weaknesses.

Scientific Investigations Report

An approach for decomposing river water-quality trends into different flow classes

A number of statistical approaches have been developed to quantify the overall trend in river water quality, but most approaches are not intended for reporting separate trends for different flow conditions. We propose an approach called FN 2Q , which is an extension of the flow-normalization (FN) procedure of the well-established WRTDS (“Weighted Regressions on Time, Discharge, and Season”) method. The FN 2Q approach provides a daily time series of low-flow and high-flow FN flux estimates that represent the lower and upper half of daily riverflow observations that occurred on each calendar day across the period of record. These daily estimates can be summarized into any time period of interest (e.g., monthly, seasonal, or annual) for quantifying trends. The proposed approach is illustrated with an application to a record of total nitrogen concentration (632 samples) collected between 1985 and 2018 from the South Fork Shenandoah River at Front Royal, Virginia (USA). Results show that the overall FN flux of total nitrogen has declined in the period of 1985–2018, which is mainly attributable to FN flux decline in the low-flow class. Furthermore, the decline in the low-flow class was highly correlated with wastewater effluent loads, indicating that the upgrades of treatment technology at wastewater treatment facilities have likely led to water-quality improvement under low-flow conditions. The high-flow FN flux showed a spike around 2007, which was likely caused by increased delivery of particulate nitrogen associated with sediment transport. The case study demonstrates the utility of the FN 2Q approach toward not only characterizing the changes in river water quality but also guiding the direction of additional analysis for capturing the underlying drivers. The FN 2Q approach (and the published code) can easily be applied to widely available river monitoring records to quantify water-quality trends under different flow conditions to enhance understanding of river water-quality dynamics.

Virginia

Global cropland-extent product at 30-m resolution (GCEP30) derived from Landsat satellite time-series data for the year 2015 using multiple machine-learning algorithms on Google Earth Engine cloud

Executive Summary Global food and water security analysis and management require precise and accurate global cropland-extent maps. Existing maps have limitations, in that they are (1) mapped using coarse-resolution remote-sensing data, resulting in the lack of precise mapping location of croplands and their accuracies; (2) derived by collecting and collating national statistical data that are often subjective, leading to substantial uncertainties in cropland-area estimates, as well as their locations; and (3) extracted from one or more classes of a land use–land cover product in which cropland classes are not the focus of mapping, leading to their mixing with other classes and creating significant errors of omission and commission. These limitations can be overcome by producing high-resolution cropland-extent maps using satellite-sensor data, such as Landsat 30-m resolution or higher. The most fundamental cropland product is the high-resolution cropland-extent map because all higher level cropland products, such as crop-watering method (that is, whether crops are irrigated or rainfed), crop types, cropping intensities, cropland fallows, crop productivity, and crop-water productivity, are dependent on a precise and accurate cropland-extent product. Given these realities, the overarching goal of this study was to produce a Landsat satellite-derived global cropland-extent product at 30-m resolution. The work, which involved a paradigm shift in how global cropland-extent maps are produced, involved the following five key steps: (1) petabyte-scale computing that involved multiyear, 8- to 16-day, time-series Landsat 30-m resolution data for the global land surface; (2) composition of analysis-ready data (ARD) cubes; (3) creation of a large global-reference data hub for machine learning; (4) use of multiple machine-learning algorithms (MLAs) by writing software and computing in the cloud; and (5) Google Earth Engine (GEE) cloud computing. The five key steps involved nine distinct phases. First, the world was segmented into 74 agroecological zones (AEZs). Second, Landsat 8- to 16-day data were used to time-composite 10-band (blue, green, red, near-infrared, short-wave infrared band 1, short-wave infrared band 2, thermal infrared, enhanced vegetation index, normalized difference water index, and normalized difference vegetation index) Landsat 30-m resolution data cubes for every 2- to 4-month time period during 3- to 4-year periods (stated as nominal-year 2015 or, simply, 2015), along with two additional 30-m resolution bands (Shuttle Radar Topography Mission elevation, and slope) in each of the 74 AEZs. Third, more than 100,000 reference-training data samples were collected using ground data (some of which were collected using a mobile application), as well as submeter- to 5-m-resolution, very high-resolution imagery sourced from other reliable sources. Fourth, reference-training data were used to create a knowledge base for separating cropland from noncropland. Fifth, MLAs such as the pixel-based supervised random forest and support-vector machines were written on the GEE using Python and JavaScript. Sixth, object-based recursive hierarchical segmentation algorithm was used, in addition to MLAs, to overcome uncertainties. Seventh, MLAs used the knowledge base to classify and separate cropland from noncropland. Eighth, accuracy assessment was conducted by generating error matrices for each of the 74 AEZs using 19,171 independent validation-data samples. Ninth, cropland areas were computed for all countries of the world and compared with United Nation’s (UN’s) Food and Agricultural Organization (FAO) and other national statistics. The outcome was a Landsat-derived global cropland-extent product at 30-m resolution (GCEP30), which has an overall accuracy of 91.7 percent. For the cropland class, producer’s accuracy was 83.4 percent, and user’s accuracy was 78.3 percent. GCEP30 calculated (using direct pixel count) the global net-cropland area (GNCA) for the year 2015 as 1.873 billion hectares (~12.6 percent of the Earth’s terrestrial area). The continental cropland distribution as a percentage of GNCA was Asia, 33 percent; Europe, 25.5 percent; Africa, 16.7 percent; North America, 14.4 percent; South America, 8.1 percent; and Australia and Oceania, 2.4 percent. The worldwide cropland areas in GCEP30 for 2015 were higher by 236 to 299 million hectares (Mha) compared to national statistics reported elsewhere for the same year (for example, in Food and Agriculture Organization’s corporate statistical database [FAOSTAT] and in the monthly irrigated and rainfed crop areas [MIRCA] database). The global cropland area reported for 2015 increased by 344 Mha (22.5 percent), compared to the year 2000. During the same period (2000–2015), the world’s population increased by 20 percent. Whereas some of these areal increases are real increases in cropland areas, others are due to the types of data, methods, and approaches used. Using the highest known resolution (compared to previous coarse-resolution global products) enabled this study to capture fragmented croplands. Coarse-resolution data compute areas on the basis of subpixels, which, for a large proportion of certain land use–land cover classes, will show only a certain percentage of the total pixel area as actual area. Subpixel areas can lead to substantial uncertainties in area computation, as determining the exact fraction of cropland areas within a coarse-resolution pixel is resource intensive and subject to errors. Other innovations in GCEP30 include reference-data hubs, machine learning, and cloud computing. Cropland areas in 214 countries, territories, departments, and regions were calculated for the year 2015 using GCEP30, on the basis of UN’s global administrative unit layers (GAUL) boundaries. The 10 leading countries in terms of cropland area (as a percentage of the GNCA) were India (9.6 percent), United States (8.95 percent), China (8.82 percent), Russia (8.32 percent), Brazil (3.42 percent), Ukraine (2.32 percent), Canada (2.29 percent), Argentina (2.05 percent), Indonesia (2 percent), and Nigeria (1.91 percent). Together, these 10 countries occupy 50 percent of the global cropland, and they have 52 percent of the global population. Their combined cropland area increased by 2 percent between 2000 and 2015, compared to the substantial increase in population of 517 million (15.5 percent). Together, India, United States, China, and Russia encompass 36 percent of the total area. In the United States and Canada, from 2000 to 2015, cropland decreased by about 2 percent, whereas their populations increased by 14 and 13 percent, respectively. The additional food requirements in these 10 countries, which are caused by increased populations, as well as increasing nutritional demands, are met by production increases in existing cropland or through virtual food trade, or both. More than 18 countries, territories, departments, or regions had 60 percent or more of their geographic area as cropland: Republic of Moldova, San Marino, and Hungary had more than 80 percent of the country’s area as cropland; Denmark, Ukraine, Ireland, and Bangladesh, 70 to 80 percent; and Uruguay, Netherlands, United Kingdom, Spain, Lithuania, Poland, Gaza Strip, Czechia, Italy, India, and Azerbaijan, 60 to 70 percent. Europe and South Asia can be considered agricultural capitals of the world, on the basis of their percentages of geographic area as cropland. United States, China, and Russia, which all have high cropland areas, are ranked second, third, and fourth in the world; India is ranked first. However, the amount of cropland as a percentage of the country’s geographic area is relatively very low for United States (18.3 percent), China (17.7 percent), and Russia (9.5 percent), whereas it is 60.5 percent for India. Most African and South American countries, territories, departments, or regions have less than 15 percent of their geographic area as cropland. China and India together house 36 percent of the world’s population; however, between 2000 and 2015, the amount of China’s cropland area fell by 18.9 percent, owing to urban expansion and the abandonment of farmlands caused by demographic changes (that is, the movement of population from villages to cities). In contrast, China’s population grew by 10 percent. The amount of India’s cropland increased by 8.5 percent, whereas its population grew by 20 percent. This study showed that, out of the 10 leading cropland countries, Ukraine, Nigeria, Russia, and Indonesia showed an 18 to 31 percent increase in cropland areas, on the basis of GCEP30 by the year 2015, compared to 2000. Nigeria’s cropland area increased by 25 percent, and its population increased by 31 percent in the same period. In these countries, food security is maintained by cropland expansion, productivity increases, and virtual food trade. Nevertheless, this trend of increasing net-cropland area and productivity will likely become difficult to maintain, owing to diminishing arable lands and plateauing of 50 years of continual yield increases, requiring policymakers to explore novel and data-supported approaches to solving future food security issues. The GCEP30 product, which can be browsed at full resolution at www.croplands.org , has been released for public download and use through U.S. Geological Survey (USGS)–National Aeronautics and Space Administration (NASA) Land Processes Distributed Active Archive Center (see https://lpdaac.usgs.gov/news/release-of-gfsad-30-meter-cropland-extent-products/ ).

Professional Paper

Methods to determine streamflow statistics based on data through water year 2021 for selected streamgages in or near Wyoming

The U.S. Geological Survey (USGS), in cooperation with the Wyoming Water Development Office, developed streamflow statistics for streamgages in and near Wyoming. Statistics were computed for active (through September 30, 2021) and discontinued USGS streamgages with 10 or more years of daily mean streamflow record. Streamflow at each streamgage was assessed for degree of human alteration owing to dams and diversions before streamflow statistics were computed. Streamflow records from 615 streamgages were used to compute basic, seasonal, and flow-duration statistics; streamflow records from 387 streamgages were used to compute n -day statistics, which are streamflow statistics describing streamflow over a number of days ( n ), and statistics that can be used for regional regression. The streamflow statistics are provided in a USGS data publication that accompanies this report and through the USGS StreamStats web-based application ( https://www.usgs.gov/streamstats ).

Wyoming

A test to evaluate the earthquake prediction algorithm, M8

A test of the algorithm M8 is described. The test is constructed to meet four rules, which we propose to be applicable to the test of any method for earthquake prediction: 1. An earthquake prediction technique should be presented as a well documented, logical algorithm that can be used by investigators without restrictions. 2. The algorithm should be coded in a common programming language and implementable on widely available computer systems. 3. A test of the earthquake prediction technique should involve future predictions with a black box version of the algorithm in which potentially adjustable parameters are fixed in advance. The source of the input data must be defined and ambiguities in these data must be resolved automatically by the algorithm. 4. At least one reasonable null hypothesis should be stated  in advance of testing the earthquake prediction method, and it should be stated how this null hypothesis will be used to estimate the statistical significance of the earthquake predictions. The M8 algorithm has successfully predicted several destructive  earthquakes, in the sense that the earthquakes occurred inside regions with linear dimensions from 384 to 854 km that the algorithm had identified as being in times of increased probability for strong earthquakes. In addition, M8 has successfully "post predicted" high percentages of strong earthquakes in regions to which it has been applied in retroactive studies. The statistical significance of previous predictions has not been established, however, and post-prediction studies in general are notoriously subject to success-enhancement through hindsight. Nor has it been determined how much more precise an M8 prediction might be than forecasts and probability-of-occurrence estimates made by other techniques. We view our test of M8 both as a means to better determine the effectiveness of M8 and as an experimental structure within which to make observations that might lead to improvements in the algorithm or conceivably lead to a radically different approach to earthquake prediction.

Open-File Report

Regional analysis of annual precipitation maxima in Montana

Dimensionless precipitation-frequency curves for estimating precipitation depths having large recurrence intervals were developed for 2-, 6-, and 24-hour storm durations for three homogeneous regions in Montana. Within each homogeneous region, at-site annual precipitation maxima were made dimensionless by dividing by the at-site mean and grouped so that a single frequency curve would be applicable for each duration. L-moment statistics were used to help define the homogeneous regions and to develop the dimensionless precipitation- frequency curves. Data from 459 precipitation stations were used after application of statistical tests to ensure that the data were not serially correlated and were stationary over the general period of data collection (1900-92). The data were found to have a small, but significant, degree of interstation correlation. The GEV distribution was used to construct dimensionless frequency curves of annual precipitation maxima for each duration within each region. Each dimensionless frequency curve was considered to be reliable for recurrence intervals up to the effective record length. Because of significant, though small, interstation correlation in all regions for all durations, and because the selected regions exhibited some heterogeneity, the effective record length was considered to be less than the total number of station-years of data. The effective record length for each duration in each region was estimated using a graphical method and found to range from 500 years for 6-hour duration data in Region 2 to 5,100 years for 24-hour duration data in Region 3.

Water-Resources Investigations Report

GenEst statistical models—A generalized estimator of mortality

Introduction GenEst (a generalized estimator of mortality) is a suite of statistical models and software tools for generalized mortality estimation. It was specifically designed for estimating the number of bird and bat fatalities at solar and wind power facilities, but both the software (Dalthorp and others, 2018) and the underlying statistical models are general enough to be useful in various situations to estimate the size of open populations when detection probabilities and search coverages are less than 1. In this report, we outline the statistical models and data structures underlying the estimator. The models are numerous, complex, and intricately interwoven. Discussion begins with broad, high-level overviews of the general models. The lower-level technical details are then gradually added. Broader and less technical discussions on the general context and applications of the models and the use of the software are available in the software user guide (Simonis and others, 2018), vignettes bundled with the software, and the help files within the software itself.

Techniques and Methods

Effects of hydrologic variability and remedial actions on first flush and metal loading from streams draining the Silverton caldera, 1992–2014

This study examined water quality in the upper Animas River watershed, a mined watershed that gained notoriety following the 2015 Gold King mine release of acid mine drainage to downstream communities. Water-quality data were used to evaluate trends in metal concentrations and loads over a two-decade period. Selected sites included three sites on tributary streams and one main-stem site on the Animas River downstream from the tributary confluences. During the study period, metal concentrations and loads varied seasonally and annually because of hydrologic variability and remedial actions designed to ameliorate the effects of acid mine drainage. Water-quality data were divided into two periods based on the timing of remedial activities in the watershed. The first period includes active water treatment, surface reclamation and installation of bulkheads in adits; the second period includes the decade following these activities. Water-quality data were used to estimate annual and monthly zinc loads using the Adjusted Maximum Likelihood Method (using LOADEST software) and U.S. Geological Survey streamflow data. This study presents one of the first applications of LOADEST focused on metal loads. Monthly flow-weighted concentrations were analysed using a Mann-Kendall trend test to determine the direction, magnitude, and significance of temporal trends in zinc loading in any given month and using t -test comparisons between the two periods. Zinc loads estimated for the Animas River below the tributaries indicate decreased zinc loading during the rising limb of the hydrograph in the second period, perhaps reflecting a reduction of snowmelt-derived zinc load following surface reclamation activities. In contrast, base-flow zinc loading increased at the main-stem site, perhaps because of the cessation of water treatment in tributary streams. Flow weighting of monthly load estimates yielded increased statistical significance and enabled more nuanced differentiation between the effects of hydrologic variability and remedial activities on zinc loading.

Colorado

Flow Durations, Low-Flow Frequencies, and Monthly Median Flows for Selected Streams in Connecticut through 2005

Flow durations, low-flow frequencies, and monthly median streamflows were computed for 91 continuous-record, streamflow-gaging stations in Connecticut with 10 or more years of record. Flow durations include the 99-, 98-, 97-, 95-, 90-, 85-, 80-, 75-, 70-, 60-, 50-, 40-, 30-, 25-, 20-, 10-, 5-, and 1-percent exceedances. Low-flow frequencies include the 7-day, 10-year (7Q10) low flow; 7-day, 2-year (7Q2) low flow; and 30-day, 2-year (30Q2) low flow. Streamflow estimates were computed for each station using data for the period of record through water year 2005. Estimates of low-flow statistics for 7 short-term (operated between 3 and 10 years) streamflow-gaging stations and 31 partial-record sites were computed. Low-flow estimates were made on the basis of the relation between base flows at a short-term station or partial-record site and concurrent daily mean streamflows at a nearby index station. The relation is defined by the Maintenance of Variance Extension, type 3 (MOVE.3) method. Several short-term stations and partial-record sites had poorly defined relations with nearby index stations; therefore, no low-flow statistics were derived for these sites. The estimated low-flow statistics for the short-term stations and partial-record sites include the 99-, 98-, 97-, 95-, 90-, and 85-percent flow durations; the 7-day, 10-year (7Q10) low flow; 7-day, 2-year (7Q2) low flow; and 30-day, 2-year (30Q2) low-flow frequencies; and the August median flow. Descriptive information on location and record length, measured basin characteristics, index stations correlated to the short-term station and partial-record sites, and estimated flow statistics are provided in this report for each station. Streamflow estimates from this study are stored on USGS's World Wide Web application 'StreamStats' (http://water.usgs.gov/osw/streamstats/connecticut.html).

Scientific Investigations Report

Estimating concentrations of fine-grained and total suspended sediment from close-range remote sensing imagery

Fluvial sediment, a vital surface water resource, is hazardous in excess. Suspended sediment, the most prevalent source of impairment of river systems, can adversely affect flood control, navigation, fisheries and aquatic ecosystems, recreation, and water supply (e.g., Rasmussen et al., 2009; Qu, 2014). Monitoring programs typically focus on suspended-sediment concentration (SSC) and discharge (SSQ). These time-series data are used to study changes to basin hydrology, geomorphology, and ecology caused by disturbances. The U.S. Geological Survey (USGS) has traditionally used physical sediment sample-based methods (Edwards and Glysson, 1999; Nolan et al., 2005; Gray et al., 2008) to compute SSC and SSQ from continuous streamflow data using a sediment transport-curve (e.g., Walling, 1977) or hydrologic interpretation (Porterfield, 1972). Accuracy of these data is typically constrained by the resources required to collect and analyze intermittent physical samples. Quantifying SSC using continuous instream turbidity is rapidly becoming common practice among sediment monitoring programs. Estimations of SSC and SSQ are modeled from linear regression analysis of concurrent turbidity and physical samples. Sediment-surrogate technologies such as turbidity promise near real-time information, increased accuracy, and reduced cost compared to traditional physical sample-based methods (Walling, 1977; Uhrich and Bragg, 2003; Gray and Gartner, 2009; Rasmussen et al., 2009; Landers et al., 2012; Landers and Sturm, 2013; Uhrich et al., 2014). Statistical comparisons among SSQ computation methods show that turbidity-SSC regression models can have much less uncertainty than streamflow-based sediment transport-curves or hydrologic interpretation (Walling, 1977; Lewis, 1996; Glysson et al., 2001; Lee et al., 2008). However, computation of SSC and SSQ records from continuous instream turbidity data is not without challenges; some of these include environmental fouling, calibration, and data range among sensors. Of greatest interest to many programs is a hysteresis in the relationship between turbidity and SSC, attributed to temporal variation of particle size distribution (Landers and Sturm, 2013; Uhrich et al., 2014). This phenomenon causes increased uncertainty in regression-estimated values of SSC, due to changes in nephelometric reflectance off the varying grain sizes in suspension (Uhrich et al., 2014). Here, we assess the feasibility and application of close-range remote sensing to quantify SSC and particle size distribution of a disturbed, and highly-turbid, river system. We use a consumer-grade digital camera to acquire imagery of the river surface and a depth-integrating sampler to collect concurrent suspended-sediment samples. We then develop two empirical linear regression models to relate image spectral information to concentrations of fine sediment (clay to silt) and total suspended sediment. Before presenting our regression model development, we briefly summarize each data-acquisition method.

Conference Paper

A critique of the use of indicator-species scores for identifying thresholds in species responses

Identification of ecological thresholds is important both for theoretical and applied ecology. Recently, Baker and King (2010, King and Baker 2010) proposed a method, threshold indicator analysis (TITAN), to calculate species and community thresholds based on indicator species scores adapted from Dufrêne and Legendre (1997). We tested the ability of TITAN to detect thresholds using models with (broken-stick, disjointed broken-stick, dose-response, step-function, Gaussian) and without (linear) definitive thresholds. TITAN accurately and consistently detected thresholds in step-function models, but not in models characterized by abrupt changes in response slopes or response direction. Threshold detection in TITAN was very sensitive to the distribution of 0 values, which caused TITAN to identify thresholds associated with relatively small differences in the distribution of 0 values while ignoring thresholds associated with large changes in abundance. Threshold identification and tests of statistical significance were based on the same data permutations resulting in inflated estimates of statistical significance. Application of bootstrapping to the split-point problem that underlies TITAN led to underestimates of the confidence intervals of thresholds. Bias in the derivation of the z-scores used to identify TITAN thresholds and skewedness in the distribution of data along the gradient produced TITAN thresholds that were much more similar than the actual thresholds. This tendency may account for the synchronicity of thresholds reported in TITAN analyses. The thresholds identified by TITAN represented disparate characteristics of species responses that, when coupled with the inability of TITAN to identify thresholds accurately and consistently, does not support the aggregation of individual species thresholds into a community threshold.

Freshwater Science

Application of empirical predictive modeling using conventional and alternative fecal indicator bacteria in eastern North Carolina waters

Coastal and estuarine waters are the site of intense anthropogenic influence with concomitant use for recreation and seafood harvesting. Therefore, coastal and estuarine water quality has a direct impact on human health. In eastern North Carolina (NC) there are over 240 recreational and 1025 shellfish harvesting water quality monitoring sites that are regularly assessed. Because of the large number of sites, sampling frequency is often only on a weekly basis. This frequency, along with an 18–24 h incubation time for fecal indicator bacteria (FIB) enumeration via culture-based methods, reduces the efficiency of the public notification process. In states like NC where beach monitoring resources are limited but historical data are plentiful, predictive models may offer an improvement for monitoring and notification by providing real-time FIB estimates. In this study, water samples were collected during 12 dry (n = 88) and 13 wet (n = 66) weather events at up to 10 sites. Statistical predictive models for Escherichiacoli (EC), enterococci (ENT), and members of the Bacteroidales group were created and subsequently validated. Our results showed that models for EC and ENT (adjusted R2 were 0.61 and 0.64, respectively) incorporated a range of antecedent rainfall, climate, and environmental variables. The most important variables for EC and ENT models were 5-day antecedent rainfall, dissolved oxygen, and salinity. These models successfully predicted FIB levels over a wide range of conditions with a 3% (EC model) and 9% (ENT model) overall error rate for recreational threshold values and a 0% (EC model) overall error rate for shellfish threshold values. Though modeling of members of the Bacteroidales group had less predictive ability (adjusted R 2 were 0.56 and 0.53 for fecal Bacteroides spp. and human Bacteroides spp., respectively), the modeling approach and testing provided information on Bacteroidales ecology. This is the first example of a set of successful statistical predictive models appropriate for assessment of both recreational and shellfish harvesting water quality in estuarine waters.

North Carolina

A multispecies statistical age-structured model to assess predator-prey balance: application to an intensively managed Lake Michigan pelagic fish community

Using a Bayesian model fitting approach, we developed a multispecies statistical catch-at-age model to assess trade-offs between predatory demands and prey productivities, focusing on the Lake Michigan pelagic fish community. We assessed these trade-offs in terms of predation mortalities and productivities of alewife (Alosa pseudoharengus) and rainbow smelt (Osmerus mordax) and functional responses of salmonines. Our predation mortality estimates suggest that salmonine consumption has been a major driver of historical fluctuations in prey abundance, with sharp declines in alewife abundance in the 1980s and 2000s coinciding with estimated increases in predation mortalities. While Chinook salmon (Oncorhynchus tshawytscha) were food limited during periods of low alewife abundance, other salmonines appeared to maintain a (near) maximum per-predator consumption across all observed prey densities, suggesting that feedback mechanisms are unlikely to help maintain a balance between predator consumption and prey productivity in Lake Michigan. This study demonstrates that a multispecies modeling approach that combines stock assessment methods with explicit consideration of predator–prey interactions could provide the basis for tactical decision-making from a broader ecosystem perspective.

Lake Michigan