Search USGSSearch

SEARCH · Search USGS

Results for “Applied Stochastic Models and Data Analysis”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

18 recordsLinked to original sources

The use of analysis of variance procedures in biological studies

The analysis of variance (ANOVA) is widely used in biological studies, yet there remains considerable confusion among researchers about the interpretation of hypotheses being tested. Ambiguities arise when statistical designs are unbalanced, and in particular when not all combinations of design factors are represented in the data. This paper clarifies the relationship among hypothesis testing, statistical modelling and computing procedures in ANOVA for unbalanced data. A simple two-factor fixed effects design is used to illustrate three common parametrizations for ANOVA models, and some associations among these parametrizations are developed. Biologically meaningful hypotheses for main effects and interactions are given in terms of each parametrization, and procedures for testing the hypotheses are described. The standard statistical computing procedures in ANOVA are given along with their corresponding hypotheses. Throughout the development unbalanced designs are assumed and attention is given to problems that arise with missing cells.

Applied Stochastic Models and Data Analysis

MARKOV: A methodology for the solution of infinite time horizon MARKOV decision processes

Algorithms are described for determining optimal policies for finite state, finite action, infinite discrete time horizon Markov decision processes. Both value-improvement and policy-improvement techniques are used in the algorithms. Computing procedures are also described. The algorithms are appropriate for processes that are either finite or infinite, deterministic or stochastic, discounted or undiscounted, in any meaningful combination of these features. Computing procedures are described in terms of initial data processing, bound improvements, process reduction, and testing and solution. Application of the methodology is illustrated with an example involving natural resource management. Management implications of certain hypothesized relationships between mallard survival and harvest rates are addressed by applying the optimality procedures to mallard population models.

Applied Stochastic Models and Data Analysis

Developing population models with data from marked individuals

Population viability analysis (PVA) is a powerful tool for biodiversity assessments, but its use has been limited because of the requirements for fully specified population models such as demographic structure, density-dependence, environmental stochasticity, and specification of uncertainties. Developing a fully specified population model from commonly available data sources – notably, mark–recapture studies – remains complicated due to lack of practical methods for estimating fecundity, true survival (as opposed to apparent survival), natural temporal variability in both survival and fecundity, density-dependence in the demographic parameters, and uncertainty in model parameters. We present a general method that estimates all the key parameters required to specify a stochastic, matrix-based population model, constructed using a long-term mark–recapture dataset. Unlike standard mark–recapture analyses, our approach provides estimates of true survival rates and fecundities, their respective natural temporal variabilities, and density-dependence functions, making it possible to construct a population model for long-term projection of population dynamics. Furthermore, our method includes a formal quantification of parameter uncertainty for global (multivariate) sensitivity analysis. We apply this approach to 9 bird species and demonstrate the feasibility of using data from the Monitoring Avian Productivity and Survivorship (MAPS) program. Bias-correction factors for raw estimates of survival and fecundity derived from mark–recapture data (apparent survival and juvenile:adult ratio, respectively) were non-negligible, and corrected parameters were generally more biologically reasonable than their uncorrected counterparts. Our method allows the development of fully specified stochastic population models using a single, widely available data source, substantially reducing the barriers that have until now limited the widespread application of PVA. This method is expected to greatly enhance our understanding of the processes underlying population dynamics and our ability to analyze viability and project trends for species of conservation concern.

Biological Conservation

A statistical forecasting approach to metapopulation viability analysis

Conservation of at‐risk species is aided by reliable forecasts of the consequences of environmental change and management actions on population viability. Forecasts from conventional population viability analysis (PVA) are made using a two‐step procedure in which parameters are estimated, or elicited from expert opinion, and then plugged into a stochastic population model without accounting for parameter uncertainty. Recently developed statistical PVAs differ because forecasts are made conditional on models fitted to empirical data. The statistical forecasting approach allows for uncertainty about parameters, but it has rarely been applied in metapopulation contexts where spatially explicit inference is needed about colonization and extinction dynamics and other forms of stochasticity that influence metapopulation viability. We conducted a statistical metapopulation viability analysis (MPVA) using 11 yr of data on the federally threatened Chiricahua leopard frog ( Lithobates chiricahuensis ) to forecast responses to landscape heterogeneity, drought, environmental stochasticity, and management. We evaluated several future environmental scenarios and pond restoration options designed to reduce extinction risk. Forecasts over a 50‐yr time horizon indicated that metapopulation extinction risk was <4% for all scenarios, but uncertainty was high. Without pond restoration, extinction risk is forecasted to be 3.9% (95% CI 0–37%) by year 2066. Restoring six ponds by increasing their hydroperiod reduced extinction risk to <1% and greatly reduced uncertainty (95% CI 0–2%). Our results suggest that managers can mitigate the impacts of drought and environmental stochasticity on metapopulation viability by maintaining ponds that hold water throughout the year and keeping them free of invasive predators. Our study illustrates the utility of the spatially explicit statistical forecasting approach to MPVA in conservation planning efforts.

Ecological Applications

Assessing roadway contributions to stormwater flows, concentrations, and loads with the StreamStats application

The Oregon Department of Transportation (ODOT) and other state departments of transportation need quantitative information about the percentages of different land cover categories above any given stream crossing in the state to assess and address roadway contributions to water-quality impairments and resulting total maximum daily loads. The U.S. Geological Survey, in cooperation with ODOT and the FHWA, added roadway and land cover information to the online StreamStats application to facilitate analysis of stormwater runoff contributions from different land covers. Analysis of 25 delineated basins with drainage areas of about 100 mi2 indicates the diversity of land covers in the Willamette Valley, Oregon. On average, agricultural, developed, and undeveloped land covers comprise 15%, 2.3%, and 82% of these basin areas. On average, these basins contained about 10 mi of state highways and 222 mi of non-state roads. The Stochastic Empirical Loading and Dilution Model was used with available water-quality data to simulate long-term yields of total phosphorus from highways, non-highway roadways, and agricultural, developed, and undeveloped areas. These yields were applied to land cover areas obtained from StreamStats for the Willamette River above Wilsonville, Oregon. This analysis indicated that highway yields were larger than yields from other land covers because highway runoff concentrations were higher than other land covers and the highway is fully impervious. However, the total highway area was a fraction of the other land covers. Accordingly, highway runoff mitigation measures can be effective for managing water quality locally, they may have limited effect on achieving basin-wide stormwater reduction goals.

Transportation Research Record

Longitudinal analysis of bioaccumulative contaminants in freshwater fishes

The National Contaminant Biomonitoring Program (NCBP) was initiated in 1967 as a component of the National Pesticide Monitoring program. It consists of periodic collection of freshwater fish and other samples and the analysis of the concentrations of persistent environmental contaminants in these samples. For the analysis, the common approach has been to apply the mixed two-way ANOVA model to combined data. A main disadvantage of this method is that it cannot give a detailed temporal trend of the concentrations since the data are grouped. In this paper, we present an alternative approach that performs a longitudinal analysis of the information using random effects models. In the new approach, no grouping is needed and the data are treated as samples from continuous stochastic processes, which seems more appropriate than ANOVA for the problem.

Environmental and Ecological Statistics

Modeling unobserved sources of heterogeneity in animal abundance using a Dirichlet process prior

In surveys of natural populations of animals, a sampling protocol is often spatially replicated to collect a representative sample of the population. In these surveys, differences in abundance of animals among sample locations may induce spatial heterogeneity in the counts associated with a particular sampling protocol. For some species, the sources of heterogeneity in abundance may be unknown or unmeasurable, leading one to specify the variation in abundance among sample locations stochastically. However, choosing a parametric model for the distribution of unmeasured heterogeneity is potentially subject to error and can have profound effects on predictions of abundance at unsampled locations. In this article, we develop an alternative approach wherein a Dirichlet process prior is assumed for the distribution of latent abundances. This approach allows for uncertainty in model specification and for natural clustering in the distribution of abundances in a data-adaptive way. We apply this approach in an analysis of counts based on removal samples of an endangered fish species, the Okaloosa darter. Results of our data analysis and simulation studies suggest that our implementation of the Dirichlet process prior has several attractive features not shared by conventional, fully parametric alternatives. ?? 2008, The International Biometric Society.

Biometrics

Developing species-age cohorts from forest inventory and analysis data to parameterize a forest landscape model

Simulating long-term, landscape level changes in forest composition requires estimates of stand age to initialize succession models. Detailed stand ages are rarely available, and even general information on stand history often is lacking. We used data from USDA Forest Service Forest Inventory and Analysis (FIA) database to estimate broad age classes for a forested landscape to simulate changes in landscape composition and structure relative to climate change at Fort Drum, a 43,000 ha U.S. Army installation in northwestern New York. Using simple linear regression, we developed relationships between tree diameter and age for FIA site trees from the host and adjacent ecoregions and applied those relationships to forest stands at Fort Drum. We observed that approximately half of the variation in age was explained by diameter breast height (DBH) across all species studied ( r 2 = 0.42 for sugar maple Acer saccharum to 0.63 for white ash Fraxinus americana ). We then used age-diameter relationships from published research on northern hardwood species to calibrate results from the FIA-based analysis. With predicted stand age, we used tree species life histories and environmental conditions represented by ecological site types to parameterize a stochastic forest landscape model (LANDIS-II) to spatially and temporally model successional changes in forest communities at Fort Drum. Forest stands modeled over 100 years without significant disturbance appeared to reflect expected patterns of increasing dominance by shade-tolerant mesophytic tree species such as sugar maple, red maple ( Acer rubrum ), and eastern hemlock ( Tsuga canadensis ) where soil moisture was sufficient. On drier sandy soils, eastern white pine ( Pinus strobus ), red pine ( P. resinosa ), northern red oak ( Quercus rubra ), and white oak ( Q. alba ) continued to be important components throughout the modeling period with no net loss at the landscape scale. Our results suggest that despite abundant precipitation and relatively low evapotranspiration rates for the region, low soil water holding capacity and fertility may be limiting factors for the spread of mesophytic species on excessively drained soils in the region. Increasing atmospheric temperatures projected for the region could alter moisture regimes for many coarse-textured soils providing a possible mechanism for expansion of xerophytic tree species.

New York

Mapping of coal quality using stochastic simulation and isometric logratio transformation with an application to a Texas lignite

Coal is a chemically complex commodity that often contains most of the natural elements in the periodic table. Coal constituents are conventionally grouped into four components (proximate analysis): fixed carbon, ash, inherent moisture, and volatile matter. These four parts, customarily measured as weight losses and expressed as percentages, share all properties and statistical challenges of compositional data. Consequently, adequate modeling should be done in terms of a logratio transformation, a requirement that is commonly overlooked by modelers. The transformation of choice is the isometric logratio transformation because of its geometrical and statistical advantages. The modeling is done through a series of realizations prepared by applying sequential simulation for the purpose of displaying the parts in maps incorporating uncertainty. The approach makes realistic assumptions and the results honor the data and basic considerations, such as percentages between 0 and 100, all four parts adding to 100% at any location in the study area, and a style of spatial fluctuation in the realizations equal to that of the data. The realizations are used to prepare different results, including probability distributions across a deposit, E-type maps displaying average properties, and probability maps summarizing joint fluctuations of several parts. Application of these maps to a lignite bed clearly delineates the deposit boundary, reveals a channel cutting across, and shows that the most favorable coal quality is to the north and deteriorates toward the southeast.

International Journal of Coal Geology

A theory for modeling ground-water flow in heterogeneous media

Construction of a ground-water model for a field area is not a straightforward process. Data are virtually never complete or detailed enough to allow substitution into the model equations and direct computation of the results of interest. Formal model calibration through optimization, statistical, and geostatistical methods is being applied to an increasing extent to deal with this problem and provide for quantitative evaluation and uncertainty analysis of the model. However, these approaches are hampered by two pervasive problems: 1) nonlinearity of the solution of the model equations with respect to some of the model (or hydrogeologic) input variables (termed in this report system characteristics) and 2) detailed and generally unknown spatial variability (heterogeneity) of some of the system characteristics such as log hydraulic conductivity, specific storage, recharge and discharge, and boundary conditions. A theory is developed in this report to address these problems. The theory allows construction and analysis of a ground-water model of flow (and, by extension, transport) in heterogeneous media using a small number of lumped or smoothed system characteristics (termed parameters). The theory fully addresses both nonlinearity and heterogeneity in such a way that the parameters are not assumed to be effective values. The ground-water flow system is assumed to be adequately characterized by a set of spatially and temporally distributed discrete values, ?, of the system characteristics. This set contains both small-scale variability that cannot be described in a model and large-scale variability that can. The spatial and temporal variability in ? are accounted for by imagining ? to be generated by a stochastic process wherein ? is normally distributed, although normality is not essential. Because ? has too large a dimension to be estimated using the data normally available, for modeling purposes ? is replaced by a smoothed or lumped approximation y?. (where y is a spatial and temporal interpolation matrix). Set y?. has the same form as the expected value of ?, y 'line' ? , where 'line' ? is the set of drift parameters of the stochastic process; ?. is a best-fit vector to ?. A model function f(?), such as a computed hydraulic head or flux, is assumed to accurately represent an actual field quantity, but the same function written using y?., f(y?.), contains error from lumping or smoothing of ? using y?.. Thus, the replacement of ? by y?. yields nonzero mean model errors of the form E(f(?)-f(y?.)) throughout the model and covariances between model errors at points throughout the model. These nonzero means and covariances are evaluated through third and fifth-order accuracy, respectively, using Taylor series expansions. They can have a significant effect on construction and interpretation of a model that is calibrated by estimating ?.. Vector ?.. is estimated as 'hat' ? using weighted nonlinear least squares techniques to fit a set of model functions f(y'hat' ?) to a. corresponding set of observations of f(?), Y. These observations are assumed to be corrupted by zero-mean, normally distributed observation errors, although, as for ?, normality is not essential. An analytical approximation of the nonlinear least squares solution is obtained using Taylor series expansions and perturbation techniques that assume model and observation errors to be small. This solution is used to evaluate biases and other results to second-order accuracy in the errors. The correct weight matrix to use in the analysis is shown to be the inverse of the second-moment matrix E(Y-f(y?.))(Y-f(y?.))', but the weight matrix is assumed to be arbitrary in most developments. The best diagonal approximation is the inverse of the matrix of diagonal elements of E(Y-f(y?.))(Y-f(y?.))', and a method of estimating this diagonal matrix when it is unknown is developed using a special objective function to compute 'hat' ?. When considered to be an estimate of f

Professional Paper

Estimating age from recapture data: Integrating incremental growth measures with ancillary data to infer age-at-length

Estimating the age of individuals in wild populations can be of fundamental importance for answering ecological questions, modeling population demographics, and managing exploited or threatened species. Significant effort has been devoted to determining age through the use of growth annuli, secondary physical characteristics related to age, and growth models. Many species, however, either do not exhibit physical characteristics useful for independent age validation or are too rare to justify sacrificing a large number of individuals to establish the relationship between size and age. Length‐at‐age models are well represented in the fisheries and other wildlife management literature. Many of these models overlook variation in growth rates of individuals and consider growth parameters as population parameters. More recent models have taken advantage of hierarchical structuring of parameters and Bayesian inference methods to allow for variation among individuals as functions of environmental covariates or individual‐specific random effects. Here, we describe hierarchical models in which growth curves vary as individual‐specific stochastic processes, and we show how these models can be fit using capture–recapture data for animals of unknown age along with data for animals of known age. We combine these independent data sources in a Bayesian analysis, distinguishing natural variation (among and within individuals) from measurement error. We illustrate using data for African dwarf crocodiles, comparing von Bertalanffy and logistic growth models. The analysis provides the means of predicting crocodile age, given a single measurement of head length. The von Bertalanffy was much better supported than the logistic growth model and predicted that dwarf crocodiles grow from 19.4 cm total length at birth to 32.9 cm in the first year and 45.3 cm by the end of their second year. Based on the minimum size of females observed with hatchlings, reproductive maturity was estimated to be at nine years. These size benchmarks are believed to represent thresholds for important demographic parameters; improved estimates of age, therefore, will increase the precision of population projection models. The modeling approach that we present can be applied to other species and offers significant advantages when multiple sources of data are available and traditional aging techniques are not practical.

Loango National Park

Statistical relations among earthquake magnitude, surface rupture length, and surface fault displacement

In order to refine correlations of surface-wave magnitude, fault rupture length at the ground surface, and fault displacement at the surface by including the uncertainties in these variables, the existing data were critically reviewed and a new data base was compiled. Earthquake magnitudes were redetermined as necessary to make them as consistent as possible with the Gutenberg methods and results, which make up much of the data base. Measurement errors were estimated for the three variables for 58 moderate to large shallow-focus earthquakes. Regression analyses were then made utilizing the estimated measurement errors. The regression analysis demonstrates that the relations among the variables magnitude, length, and displacement are stochastic in nature. The stochastic variance, introduced in part by incomplete surface expression of seismogenic faulting, variation in shear modulus, and regional factors, dominates the estimated measurement errors. Thus, it is appropriate to use ordinary least squares for the regression models, rather than regression models based upon an underlying deterministic relation in which the variance results primarily from measurement errors. Significant differences exist in correlations of certain combinations of length, displacement, and magnitude when events are grouped by fault type or by region, including attenuation regions delineated by Evernden and others. Estimates of the magnitude and the standard deviation of the magnitude of a prehistoric or future earthquake associated with a fault can be made by correlating M s with the logarithms of rupture length, fault displacement, or the product of length and displacement. Fault rupture area could be reliably estimated for about 20 of the events in the data set. Regression of M s on rupture area did not result in a marked improvement over regressions that did not involve rupture area. Because no subduction-zone earthquakes are included in this study, the reported results do not apply to such zones.

Bulletin of the Seismological Society of America

Habitat suitability index model for brook trout in streams of the Southern Blue Ridge Province: Surrogate variables, model evaluation, and suggested improvements

Data from several sources were collated and analyzed by correlation, regression, and principal components analysis to define surrrogate variables for use in the brook trout (Salvelinus fontinalis) habitat suitability index (HSI) model, and to evaluate the applicability of the model for assessing habitat in high elevation streams of the southern Blue Ridge Province (SBRP). In all data sets examined, pH and alkalinity were highly correlated, and both declined with increasing elevation; however, the magnitude of the decline varied with underlying rock formations and other factors, thereby restricting the utility of elevation as a surrogate for pH. In the data sets that contained biological information, brook trout abundance (as biomass, density, or both) tended to increase with elevation and decrease with the abundance of rainbow trout (Oncorhynchus mykiss), and was not significantly correlated (P >0.05) with the abundance of most benthic macroinvertebrate taxa normally construed as important in the diet of brook trout. Using multiple linear regression, the authors formulated an alternative HSI model A? based on point estimates of gradient, pH, elevation, stream width, and rainbow trout density A? which explained 40 to 50 percent of the variance in brook trout density in 256 stream reaches. Although logically developed, the present U.S. Fish and Wildlife Service HSI model, proposed in 1982, seems deficient in several areas, especially when applied to SBRP streams. The authors recommend that the water quality component in the model be updated and reevaluated, focusing on the differential sensitivities of each life stage, the stochastic nature of the water quality variables, and the possible existence of habitat requirements that differ among brook trout strains.

Biological Report

Using population models to evaluate management alternatives for Gulf Striped Bass

Interstate management of Gulf Striped Bass Morone saxatilis has involved a thirty-year cooperative effort involving Federal and State agencies in Georgia, Florida and Alabama (Apalachicola-Chattahoochee-Flint Gulf Striped Bass Technical Committee). The Committee has recently focused on developing an adaptive framework for conserving and restoring Gulf Striped Bass in the Apalachicola, Chattahoochee, and Flint River (ACF) system. To evaluate the consequences and tradeoffs among management activities, population models were used to inform management decisions. Stochastic matrix models were constructed with varying recruitment and stocking rates to simulate effects of management alternatives on Gulf Striped Bass population objectives. An age-classified matrix model that incorporated stock fecundity estimates and survival estimates was used to project population growth rate. In addition, combinations of management alternatives (stocking rates, Hydrilla control, harvest regulations) were evaluated with respect to how they influenced Gulf Striped Bass population growth. Annual survival and mortality rates were estimated from catch-curve analysis, while fecundity was estimated and predicted using a linear least squares regression analysis of fish length versus egg number from hatchery brood fish data. Stocking rates and stocked-fish survival rates were estimated from census data. Results indicated that management alternatives could be an effective approach to increasing the Gulf Striped Bass population. Population abundance was greatest under maximum stocking effort, maximum Hydrilla control and a moratorium. Conversely, population abundance was lowest under no stocking, no Hydrilla control and the current harvest regulation. Stocking rates proved to be an effective management strategy; however, low survival estimates of stocked fish (1%) limited the potential for population growth. Hydrilla control increased the survival rate of stocked fish and provided higher estimates of population abundances than maximizing the stocking rate. A change in the current harvest regulation (50% harvest regulation) was not an effective alternative to increasing the Gulf Striped Bass population size. Applying a moratorium to the Gulf Striped Bass fishery increased survival rates from 50% to 74% and resulted in the largest population growth of the individual management alternatives. These results could be used by the Committee to inform management decisions for other populations of Striped Bass in the Gulf Region.

Cooperator Science Series

Effects of Lead Exposure, Environmental Conditions, and Metapopulation Processes on Population Dynamics of Spectacled Eiders.

Spectacled eider Somateria fischeri numbers have declined and they are considered threatened in accordance with the US Endangered Species Act throughout their range. We synthesized the available information for spectacled eiders to construct deterministic, stochastic, and metapopulation models for this species that incorporated current estimates of vital rates such as nest success, adult survival, and the impact of lead poisoning on survival. Elasticities of our deterministic models suggested that the populations would respond most dramatically to changes in adult female survival and that the reductions in adult female survival related to lead poisoning were locally important. We also examined the sensitivity of the population to changes in lead exposure rates. With the knowledge that some vital rates vary with environmental conditions, we cast stochastic models that mimicked observed variation in productivity. We also used the stochastic model to examine the probability that a specific population will persist for periods of up to 50 y. Elasticity analysis of these models was consistent with that for the deterministic models, with perturbations to adult female survival having the greatest effect on population projections. When used in single population models, demographic data for some localities predicted rapid declines that were inconsistent with our observations in the field. Thus, we constructed a metapopulation model and examined the predictions for local subpopulations and the metapopulation over a wide range of dispersal rates. Using the metapopulation model, we were able to simulate the observed stability of local subpopulations as well as that of the metapopulation. Finally, we developed a global metapopulation model that simulates periodic winter habitat limitation, similar to that which might be experienced in years of heavy sea ice in the core wintering area of spectacled eiders in the central Bering Sea. Our metapopulation analyses suggested that no subpopulation is independent and that future management actions may be improved through a metapopulation framework. For example, management actions could include displacement of breeding females from"sink" areas that reduce the growth potential of the population as a whole. However, this action is contingent upon dispersal among local populations, for which there is limited information. Thus, we recommend that researchers examine dispersal behavior among areas on the Yukon-Kuskokwim Delta in western Alaska. The metapopulation framework could also be applied at the rangewide scale to address the density-dependent limitation of available polynya habitat during winter that may limit the recovery of small subpopulations, such as that on the Yukon-Kuskokwim Delta. Reductions in other subpopulations may be necessary to ensure an increase in the Yukon-Kuskokwim Delta population. Thus, we recommend that managers consider the interpopulation dynamics of spectacled eiders at different spatial scales in future management actions.

Alaska

Building a state-space life cycle model for naturally produced Snake River fall Chinook salmon

In 1992, Snake River basin fall Chinook salmon (Oncorhynchus tshawytscha) were listed for protection under the U.S. Endangered Species Act (NMFS 1992) and the population remained below 1000 individuals until 2000. Since then, returns from natural production has rebounded to over 20,000 spawners owing to a host of factors including reduced harvest (Peters et al. 2001), stable minimum spawning flows (Groves and Chandler 1999), summer flow augmentation (Connor et al. 2003), predator control (Beamesderfer et al. 1996), hatchery supplementation (Rosenberger et al. 2017), improved juvenile passage structures (Adams et al. 2014), summer spill operations (Perry et al. 2006; Adams et al. 2008), and periods of favorable ocean conditions and food availability (Logerwell et al. 2003; Peterson et al. 2014). Given this change in abundance coincident with numerous management actions and fluctuation in environmental drivers, quantifying which factors contributed to the observed rebound in natural production can provide critical insights into future management actions for this at-risk population. Multistage life cycle models provide a powerful analytical framework for understating how each life stage of a population contributes to population growth rate (Moussalli and Hilborn 1986; Greene and Beechie 2004). Multistage models may also be used as an analytical framework to explicitly estimate demographic parameters of a population model. This approach has an advantage over single-stage stock-recruitment models by allowing population growth rates to be partitioned among life stages rather than aggregated over an entire life cycle. Such partitioning allows for estimating 1) stage-specific density dependence, and 2) stage-specific effects of environmental factors or management actions. For example, Zabel et al. (2006) estimated parameters of a multistage model used in the context of a population viability analysis for spring/summer Chinook salmon in the Snake River, but such an approach has yet to be applied to fall Chinook salmon in the Snake River basin. Typically, data informing estimates of abundance at particular “check points” in the life cycle determines the complexity of the multistage model that can be fit to the data. For fall Chinook salmon, we are developing a two-stage model that encompasses: 1) upstream passage of spawners at Lower Granite Dam (LGR) to the subsequent downstream passage of their progeny at the dam, and 2) downstream passage of juveniles at LGR to their subsequent return from the ocean and passage at the Dam 2‒6 years later. This approach partitions the life cycle of fall Chinook salmon both spatially and temporally, which allows us to fit and compare alternative models with covariates specific to each stage. Our previous report to the ISAB (Zabel et al. 2013) detailed methods for estimating abundance of naturally produced adults and juveniles passing Lower Granite Dam, which provides the requisite data for fitting a two-stage model. The intent of this report is to describe the structure of the two-stage life cycle model, present preliminary results from fitting the model to data, and outline future directions and developments. As is clear from the diversity of models presented in this report, “life cycle models” range from very simple theoretically based population models (e.g., the Beverton-Holt stock- recruitment model) to very complex spatially explicit simulation models linked to hydrosystem hydrodynamic models (e.g., the COMPASS model for a single transition in a life cycle model, Zabel et al. 2008). We chose to develop a model of intermediate complexity that casts the two- stage life cycle model in a state-space framework (Newman et al. 2014). We chose to use a state-space framework implemented in a Bayesian framework because: • It provides both a statistical estimation framework for retrospective statistical analysis and a stochastic simulation framework for prospective analysis to evaluate alternative management actions. • Abundance estimates are uncertain. A state-space framework accounts for observation uncertainty in the abundance estimates and other data (e.g., age structure) while simultaneously estimating process uncertainty. • It allows for missing data. By drawing missing data from an appropriate probability model, uncertainty owing to missing data can be propagated without having to omit data or assume fixed values for missing data. Thus, a two-stage state-space life cycle model for fall Chinook salmon strikes an appropriate balance between model complexity, tractability, and applicability given the goals of performing both retrospective and prospective analysis to guide future management of this population.

Idaho, Oregon, Washington, Wyoming

A practical primer on geostatistics

Introduction The Challenge —Most geological phenomena are extraordinarily complex in their interrelationships and vast in their geographical extension. Ordinarily, engineers and geoscientists are faced with corporate or scientific requirements to properly prepare geological models with measurements involving a small fraction of the entire area or volume of interest. Exact description of a system such as an oil reservoir is neither feasible nor economically possible. The results are necessarily uncertain. Note that the uncertainty is not an intrinsic property of the systems; it is the result of incomplete knowledge by the observer. The Aim of Geostatistics —The main objective of geostatistics is the characterization of spatial systems that are incompletely known, systems that are common in geology. A key difference from classical statistics is that geostatistics uses the sampling location of every measurement. Unless the measurements show spatial correlation, the application of geostatistics is pointless. Ordinarily the need for additional knowledge goes beyond a few points, which explains the display of results graphically as fishnet plots, block diagrams, and maps. Geostatistical Methods —Geostatistics is a collection of numerical techniques for the characterization of spatial attributes using primarily two tools: probabilistic models, which are used for spatial data in a manner similar to the way in which time-series analysis characterizes temporal data, or pattern recognition techniques. The probabilistic models are used as a way to handle uncertainty in results away from sampling locations, making a radical departure from alternative approaches like inverse distance estimation methods. Differences with Time Series —On dealing with time-series analysis, users frequently concentrate their attention on extrapolations for making forecasts. Although users of geostatistics may be interested in extrapolation, the methods work at their best interpolating. This simple difference has significant methodological implications. Historical Remarks —As a discipline, geostatistics was firmly established in the 1960s by the French engineer Georges Matheron, who was interested in the appraisal of ore reserves in mining. Geostatistics did not develop overnight. Like other disciplines, it has built on previous results, many of which were formulated with different objectives in various fields. Pioneers —Seminal ideas conceptually related to what today we call geostatistics or spatial statistics are found in the work of several pioneers, including: 1940s: A.N. Kolmogorov in turbulent flow and N. Wiener in stochastic processing; 1950s: D. Krige in mining; 1960s: B. Mathern in forestry and L.S. Gandin in meteorology Calculations —Serious applications of geostatistics require the use of digital computers. Although for most geostatistical techniques rudimentary implementation from scratch is fairly straightforward, coding programs from scratch is recommended only as part of a practice that may help users to gain a better grasp of the formulations. Software —For professional work, the reader should employ software packages that have been thoroughly tested to handle any sampling scheme, that run as efficiently as possible, and that offer graphic capabilities for the analysis and display of results. This primer employs primarily the package Stanford Geomodeling Software (SGeMS) - recently developed at the Energy Resources Engineering Department at Stanford University - as a way to show how to obtain results practically. This applied side of the primer should not be interpreted as the notes being a manual for the use of SGeMS. The main objective of the primer is to help the reader gain an understanding of the fundamental concepts and tools in geostatistics. Organization of the Primer —The chapters of greatest importance are those covering kriging and simulation. All other materials are peripheral and are included for better comprehension of these main geostatistical modeling tools. The choice of kriging versus simulation is often a big puzzle to the uninitiated, let alone the different variants of both of them. Chapters 14, 18, and 19 are intended to shed light on those subjects. The critical aspect of assessing and modeling spatial correlation is covered in chapter 7. Chapters 2 and 3 review relevant concepts in classical statistics. Course Objectives —This course offers stochastic solutions to common problems in the characterization of complex geological systems. At the end of the course, participants should have: an understanding of the theoretical foundations of geostatistics; a good grasp of its possibilities and limitations; and reasonable familiarity with the SGeMS software, thus opening the possibility of practically applying geostatistics.

Open-File Report

Statistics for stochastic modeling of volume reduction, hydrograph extension, and water-quality treatment by structural stormwater runoff best management practices (BMPs)

The U.S. Geological Survey (USGS) developed the Stochastic Empirical Loading and Dilution Model (SELDM) in cooperation with the Federal Highway Administration (FHWA) to indicate the risk for stormwater concentrations, flows, and loads to be above user-selected water-quality goals and the potential effectiveness of mitigation measures to reduce such risks. SELDM models the potential effect of mitigation measures by using Monte Carlo methods with statistics that approximate the net effects of structural and nonstructural best management practices (BMPs). In this report, structural BMPs are defined as the components of the drainage pathway between the source of runoff and a stormwater discharge location that affect the volume, timing, or quality of runoff. SELDM uses a simple stochastic statistical model of BMP performance to develop planning-level estimates of runoff-event characteristics. This statistical approach can be used to represent a single BMP or an assemblage of BMPs. The SELDM BMP-treatment module has provisions for stochastic modeling of three stormwater treatments: volume reduction, hydrograph extension, and water-quality treatment. In SELDM, these three treatment variables are modeled by using the trapezoidal distribution and the rank correlation with the associated highway-runoff variables. This report describes methods for calculating the trapezoidal-distribution statistics and rank correlation coefficients for stochastic modeling of volume reduction, hydrograph extension, and water-quality treatment by structural stormwater BMPs and provides the calculated values for these variables. This report also provides robust methods for estimating the minimum irreducible concentration (MIC), which is the lowest expected effluent concentration from a particular BMP site or a class of BMPs. These statistics are different from the statistics commonly used to characterize or compare BMPs. They are designed to provide a stochastic transfer function to approximate the quantity, duration, and quality of BMP effluent given the associated inflow values for a population of storm events. A database application and several spreadsheet tools are included in the digital media accompanying this report for further documentation of methods and for future use. In this study, analyses were done with data extracted from a modified copy of the January 2012 version of International Stormwater Best Management Practices Database, designated herein as the January 2012a version. Statistics for volume reduction, hydrograph extension, and water-quality treatment were developed with selected data. Sufficient data were available to estimate statistics for 5 to 10 BMP categories by using data from 40 to more than 165 monitoring sites. Water-quality treatment statistics were developed for 13 runoff-quality constituents commonly measured in highway and urban runoff studies including turbidity, sediment and solids; nutrients; total metals; organic carbon; and fecal coliforms. The medians of the best-fit statistics for each category were selected to construct generalized cumulative distribution functions for the three treatment variables. For volume reduction and hydrograph extension, interpretation of available data indicates that selection of a Spearman’s rho value that is the average of the median and maximum values for the BMP category may help generate realistic simulation results in SELDM. The median rho value may be selected to help generate realistic simulation results for water-quality treatment variables. MIC statistics were developed for 12 runoff-quality constituents commonly measured in highway and urban runoff studies by using data from 11 BMP categories and more than 167 monitoring sites. Four statistical techniques were applied for estimating MIC values with monitoring data from each site. These techniques produce a range of lower-bound estimates for each site. Four MIC estimators are proposed as alternatives for selecting a value from among the estimates from multiple sites. Correlation analysis indicates that the MIC estimates from multiple sites were weakly correlated with the geometric mean of inflow values, which indicates that there may be a qualitative or semiquantitative link between the inflow quality and the MIC. Correlations probably are weak because the MIC is influenced by the inflow water quality and the capability of each individual BMP site to reduce inflow concentrations.

Scientific Investigations Report