Search USGSSearch

SEARCH · Search USGS

Results for “Journal of Applied Statistics”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Importance of sampling design and analysis in animal population studies: a comment on Sergio et al

1. The use of predators as indicators and umbrellas in conservation has been criticized. In the Trentino region, Sergio et al. (2006; hereafter SEA) counted almost twice as many bird species in quadrats located in raptor territories than in controls. However, SEA detected astonishingly few species. We used contemporary Swiss Breeding Bird Survey data from an adjacent region and a novel statistical model that corrects for overlooked species to estimate the expected number of bird species per quadrat in that region. 2. There are two anomalies in SEA which render their results ambiguous. First, SEA detected on average only 6.8 species, whereas a value of 32 might be expected. Hence, they probably overlooked almost 80% of all species. Secondly, the precision of their mean species counts was greater in two-thirds of cases than in the unlikely case that all quadrats harboured exactly the same number of equally detectable species. This suggests that they detected consistently only a biased, unrepresentative subset of species. 3. Conceptually, expected species counts are the product of true species number and species detectability p. Plenty of factors may affect p, including date, hour, observer, previous knowledge of a site and mobbing behaviour of passerines in the presence of predators. Such differences in p between raptor and control quadrats could have easily created the observed effects. Without a method that corrects for such biases, or without quantitative evidence that species detectability was indeed similar between raptor and control quadrats, the meaning of SEA's counts is hard to evaluate. Therefore, the evidence presented by SEA in favour of raptors as indicator species for enhanced levels of biodiversity remains inconclusive. 4. Synthesis and application. Ecologists should pay greater attention to sampling design and analysis in animal population estimation. Species richness estimation means sampling a community. Samples should be representative for the community studied and the sampling fraction among communities compared should be the same on average, otherwise formal estimation approaches must be applied to avoid misleading inference.

Journal of Applied Ecology

The joint effect of changes in urbanization and climate on trends in floods: A comparison of panel and single-station quantile regression approaches

Estimates of annual maximum (peak) flow quantiles are needed for basins undergoing changes in both urbanization and climate. Most previous work on the effect of urbanization on peak flows has considered urbanization alone and only the spatial variation in flood quantiles or its mean temporal effect, and most work on the effect of nonstationarity in climate has focused on single-station analyses, which give uncertain results for extreme quantiles. To address these gaps, three approaches to the statistical estimation of the joint effects of changes in impervious cover and climate on the estimation of peak-flow quantiles were compared: single-station quantile regression; a fixed effect panel-quantile regression (pQR) method using a location (mean) shift to homogenize the panel; and a location-scale panel regression model (pQRmom), which accounts for both scale (variance) and location effects. The different approaches were applied to a dataset consisting of instantaneous annual peak flows from 127 minimally nested basins in the midwestern United States with at least 4 % change in imperviousness. The annual maximum daily discharge from a water-balance model was selected as the primary climate predictor; in addition, to provide a comparison of climate predictors, precipitation was also considered. The coefficients from single-station regressions were usually sufficiently certain to determine the effects of climate variation but usually too uncertain to estimate the effects of urbanization. The panel-quantile regression approaches give much more certain results, but their estimates of quantile dependence differ: although both indicate urbanization effects decreasing with decreasing annual exceedance probability (AEP), the pQRmom urbanization coefficients are insignificantly different from zero for AEPs less than 0.10, whereas the pQR coefficients remain positive and are significant except for AEP = 0.01, the smallest AEP value considered. Although the location-scale structure of the pQRmom approach has less flexible quantile dependence than the pQR approach, the pQRmom approach has somewhat lower overall error, and it is found that by subsetting the dataset to homogenize the scale effects, the pQR and pQRmom results become similar, indicating the insignificant urbanization coefficients for small AEPs of the pQRmom results are likely correct for the study dataset.

Arkansas, Illinois, Indiana, Iowa, Michigan, Minne

Modeling vegetation heights from high resolution stereo aerial photography: an application for broad-scale rangeland monitoring

Vertical vegetation structure in rangeland ecosystems can be a valuable indicator for assessing rangeland health and monitoring riparian areas, post-fire recovery, available forage for livestock, and wildlife habitat. Federal land management agencies are directed to monitor and manage rangelands at landscapes scales, but traditional field methods for measuring vegetation heights are often too costly and time consuming to apply at these broad scales. Most emerging remote sensing techniques capable of measuring surface and vegetation height (e.g., LiDAR or synthetic aperture radar) are often too expensive, and require specialized sensors. An alternative remote sensing approach that is potentially more practical for managers is to measure vegetation heights from digital stereo aerial photographs. As aerial photography is already commonly used for rangeland monitoring, acquiring it in stereo enables three-dimensional modeling and estimation of vegetation height. The purpose of this study was to test the feasibility and accuracy of estimating shrub heights from high-resolution (HR, 3-cm ground sampling distance) digital stereo-pair aerial images. Overlapping HR imagery was taken in March 2009 near Lake Mead, Nevada and 5-cm resolution digital surface models (DSMs) were created by photogrammetric methods (aerial triangulation, digital image matching) for twenty-six test plots. We compared the heights of individual shrubs and plot averages derived from the DSMs to field measurements. We found strong positive correlations between field and image measurements for several metrics. Individual shrub heights tended to be underestimated in the imagery, however, accuracy was higher for dense, compact shrubs compared with shrubs with thin branches. Plot averages of shrub height from DSMs were also strongly correlated to field measurements but consistently underestimated. Grasses and forbs were generally too small to be detected with the resolution of the DSMs. Estimates of vertical structure will be more accurate in plots having low herbaceous cover and high amounts of dense shrubs. Through the use of statistically derived correction factors or choosing field methods that better correlate with the imagery, vegetation heights from HR DSMs could be a valuable technique for broad-scale rangeland monitoring needs.

California;Nevada

Pore-pressure sensitivities to dynamic strains: observations in active tectonic regions

Triggered seismicity arising from dynamic stresses is often explained by the Mohr-Coulomb failure criterion, where elevated pore pressures reduce the effective strength of faults in fluid-saturated rock. The seismic response of a fluid-rock system naturally depends on its hydro-mechanical properties, but accurately assessing how pore-fluid pressure responds to applied stress over large scales in situ remains a challenging task; hence, spatial variations in response are not well understood, especially around active faults. Here I analyze previously unutilized records of dynamic strain and pore-pressure from regional and teleseismic earthquakes at Plate Boundary Observatory (PBO) stations from 2006 through 2012 to investigate variations in response along the Pacific/North American tectonic plate boundary. I find robust scaling-response coefficients between excess pore pressure and dynamic strain at each station that are spatially correlated: around the San Andreas and San Jacinto fault systems, the response is lowest in regions of the crust undergoing the highest rates of secular shear strain. PBO stations in the Parkfield instrument cluster are at comparable distances to the San Andreas fault (SAF), and spatial variations there follow patterns in dextral creep rates along the fault, with the highest response in the actively creeping section, which is consistent with a narrowing zone of strain accumulation seen in geodetic velocity profiles. At stations in the San Juan Bautista (SJB) and Anza instrument clusters, the response depends non-linearly on the inverse fault-perpendicular distance, with the response decreasing towards the fault; the SJB cluster is at the northern transition from creeping-to-locked behavior along the SAF, where creep rates are at moderate to low levels, and the Anza cluster is around the San Jacinto fault, where to date there have been no statistically significant creep rates observed at the surface. These results suggest that the strength of the pore pressure response in fluid-saturated rock near active faults is controlled by shear strain accumulation associated with tectonic loading, which implies a strong feedback between fault strength and permeability: dynamic triggering susceptibilities may vary in space and also in time.

Journal of Geophysical Research B: Solid Earth

A meta-analysis of global crop water productivity of three leading world crops (wheat, corn, and rice) in the irrigated areas over three decades

The overarching goal of this study was to perform a comprehensive meta-analysis of irrigated agricultural Crop Water Productivity (CWP) of the world’s three leading crops: wheat, corn, and rice based on three decades of remote sensing and non-remote sensing-based studies. Overall, CWP data from 148 crop growing study sites (60 wheat, 43 corn, and 45 rice) spread across the world were gathered from published articles spanning 31 different countries. There was overwhelming evidence of a significant increase in CWP with an increase in latitude for predominately northern hemisphere datasets. For example, corn grown in latitude 40–50° had much higher mean CWP (2.45 kg/m³) compared to mean CWP of corn grown in other latitudes such as 30–40° (1.67 kg/m³) or 20–30° (0.94 kg/m³). The same trend existed for wheat and rice as well. For soils, none of the CWP values, for any of the three crops, were statistically different. However, mean CWP in higher latitudes for the same soil was significantly higher than the mean CWP for the same soil in lower latitudes. This applied for all three crops studied. For wheat, the global CWP categories were low (≤0.75 kg/m³), medium (>0.75 to <1.10 kg/m³), and high CWP (≥1.10 kg/m³). For corn the global CWP categories were low (≤1.25 kg/m³), medium (>1.25 to ≤1.75 kg/m³), and high (>1.75 kg/m³). For rice the global CWP categories were low (≤0.70 kg/m³), medium (>0.70 to ≤1.25 kg/m³), and high (>1.25 kg/m³). USA and China are the only two countries that have consistently high CWP for wheat, corn, and rice. Australia and India have medium CWP for wheat and rice. India’s corn, however, has low CWP. Egypt, Turkey, Netherlands, Mexico, and Israel have high CWP for wheat. Romania, Argentina, and Hungary have high CWP for corn, and Philippines has high CWP for rice. All other countries have either low or medium CWP for all three crops. Based on data in this study, the highest consumers of water for crop production also have the most potential for water savings. These countries are USA, India, and China for wheat; USA, China, and Brazil for corn; India, China, and Pakistan for rice. For example, even just a 10% increase in CWP of wheat grown in India can save 6974 billion liters of water. This is equivalent to creating 6974 lakes each of 100 m³ in volume that leads to many benefits such as acting as ‘water banks’ for lean season, recreation, and numerous ecological services. This study establishes the volume of water that can be saved for each crop in each country when there is an increase in CWP by 10%, 20%, and 30%.

International Journal of Digital Earth

Relating net nitrogen input in the Mississippi River Basin to nitrate flux in the Lower Mississippi River--A comparison of approaches

A quantitative understanding of the relationship between terrestrial N inputs and riverine N flux can help guide conservation, policy, and adaptive management efforts aimed at preserving or restoring water quality. The objective of this study was to compare recently published approaches for relating terrestrial N inputs to the Mississippi River basin (MRB) with measured nitrate flux in the lower Mississippi River. Nitrogen inputs to and outputs from the MRB (1951 to 1996) were estimated from state-level annual agricultural production statistics and NO y (inorganic oxides of N) deposition estimates for 20 states that comprise 90% of the MRB. A model with water yield and gross N inputs accounted for 85% of the variation in observed annual nitrate flux in the lower Mississippi River, from 1960 to 1998, but tended to underestimate high nitrate flux and overestimate low nitrate flux. A model that used water yield and net anthropogenic nitrogen inputs (NANI) accounted for 95% of the variation in riverine N flux. The NANI approach accounted for N harvested in crops and assumed that crop harvest in excess of the nutritional needs of the humans and livestock in the basin would be exported from the basin. The U.S. White House Committee on Natural Resources and Environment (CENR) developed a more comprehensive N budget that included estimates of ammonia volatilization, denitrification, and exchanges with soil organic matter. The residual N in the CENR budget was weakly and negatively correlated with observed riverine nitrate flux. The CENR estimates of soil N mineralization and immobilization suggested that there were large (2000 kg N ha −1 ) net losses of soil organic N between 1951 and 1996. When the CENR N budget was modified by assuming that soil organic N levels have been relatively constant after 1950, and ammonia volatilization losses are redeposited within the basin, the trend of residual N closely matched temporal variation in NANI and was positively correlated with riverine nitrate flux in the lower Mississippi River. Based on results from applying these three modeling approaches, we conclude that although the NANI approach does not address several processes that influence the N cycle, it appears to focus on the terms that can be estimated with reasonable certainty and that are correlated with riverine N flux.

Journal of Environmental Quality

Prediction and visualization of redox conditions in the groundwater of Central Valley, California

Regional-scale, three-dimensional continuous probability models, were constructed for aspects of redox conditions in the groundwater system of the Central Valley, California. These models yield grids depicting the probability that groundwater in a particular location will have dissolved oxygen (DO) concentrations less than selected threshold values representing anoxic groundwater conditions, or will have dissolved manganese (Mn) concentrations greater than selected threshold values representing secondary drinking water-quality contaminant levels (SMCL) and health-based screening levels (HBSL). The probability models were constrained by the alluvial boundary of the Central Valley to a depth of approximately 300 m. Probability distribution grids can be extracted from the 3-D models at any desired depth, and are of interest to water-resource managers, water-quality researchers, and groundwater modelers concerned with the occurrence of natural and anthropogenic contaminants related to anoxic conditions. Models were constructed using a Boosted Regression Trees (BRT) machine learning technique that produces many trees as part of an additive model and has the ability to handle many variables, automatically incorporate interactions, and is resistant to collinearity. Machine learning methods for statistical prediction are becoming increasing popular in that they do not require assumptions associated with traditional hypothesis testing. Models were constructed using measured dissolved oxygen and manganese concentrations sampled from 2767 wells within the alluvial boundary of the Central Valley, and over 60 explanatory variables representing regional-scale soil properties, soil chemistry, land use, aquifer textures, and aquifer hydrologic properties. Models were trained on a USGS dataset of 932 wells, and evaluated on an independent hold-out dataset of 1835 wells from the California Division of Drinking Water. We used cross-validation to assess the predictive performance of models of varying complexity, as a basis for selecting final models. Trained models were applied to cross-validation testing data and a separate hold-out dataset to evaluate model predictive performance by emphasizing three model metrics of fit: Kappa; accuracy; and the area under the receiver operator characteristic curve (ROC). The final trained models were used for mapping predictions at discrete depths to a depth of 304.8 m. Trained DO and Mn models had accuracies of 86–100%, Kappa values of 0.69–0.99, and ROC values of 0.92–1.0. Model accuracies for cross-validation testing datasets were 82–95% and ROC values were 0.87–0.91, indicating good predictive performance. Kappas for the cross-validation testing dataset were 0.30–0.69, indicating fair to substantial agreement between testing observations and model predictions. Hold-out data were available for the manganese model only and indicated accuracies of 89–97%, ROC values of 0.73–0.75, and Kappa values of 0.06–0.30. The predictive performance of both the DO and Mn models was reasonable, considering all three of these fit metrics and the low percentages of low-DO and high-Mn events in the data.

California

Climate change implications for tropical islands: Interpolating and interpreting statistically downscaled GCM projections for management and planning

The potential ecological and economic effects of climate change for tropical islands were studied using output from 12 statistically downscaled general circulation models (GCMs) taking Puerto Rico as a test case. Two model selection/model averaging strategies were used: the average of all available GCMs and the average of the models that are able to reproduce the observed large-scale dynamics that control precipitation over the Caribbean. Five island-wide and multidecadal averages of daily precipitation and temperature were estimated by way of a climatology-informed interpolation of the site-specific downscaled climate model output. Annual cooling degree-days (CDD) were calculated as a proxy index for air-conditioning energy demand, and two measures of annual no-rainfall days were used as drought indices. Holdridge life zone classification was used to map the possible ecological effects of climate change. Precipitation is predicted to decline in both model ensembles, but the decrease was more severe in the “regionally consistent” models. The precipitation declines cause gradual and linear increases in drought intensity and extremes. The warming from the 1960–90 period to the 2071–99 period was 4.6°–9°C depending on the global emission scenarios and location. This warming may cause increases in CDD, and consequently increasing energy demands. Life zones may shift from wetter to drier zones with the possibility of losing most, if not all, of the subtropical rain forests and extinction risks to rain forest specialists or obligates.

Puerto Rico

Identifying sites for elk restoration in Arkansas

We used spatial data to identify potential areas for elk ( Cervus elaphus ) restoration in Arkansas. To assess habitat, we used locations of 239 elk groups collected from helicopter surveys in the Buffalo National River area of northwestern Arkansas, USA, from 1992 to 2002. We calculated the Mahalanobis distance ( D 2 ) statistic based on the relationship between those elk-group locations and a suite of 9 landscape variables to evaluate winter habitat in Arkansas. We tested model performance in the Buffalo National River area by comparing the D 2 values of pixels representing areas with and without elk pellets along 19 fixed-width transects surveyed in March 2002. Pixels with elk scat had lower D 2 values than pixels in which we found no pellets (logistic regression: Wald &chi; 2 = 24.37, P < 0.001), indicating that habitat characteristics were similar to those selected by the aerially surveyed elk. Our D 2 model indicated that the best elk habitat primarily occurred in northern and western Arkansas and was associated with areas of high landscape heterogeneity, heavy forest cover, gently sloping ridge tops and valleys, low human population density, and low road densities. To assess the potential for elk&ndash;human conflicts in Arkansas, we used the analytical hierarchy process to rank the importance of 8 criteria based on expert opinion from biologists involved in elk management. The biologists ranked availability of forage on public lands as having the strongest influence on the potential for elk&ndash;human conflict (33%), followed by human population growth rate (22%) and the amount of private land in row crops (18%). We then applied those rankings in a weighted linear summation to map the relative potential for elk&ndash;human conflict. Finally, we used white-tailed deer ( Odocoileus virginianus ) densities to identify areas where success of elk restoration may be hampered due to meningeal worm ( Parelaphostrongylus tenuis ) transmission. By combining results of the 3 spatial data layers (i.e., habitat model, elk&ndash;human conflict model, deer density), our model indicated that restoration sites located in west-central and north-central Arkansas were most favorable for reintroduction.

Arkansas

Coastal change analysis program implemented in Louisiana

Landsat Thematic Mapper images from 1990 to 1996 and collateral data sources were used to classify the land cover of the Mermentau River Basin (MRB) within the Chenier Plain of coastal Louisiana. Landcover classes followed the definition of the National Oceanic and Atmospheric Administration's Coastal Change Analysis Program; however, classification methods had to be developed as part of this study for attainment of these national classification standards. Classification method developments were especially important when classes were spectrally inseparable, when classes were part of spatial and spectral continuums, when the spatial resolution of the sensor included more than one landcover type, and when human activities caused abnormal transitions in the landscape. Most classification problems were overcome by using one or a combination of techniques, such as separating the MRB into subregions of commonality, applying masks to specific land mixtures, and highlighting class transitions between years that were highly unlikely. Overall, 1990, 1993, and 1996 classification accuracy percentages (associated kappa statistics) were 80% (0.79), 78% (0.76), and 86% (0.84), respectively. Most classification errors were associated with confusion between managed (cultivated land) and unmanaged grassland classes; scrub shrub, grasslands and forest classes; water, unconsolidated shore and bare land classes; and especially in 1993, between water and floating vegetation classes. Combining cultivated land and grassland classes and water and floating vegetation classes into single classes accuracies for 1990, 1993, and 1996 increased to 82%, 83%, and 90%, respectively. To improve the interpretation of landcover change, three indicators of landcover class stability were formulated. Location stability was defined as the percentage of a landcover class that remained as the same class in the same location at the beginning and the end of the monitoring period. Residence stability was defined as the percent change in each class within the entire MRB during the monitoring period. Turnover was defined as the addition of other landcover classes to the target landcover class during the defined monitoring period. These indicators allowed quick assessment of the dynamic nature of landcover classes, both in reference to a spatial location and to retaining their presence throughout the MRB. Examining the landcover changes between 1990 to 1993 and 1993 to 1996, led us to five principal findings: (1) Landcover turnover is maintaining a near stable logging cycle, although the locations of grassland, scrub shrub, and forest areas involved in the cycle appeared to change. (2) Planting of seedlings is critical to maintaining cycle stability. (3) Logging activities tend to replace woody land mixed forests with woody land evergreen forests. (4) Wetland estuarine marshes are expanding slightly. (5) Wetland palustrine marshes and mature forested wetlands in the MRB are relatively stable.

Louisiana

Incorporating harvest rates into the sex-age-kill model for white-tailed deer

Although monitoring population trends is an essential component of game species management, wildlife managers rarely have complete counts of abundance. Often, they rely on population models to monitor population trends. As imperfect representations of real-world populations, models must be rigorously evaluated to be applied appropriately. Previous research has evaluated population models for white-tailed deer ( Odocoileus virginianus ); however, the precision and reliability of these models when tested against empirical measures of variability and bias largely is untested. We were able to statistically evaluate the Pennsylvania sex-age-kill (PASAK) population model using realistic error measured using data from 1,131 radiocollared white-tailed deer in Pennsylvania from 2002 to 2008. We used these data and harvest data (number killed, age-sex structure, etc.) to estimate precision of abundance estimates, identify the most efficient harvest data collection with respect to precision of parameter estimates, and evaluate PASAK model robustness to violation of assumptions. Median coefficient of variation (CV) estimates by Wildlife Management Unit, 13.2% in the most recent year, were slightly above benchmarks recommended for managing game species populations. Doubling reporting rates by hunters or doubling the number of deer checked by personnel in the field reduced median CVs to recommended levels. The PASAK model was robust to errors in estimates for adult male harvest rates but was sensitive to errors in subadult male harvest rates, especially in populations with lower harvest rates. In particular, an error in subadult (1.5-yr-old) male harvest rates resulted in the opposite error in subadult male, adult female, and juvenile population estimates. Also, evidence of a greater harvest probability for subadult female deer when compared with adult (&ge;2.5-yr-old) female deer resulted in a 9.5% underestimate of the population using the PASAK model. Because obtaining appropriate sample sizes, by management unit, to estimate harvest rate parameters each year may be too expensive, assumptions of constant annual harvest rates may be necessary. However, if changes in harvest regulations or hunter behavior influence subadult male harvest rates, the PASAK model could provide an unreliable index to population changes.

Journal of Wildlife Management

A multi-state occupancy modelling framework for robust estimation of disease prevalence in multi-tissue disease systems

Given the public health, economic and conservation implications of zoonotic diseases, their effective surveillance is of paramount importance. The traditional approach to estimating pathogen prevalence as the proportion of infected individuals in the population is biased because it fails to account for imperfect detection. A statistically robust way to reduce bias in prevalence estimates is to obtain repeated samples (or sample many tissues in multi‐tissue disease systems) and to apply statistical methods that account for imperfect detection and permit the interdependence of the infection process across multiple tissues. We developed a multi‐state occupancy modelling framework which considers two scenarios about the infection process, one where no assumptions about the dependencies among the tissues are made (general), and another where dependence among tissues is not permitted (constrained). We applied this framework to pseudorabies virus (PrV) DNA detection data obtained from whole blood; and oral, nasal and genital mucosa of 510 feral swine Sus scrofa during the years 2014–2016 in Florida, USA. The constrained model was better supported by data. PrV prevalence estimates varied among tissues and were higher than the naïve estimates, ranging from to 0.06 (CI: 0.02–0.14) in genital to 0.54 (CI: 0.14 ‐ 0.82) in nasal tissue. Probability of PrV detection ranged from 0.11 (CI: 0.06–0.18) in nasal to 0.51 (CI: 0.21–0.81) in genital tissue. PrV prevalence was not affected by the age or sex of the animal or the year of sampling, but prevalence increased as drought severity increased. The conditional probability of detecting PrV given infection in at least one tissue type within an individual was highest for nasal tissue, suggesting that nasal is the best tissue to sample for PrV surveillance if only one tissue can be sampled, at least for systems with tissue‐specific prevalence and detection probabilities similar to ours. Synthesis and applications . We focused on inferences about pathogen prevalence in multi‐tissue disease systems, dealing with both nondetection and potential dependencies among tissues in infection status. We found strong evidence of variation in both prevalence and detection probabilities among tissues. Our results emphasize the importance of sampling multiple tissues and of applying inference methods that account for imperfect detection in the surveillance of systemic diseases. The multi‐state modelling framework is broadly applicable to the surveillance of pathogens that infect multiple tissues and can be used even when the infection status of the pathogen in one tissue may depend on the infection status of the pathogen in other tissue(s).

Journal of Applied Ecology

Adaptive monitoring in action: Reconsidering design-based estimators reveals underestimation of whitebark pine disease prevalence in the Greater Yellowstone Ecosystem

Identifying and understanding status and trends in ecological indicators motivates continual monitoring over decades. Many programs rely on probability surveys and their companion design-based estimators for status assessments (e.g. Horvitz–Thompson). Design-based estimators do not easily extend to trend estimation nor situations with observation errors. Field-based monitoring efforts inevitably have turnover of field crew members which may affect consistency and accuracy of data collection over time. Additionally, design-based estimators ignore the complexities of spatial and temporal heterogeneity in an ecological indicator and how this variability may be linked to environmental or biological dynamics. We propose monitoring programs should re-evaluate their prescribed statistical methods, consider model-based approaches and adapt their sampling designs as needed to improve inferences. The Greater Yellowstone Ecosystem, home to two of the most iconic U.S. National Parks, has experienced significant declines in whitebark pine Pinus albicaulis communities due to forest pathogens, insect outbreaks, wildland fires and drought. Whitebark pine is a keystone species found in mountainous environments throughout the Western U.S. and Canada. We assessed the design-based ratio estimator originally recommended for estimating prevalence of white pine blister rust Cronartium ribicola . We compared the design-based estimator to a model-based approach that accounts for the sampling design, imperfect detection and allows for infection probabilities to vary over space and time. Ignoring observation errors led to lower estimated prevalence of white pine blister rust in the general population. Using model-based approaches, we found that the probability of infection has increased since 2004. However, overall prevalence likely has not changed because of the mountain pine beetle Dendroctonus ponderosae -induced shift towards smaller diameter trees that have a lower probability of infection compared to their larger cohorts. Synthesis and Applications . Using a design-based approach to detect change in ecological indicators falls short because of the inability to account for observation errors or to explore environmental or biological factors explaining temporal dynamics. Inherently understanding the mechanisms leading to changes in an ecological indicator over time informs potential management actions. Our assessment underscores the need for continued evaluation and updating of a monitoring program's sampling design and analytical procedures to maintain relevancy.

Wyoming

Metamodeling and mapping of nitrate flux in the unsaturated zone and groundwater, Wisconsin, USA

Nitrate contamination of groundwater in agricultural areas poses a major challenge to the sustainability of water resources. Aquifer vulnerability models are useful tools that can help resource managers identify areas of concern, but quantifying nitrogen (N) inputs in such models is challenging, especially at large spatial scales. We sought to improve regional nitrate (NO 3 − ) input functions by characterizing unsaturated zone NO 3 − transport to groundwater through use of surrogate, machine-learning metamodels of a process-based N flux model. The metamodels used boosted regression trees (BRTs) to relate mappable landscape variables to parameters and outputs of a previous “vertical flux method” (VFM) applied at sampled wells in the Fox, Wolf, and Peshtigo (FWP) river basins in northeastern Wisconsin. In this context, the metamodels upscaled the VFM results throughout the region, and the VFM parameters and outputs are the metamodel response variables. The study area encompassed the domain of a detailed numerical model that provided additional predictor variables, including groundwater recharge, to the metamodels. We used a statistical learning framework to test a range of model complexities to identify suitable hyperparameters of the six BRT metamodels corresponding to each response variable of interest: NO 3 − source concentration factor (which determines the local NO 3 − input concentration); unsaturated zone travel time; NO 3 − concentration at the water table in 1980, 2000, and 2020 (three separate metamodels); and NO 3 − “extinction depth”, the eventual steady state depth of the NO 3 − front. The final metamodels were trained to 129 wells within the active numerical flow model area, and considered 58 mappable predictor variables compiled in a geographic information system (GIS). These metamodels had training and cross-validation testing R 2 values of 0.52 – 0.86 and 0.22 – 0.38, respectively, and predictions were compiled as maps of the above response variables. Testing performance was reasonable, considering that we limited the metamodel predictor variables to mappable factors as opposed to using all available VFM input variables. Relationships between metamodel predictor variables and mapped outputs were generally consistent with expectations, e.g. with greater source concentrations and NO 3 − at the groundwater table in areas of intensive crop use and well drained soils. Shorter unsaturated zone travel times in poorly drained areas likely indicated preferential flow through clay soils, and a tendency for fine grained deposits to collocate with areas of shallower water table. Numerical estimates of groundwater recharge were important in the metamodels and may have been a proxy for N input and redox conditions in the northern FWP, which had shallow predicted NO 3 − extinction depth. The metamodel results provide proof-of-concept for regional characterization of unsaturated zone NO 3 − transport processes in a statistical framework based on readily mappable GIS input variables.

Wisconsin

Optimal population prediction of sandhill crane recruitment based on climate-mediated habitat limitations

Prediction is fundamental to scientific enquiry and application; however, ecologists tend to favour explanatory modelling. We discuss a predictive modelling framework to evaluate ecological hypotheses and to explore novel/unobserved environmental scenarios to assist conservation and management decision-makers. We apply this framework to develop an optimal predictive model for juvenile (<1 year old) sandhill crane Grus canadensis recruitment of the Rocky Mountain Population (RMP). We consider spatial climate predictors motivated by hypotheses of how drought across multiple time-scales and spring/summer weather affects recruitment. Our predictive modelling framework focuses on developing a single model that includes all relevant predictor variables, regardless of collinearity. This model is then optimized for prediction by controlling model complexity using a data-driven approach that marginalizes or removes irrelevant predictors from the model. Specifically, we highlight two approaches of statistical regularization, Bayesian least absolute shrinkage and selection operator (LASSO) and ridge regression. Our optimal predictive Bayesian LASSO and ridge regression models were similar and on average 37% superior in predictive accuracy to an explanatory modelling approach. Our predictive models confirmed a priori hypotheses that drought and cold summers negatively affect juvenile recruitment in the RMP. The effects of long-term drought can be alleviated by short-term wet spring–summer months; however, the alleviation of long-term drought has a much greater positive effect on juvenile recruitment. The number of freezing days and snowpack during the summer months can also negatively affect recruitment, while spring snowpack has a positive effect. Breeding habitat, mediated through climate, is a limiting factor on population growth of sandhill cranes in the RMP, which could become more limiting with a changing climate (i.e. increased drought). These effects are likely not unique to cranes. The alteration of hydrological patterns and water levels by drought may impact many migratory, wetland nesting birds in the Rocky Mountains and beyond. Generalizable predictive models (trained by out-of-sample fit and based on ecological hypotheses) are needed by conservation and management decision-makers. Statistical regularization improves predictions and provides a general framework for fitting models with a large number of predictors, even those with collinearity, to simultaneously identify an optimal predictive model while conducting rigorous Bayesian model selection. Our framework is important for understanding population dynamics under a changing climate and has direct applications for making harvest and habitat management decisions.

Journal of Animal Ecology

Spatial patterns of aftershocks of shallow focus earthquakes in California and implications for deep focus earthquakes

Previous workers have pioneered statistical techniques to study the spatial distribution of aftershocks with respect to the focal mechanism of the main shock. Application of these techniques to deep focus earthquakes failed to show clustering of aftershocks near the nodal planes of the main shocks. To better understand the behavior of these statistics, this study applies them to the aftershocks of six large shallow focus earthquakes in California (August 6, 1979, Coyote Lake; May 2, 1983, Coalinga; April 24, 1984, Morgan Hill; August 4, 1985, Kettleman Hills; July 8, 1986, North Palm Springs; and October 1, 1987, Whittier Narrows). The large number of aftershocks accurately located by dense local networks allows us to treat these aftershock sequences individually instead of combining them, as was done for the deep earthquakes. The results for individual sequences show significant clustering about the closest nodal plane and the strike direction for five of the sequences and about the presumed fault plane for all six sequences. This implies that the previously developed method does work properly. Nonrandom behavior was also found about the slip directions, the P axis, the T axis, and the B axis, but this is probably caused by the lack of independence between these axes and the previously mentioned features of the focal mechanisms. Given that the method does work and that deep aftershocks were not shown to cluster about the main shock nodal planes, the shallow focus data were used to simulate the deep focus study. The goal is to determine if there are artificial factors that make clustering in the deep focus data unobservable. To more closely mimic the work on deep earthquakes, the largest aftershocks from each of the six sequences were combined and studied with respect to their respective main shock focal mechanisms. This reduced the significance of the clustering about the focal mechanism parameters, but not below 95% confidence. Gaussian noise was then added to the aftershock hypocenters in order to determine if the larger hypocentral and focal mechanism errors in the deep focus data could account for the previous negative result. The conclusion is that the following reasons are sufficient to explain the lack of clustering about the main shock nodal planes for the deep focus aftershocks: the need to combine aftershocks from several sequences, the size of the hypocentral location and focal mechanism errors, and the alignment of distant aftershocks with the Wadati-Benioff zone.

Journal of Geophysical Research Solid Earth

The effect of stress changes on time-dependent earthquake probabilities for the central Wasatch Fault Zone, Utah, USA

Static and quasi-static Coulomb stress changes produced by large earthquakes can modify the probability of occurrence of subsequent events on neighboring faults. This approach is based on physical (Coulomb stress changes) and statistical (probability calculations) models, which are influenced by the quality and quantity of data available in the study region. Here, we focus on the Wasatch Fault Zone (WFZ), a well-studied active normal fault system having abundant geologic and paleoseismological data. Paleoseismological trench investigations of the WFZ indicate that at least 24 large, surface-faulting earthquakes have ruptured the fault’s five central, 35–59-km long segments since ~7 ka. Our goal is to determine if the stress changes due to the youngest paleoevents have significantly modified the present-day probability of occurrence of large earthquakes on each of the segments. For each segment, we modeled the cumulative (coseismic + postseismic) Coulomb stress changes (∆CFScum) due to earthquakes younger than the most recent event on the segment in question and applied the resulting values to the time-dependent probability calculations. Results from the Coulomb stress modeling suggest that the Brigham City, Salt Lake City, and Provo segments have accumulated ∆CFScum larger than 10 bars, whereas the Weber segment has experienced a stress decrease of 5 bars, in the scenario of recent rupture of the Great Salt Lake fault to the west. Probability calculations predict high probability of occurrence for the Brigham City and Salt Lake City segments, due to their long elapsed times (>1-2 ka) when compared to the Weber, Provo, and Nephi segments (< 1 ka). The range of calculated coefficients of variation (CV) has a large influence on the final probabilities, mostly in the case of the Brigham City segment. Finally, when the Coulomb stress and the probability models are combined, our results indicate that the ∆CFScum resulting from earthquakes postdating the youngest events on each of the five segments significantly affects the probability calculations for three of the segments: Brigham City, Salt Lake City, and Provo. The probability of occurrence of a large earthquake in the next 50 years on these three segments may therefore be underestimated if a time-independent approach, or a time-dependent approach that does not consider ∆CFS, is adopted.

Utah

Concerns regarding a call for pluralism of information theory and hypothesis testing

1. Stephens et al. (2005) argue for 'pluralism' in statistical analysis, combining null hypothesis testing and information-theoretic (I-T) methods. We show that I-T methods are more informative even in single variable problems and we provide an ecological example. 2. I-T methods allow inferences to be made from multiple models simultaneously. We believe multimodel inference is the future of data analysis, which cannot be achieved with null hypothesis-testing approaches. 3. We argue for a stronger emphasis on critical thinking in science in general and less reliance on exploratory data analysis and data dredging. Deriving alternative hypotheses is central to science; deriving a single interesting science hypothesis and then comparing it to a default null hypothesis (e.g. 'no difference') is not an efficient strategy for gaining knowledge. We think this single-hypothesis strategy has been relied upon too often in the past. 4. We clarify misconceptions presented by Stephens et al. (2005) . 5. We think inference should be made about models, directly linked to scientific hypotheses, and their parameters conditioned on data, Prob(Hj| data). I-T methods provide a basis for this inference. Null hypothesis testing merely provides a probability statement about the data conditioned on a null model, Prob(data |H0). 6. Synthesis and applications . I-T methods provide a more informative approach to inference. I-T methods provide a direct measure of evidence for or against hypotheses and a means to consider simultaneously multiple hypotheses as a basis for rigorous inference. Progress in our science can be accelerated if modern methods can be used intelligently; this includes various I-T and Bayesian methods.

Journal of Applied Ecology