Search USGSSearch

USGS · 1015122

A permutation test for quantile regression

Abstract

A drop in dispersion, F -ratio like, permutation test ( D ) for linear quantile regression estimates (0≤τ≤1) had relative power ≥1 compared to quantile rank score tests ( T ) for hypotheses on parameters other than the intercept. Power was compared for combinations of sample sizes ( n =20−300) and quantiles (τ=0.50−0.99) where both tests maintained valid Type I error rates in simulations with p =2 and 6 parameters in homogeneous and heterogeneous error models. The D test required two modifications of permuting residuals from null, reduced parameter models to maintain correct Type I error rates when null models were constrained through the origin or included multiple parameters. A double permutation scheme was used when null models were constrained through the origin and all but 1 of the zero residuals were deleted for null models with multiple parameters. Although there was considerable overlap in sample size, quantiles, and hypotheses where both the D and rank score tests maintained correct Type I error rates, we identified regions at smaller n and more extreme quantiles where one or the other maintained better error rates. Confidence intervals on parameters for an ecological application relating Lahontan cutthroat trout densities to stream channel width:depth were estimated by test inversion, demonstrating a smoother pattern of slightly narrower intervals across quantiles than those provided by the rank score test.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Brian S. Cade, Jon D. Richards. 2006. A permutation test for quantile regression. https://doi.org/10.1198/108571106x96835

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Dynamic population models with temporal preferential sampling to infer phenology

To study population dynamics, ecologists and wildlife biologists typically use relative abundance data, which may be subject to temporal preferential sampling. Temporal preferential sampling occurs when the times at which observations are made and the latent process of interest are conditionally dependent. To account for preferential sampling, we specify a Bayesian hierarchical abundance model that considers the dependence between observation times and the ecological process of interest. The proposed model improves relative abundance estimates during periods of infrequent observation and accounts for temporal preferential sampling in discrete time. Additionally, our model facilitates posterior inference for population growth rates and mechanistic phenometrics. We apply our model to analyze both simulated data and mosquito count data collected by the National Ecological Observatory Network. In the second case study, we characterize the population growth rate and relative abundance of several mosquito species in the Aedes genus. Supplementary materials accompanying this paper appear on-line.

Journal of Agricultural, Biological, and Environme

Techniques to improve ecological interpretability of black box machine learning models

Statistical modeling of ecological data is often faced with a large number of variables as well as possible nonlinear relationships and higher-order interaction effects. Gradient boosted trees (GBT) have been successful in addressing these issues and have shown a good predictive performance in modeling nonlinear relationships, in particular in classification settings with a categorical response variable. They also tend to be robust against outliers. However, their black-box nature makes it difficult to interpret these models. We introduce several recently developed statistical tools to the environmental research community in order to advance interpretation of these black-box models. To analyze the properties of the tools, we applied gradient boosted trees to investigate biological health of streams within the contiguous USA, as measured by a benthic macroinvertebrate biotic index. Based on these data and a simulation study, we demonstrate the advantages and limitations of partial dependence plots (PDP), individual conditional expectation (ICE) curves and accumulated local effects (ALE) in their ability to identify covariate–response relationships. Additionally, interaction effects were quantified according to interaction strength (IAS) and Friedman’s H 2 "> H 2 statistic. Interpretable machine learning techniques are useful tools to open the black-box of gradient boosted trees in the environmental sciences. This finding is supported by our case study on the effect of impervious surface on the benthic condition, which agrees with previous results in the literature. Overall, the most important variables were ecoregion, bed stability, watershed area, riparian vegetation and catchment slope. These variables were also present in most identified interaction effects. In conclusion, graphical tools (PDP, ICE, ALE) enable visualization and easier interpretation of GBT but should be supported by analytical statistical measures. Future methodological research is needed to investigate the properties of interaction tests. Supplementary materials accompanying this paper appear on-line.

Journal of Agricultural, Biological, and Environme

Extreme value-based methods for modeling elk yearly movements

Species range shifts and the spread of diseases are both likely to be driven by extreme movements, but are difficult to statistically model due to their rarity. We propose a statistical approach for characterizing movement kernels that incorporate landscape covariates as well as the potential for heavy-tailed distributions. We used a spliced distribution for distance travelled paired with a resource selection function to model movements biased toward preferred habitats. As an example, we used data from 704 annual elk movements around the Greater Yellowstone Ecosystem from 2001 to 2015. Yearly elk movements were both heavy-tailed and biased away from high elevations during the winter months. We then used a simulation to illustrate how these habitat effects may alter the rate of disease spread using our estimated movement kernel relative to a more traditional approach that does not include landscape covariates. Supplementary materials accompanying this paper appear online.

Journal of Agricultural, Biological, and Environme