Search USGSSearch

USGS · 70035715

Kolmogorov-Smirnov test for spatially correlated data

Abstract

The Kolmogorov-Smirnov test is a convenient method for investigating whether two underlying univariate probability distributions can be regarded as undistinguishable from each other or whether an underlying probability distribution differs from a hypothesized distribution. Application of the test requires that the sample be unbiased and the outcomes be independent and identically distributed, conditions that are violated in several degrees by spatially continuous attributes, such as topographical elevation. A generalized form of the bootstrap method is used here for the purpose of modeling the distribution of the statistic D of the Kolmogorov-Smirnov test. The innovation is in the resampling, which in the traditional formulation of bootstrap is done by drawing from the empirical sample with replacement presuming independence. The generalization consists of preparing resamplings with the same spatial correlation as the empirical sample. This is accomplished by reading the value of unconditional stochastic realizations at the sampling locations, realizations that are generated by simulated annealing. The new approach was tested by two empirical samples taken from an exhaustive sample closely following a lognormal distribution. One sample was a regular, unbiased sample while the other one was a clustered, preferential sample that had to be preprocessed. Our results show that the p-value for the spatially correlated case is always larger that the p-value of the statistic in the absence of spatial correlation, which is in agreement with the fact that the information content of an uncorrelated sample is larger than the one for a spatially correlated sample of the same size. ?? Springer-Verlag 2008.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ricardo A. Olea, V. Pawlowsky-Glahn. 2008-07-29. Kolmogorov-Smirnov test for spatially correlated data. https://doi.org/10.1007/s00477-008-0255-1

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Characterizing projected future droughts for south Florida (2056–2095)

Balance anomalies, defined as the deviation of monthly precipitation minus reference evapotranspiration from their long-term monthly historical means (1950–2005), were computed for regions in south Florida and temporally averaged over 6- and 12-month timescales to identify meteorological drought events during a historical (1950–2005) and future (2056–2095) period of interest for 40 CMIP5 general circulation models (GCM) and scenario combinations downscaled by the Multivariate Adaptive Constructed Analogs method. Under the assumption of stomatal resistance ( r s ) remaining constant at the historical standard value (70 s/m), 81% of models project declines in monthly balance anomalies in the future compared to historical, with multimodel ensemble mean declines of 3.7 in/year under RCP4.5 and 8.7 in/year under RCP8.5. Drought events were identified from the downscaled model projections, their characteristics (duration and intensity) extracted, and their historical joint distributions validated against those derived from historical observational datasets. The future joint distributions of drought characteristics were compared across models using hierarchical clustering. A climate model summary plot and table were developed based on these methods to guide climate model selection for hydrologic modeling in support of water-supply planning at the South Florida Water Management District (SFWMD). The model summary plot for the entire SFWMD shows that 35% of GCM/training-dataset combinations have historical joint distributions of drought characteristics that are significantly different at the 10% level from those derived from observational gridded data, whereas 39% of GCM/scenario/training-dataset combinations have future joint distributions that are significantly different from historical. A sensitivity analysis was performed assuming r s increasing with increasing CO 2 .

Florida

Self-organizing maps for compositional data: coal combustion products of a Wyoming power plant

A self-organizing map (SOM) is a non-linear projection of a D-dimensional data set, where the distance among observations is approximately preserved on to a lower dimensional space. The SOM arranges multivariate data based on their similarity to each other by allowing pattern recognition leading to easier interpretation of higher dimensional data. The SOM algorithm allows for selection of different map topologies, distances and parameters, which determine how the data will be organized on the map. In the particular case of compositional data (such as elemental, mineralogical, or maceral abundance), the sample space is governed by Aitchison geometry and extra steps are required prior to their SOM analysis. Following the principle of working on log-ratio coordinates, the simplicial operations and the Aitchison distance, which are appropriate elements for the SOM, are presented. With this structure developed, a SOM using Aitchison geometry is applied to properly interpret elemental data from combustion products (bottom ash, fly ash, and economizer fly ash) in a Wyoming coal-fired power plant. Results from this effort provide knowledge about the differences between the ash composition in the coal combustion process.

Stochastic Environmental Research and Risk Assessm

Advancements in hydrochemistry mapping: methods and application to groundwater arsenic and iron concentrations in Varanasi, Uttar Pradesh, India

The area east of Varanasi is one of numerous places along the watershed of the Ganges River with groundwater concentrations of arsenic surpassing the maximum value of 10 parts per billion (ppb) recommended by the World Health Organization in drinking water. Here we apply geostatistics and compositional data analysis for the mapping of arsenic and iron to help in understanding the conditions leading to the occurrence of elevated level of arsenic in groundwater. The methodology allows for displaying concentrations of arsenic and iron as maps consistent with the limited information from 95 water wells across an area of approximately 210 km 2 ; visualization of the uncertainty associated with the sampling; and summary of the findings in the form of probability maps. For thousands of years, Varanasi has been on the erosional side in a meander of the river that is free of arsenic values above 10 ppb. Maps reveal two anomalies of high arsenic concentrations on the depositional side of the valley, which has started seeing urban development. The methodology using geostatistics combined with compositional data analysis is completely general, so this study could be used as a prototype for hydrochemistry mapping in other areas.

Uttar Pradesh