Search USGSSearch

USGS · 70018317

An objective replacement method for censored geochemical data

Abstract

Geochemical data are commonly censored, that is, concentrations for some samples are reported as "less than" or "greater than" some value. Censored data hampers statistical analysis because certain computational techniques used in statistical analysis require a complete set of uncensored data. We show that the simple substitution method for creating an uncensored dataset, e.g., replacement by 3/4 times the detection limit, has serious flaws, and we present an objective method to determine the replacement value. Our basic premise is that the replacement value should equal the mean of the actual values represented by the qualified data. We adapt the maximum likelihood approach (Cohen, 1961) to estimate this mean. This method reproduces the mean and skewness as well or better than a simple substitution method using 3/4 of the lower detection limit or 3/4 of the upper detection limit. For a small proportion of "less than" substitutions, a simple-substitution replacement factor of 0.55 is preferable to 3/4; for a small proportion of "greater than" substitutions, a simple-substitution replacement factor of 1.7 is preferable to 4/3, provided the resulting replacement value does not exceed 100%. For more than 10% replacement, a mean empirical factor may be used. However, empirically determined simple-substitution replacement factors usually vary among different data sets and are less reliable with more replacements. Therefore, a maximum likelihood method is superior in general. Theoretical and empirical analyses show that true replacement factors for "less thans" decrease in magnitude with more replacements and larger standard deviation; those for "greater thans" increase in magnitude with more replacements and larger standard deviation. In contrast to any simple substitution method, the maximum likelihood method reproduces these variations. Using the maximum likelihood method for replacing "less thans" in our sample data set, correlation coefficients were reasonably accurately estimated in 90% of the cases for as much as 40% replacement and in 60% of the cases for 80% replacement. These results suggest that censored data can be utilized more than is commonly realized. ?? 1993 International Association for Mathematical Geology.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

R.F. Sanford, C. T. Pierson, R. A. Crovelli. 1993. An objective replacement method for censored geochemical data. https://doi.org/10.1007/bf00890676

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Declustering of clustered preferential sampling for histogram and semivariogram inference

Measurements of attributes obtained more as a consequence of business ventures than sampling design frequently result in samplings that are preferential both in location and value, typically in the form of clusters along the pay. Preferential sampling requires preprocessing for the purpose of properly inferring characteristics of the parent population, such as the cumulative distribution and the semivariogram. Consideration of the distance to the nearest neighbor allows preparation of resampled sets that produce comparable results to those from previously proposed methods. Clustered sampling of size 140, taken from an exhaustive sampling, is employed to illustrate this approach. ?? International Association for Mathematical Geology 2007.

Mathematical Geology

Comparison of two probability distributions used to model sizes of undiscovered oil and gas accumulations: Does the tail wag the assessment?

Undiscovered oil and gas assessments are commonly reported as aggregate estimates of hydrocarbon volumes. Potential commercial value and discovery costs are, however, determined by accumulation size, so engineers, economists, decision makers, and sometimes policy analysts are most interested in projected discovery sizes. The lognormal and Pareto distributions have been used to model exploration target sizes. This note contrasts the outcomes of applying these alternative distributions to the play level assessments of the U.S. Geological Survey's 1995 National Oil and Gas Assessment. Using the same numbers of undiscovered accumulations and the same minimum, medium, and maximum size estimates, substitution of the shifted truncated lognormal distribution for the shifted truncated Pareto distribution reduced assessed undiscovered oil by 16% and gas by 15%. Nearly all of the volume differences resulted because the lognormal had fewer larger fields relative to the Pareto. The lognormal also resulted in a smaller number of small fields relative to the Pareto. For the Permian Basin case study presented here, reserve addition costs were 20% higher with the lognormal size assumption. ?? 2002 International Association for Mathematical Geology.

Mathematical Geology

Uncertainty estimation for resource assessment-an application to coal

The U.S. Geological Survey is conducting a national assessment of coal resources. As part of that assessment, a geostatistical procedure has been developed to estimate the uncertainty of coal resources for the historical categories of geological assurance: measured, indicated, inferred, and hypothetical coal. Data consist of spatially clustered coal thickness measurements from coal beds and/or zones that cover, in some cases, several thousand square kilometers. Our procedure involved trend removal, an examination of spatial correlation, computation of a sample semivariogram, and fitting a semivariogram model. This model provided standard deviations for the uncertainty estimates. The number of sample points (drill holes) in each historical category also was estimated. Measurement error in the thickness of the coal bed/zone was obtained from the fitted model or supplied exogenously. From this information approximate estimates of uncertainty on the historical categories were computed. We illustrate the methodology using drill hole data from the Harmon coal bed located in southwestern North Dakota. The methodology will be applied to approximately 50 coal data sets.

Mathematical Geology