Search USGSSearch

USGS · 70202999

Comparison of algorithms for replacing missing data in discriminant analysis

Abstract

We examined the impact of different methods for replacing missing data in discriminant analyses conducted on randomly generated samples from multivariate normal and non-normal distributions. The probabilities of correct classification were obtained for these discriminant analyses before and after randomly deleting data as well as after deleted data were replaced using: (1) variable means, (2) principal component projections, and (3) the EM algorithm. Populations compared were: (1) multivariate normal with covariance matrices ∑ 1 =∑ 2 , (2) multivariate normal with ∑ 1 ≠∑ 2 and (3) multivariate non-normal with ∑ 1 =∑ 2 . Differences in the probabilities of correct classification were most evident for populations with small Mahalanobis distances or high proportions of missing data. The three replacement methods performed similarly but all were better than non - replacement.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Daniel J. Twedt, D.S. Gill. 1992. Comparison of algorithms for replacing missing data in discriminant analysis. https://doi.org/10.1080/03610929208830864

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Modeling participation duration, with application to the North American Breeding Bird Survey

We consider “participation histories,” binary sequences consisting of alternating finite sequences of 1s and 0s, ending with an infinite sequence of 0s. Our work is motivated by a study of observer tenure in the North American Breeding Bird Survey (BBS). In our analysis, j indexes an observer’s years of service and X j is an indicator of participation in the survey; 0s interspersed among 1s correspond to years when observers did not participate, but subsequently returned to service. Of interest is the observer’s duration D = max { j : X j = 1}. Because observed records X = ( X 1 , X 2 ,..., X n ) 1 are of finite length, all that we can directly infer about duration is that D ⩾ max { j ⩽ n : X j = 1}; model-based analysis is required for inference about D . We propose models in which lengths of 0s and 1s sequences have distributions determined by the index j at which they begin; 0s sequences are infinite with positive probability, an estimable parameter. We found that BBS observers’ lengths of service vary greatly, with 25.3% participating for only a single year, 49.5% serving for 4 or fewer years, and an average duration of 8.7 years, producing an average of 7.7 counts.

Communications in Statistics - Theory and Methods

Line transect estimation of population size: the exponential case with grouped data

Gates, Marshall, and Olson (1968) investigated the line transect method of estimating grouse population densities in the case where sighting probabilities are exponential. This work is followed by a simulation study in Gates (1969). A general overview of line transect analysis is presented by Burnham and Anderson (1976). These articles all deal with the ungrouped data case. In the present article, an analysis of line transect data is formulated under the Gates framework of exponential sighting probabilities and in the context of grouped data.

Communications in Statistics - Theory and Methods

Efficiency and optimal allocation in the staggered entry design

The staggered entry design for survival analysis specifies that r left-truncated samples are to be used in estimation of a population survival function. The ith sample is taken at time Bi, from the subpopulation of individuals having survival time exceeding Bi. This paper investigates the performance of the staggered entry design relative to the usual design in which all samples have a common time origin. The staggered entry design is shown to be an attractive alternative, even when not necessitated by logistical constraints. The staggered entry design allows for increased precision in estimation of the right tail of the survival function, especially when some of the data may be censored. A trade-off between the range of values for which the increased precision occurs and the magnitude of the increased precision is demonstrated.

Communications in Statistics - Theory and Methods