Search USGSSearch

USGS · 70045681

A bootstrap estimation scheme for chemical compositional data with nondetects

Abstract

The bootstrap method is commonly used to estimate the distribution of estimators and their associated uncertainty when explicit analytic expressions are not available or are difficult to obtain. It has been widely applied in environmental and geochemical studies, where the data generated often represent parts of whole, typically chemical concentrations. This kind of constrained data is generically called compositional data, and they require specialised statistical methods to properly account for their particular covariance structure. On the other hand, it is not unusual in practice that those data contain labels denoting nondetects, that is, concentrations falling below detection limits. Nondetects impede the implementation of the bootstrap and represent an additional source of uncertainty that must be taken into account. In this work, a bootstrap scheme is devised that handles nondetects by adding an imputation step within the resampling process and conveniently propagates their associated uncertainly. In doing so, it considers the constrained relationships between chemical concentrations originated from their compositional nature. Bootstrap estimates using a range of imputation methods, including new stochastic proposals, are compared across scenarios of increasing difficulty. They are formulated to meet compositional principles following the log-ratio approach, and an adjustment is introduced in the multivariate case to deal with nonclosed samples. Results suggest that nondetect bootstrap based on model-based imputation is generally preferable. A robust approach based on isometric log-ratio transformations appears to be particularly suited in this context. Computer routines in the R statistical programming language are provided.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Javier Palarea-Albaladejo, J.A Martin-Fernandez, Ricardo A. Olea. 2014-04-02. A bootstrap estimation scheme for chemical compositional data with nondetects. https://doi.org/10.1002/cem.2621

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Evaluation of spectral collection strategies for identification of Dalbergia spp. using handheld laser-induced breakdown spectroscopy

The illegal timber trade has significant impact on the survival of endangered tropical hardwood species like Dalbergia spp. (rosewood), a world-wide protected genus from the Convention on International Trade in Endangered Species of Wild Fauna and Flora (CITES). Due to increased threat to Dalbergia spp., and lack of action to reduce threats, port of entry analysis methods are required to identify Dalbergia spp. Handheld laser-induced breakdown spectroscopy (LIBS) has been shown to be capable of identifying species and establishing provenance of Dalbergia spp. and other tropical hardwoods, but analysis methods for this work have yet to be investigated in detail. The present work investigates five well-known algorithms—partial least squares discriminant analysis (PLS-DA), classification and regression trees (CART), k -nearest neighbor ( k -NN), random forest (RF), and support vector machine (SVM)—two training/test set sampling regimes, and data collection at two signal-to-noise (S/N) ratios to assess the potential for handheld LIBS analyses. Additionally, imbalanced classes are addressed. For this application, SVM and RF yield near identical results (though RF takes nearly 100 longer to compute), while the S/N ratio has a significant effect on model success assuming all else is equal. It was found that forming a training set with replicate low S/N analyses can perform as well as higher precision training sets for true prediction, even if the predicted samples have low signal to noise! This work confirms handheld LIBS analyzers can provide a viable method for classification of hardwood species, even within the same genus.

Journal of Chemometrics

Pattern recognition analysis and classification modeling of selenium-producing areas

Established chemometric and geochemical techniques were applied to water quality data from 23 National Irrigation Water Quality Program (NIWQP) study areas in the Western United States. These techniques were applied to the NIWQP data set to identify common geochemical processes responsible for mobilization of selenium and to develop a classification model that uses major-ion concentrations to identify areas that contain elevated selenium concentrations in water that could pose a hazard to water fowl. Pattern recognition modeling of the simple-salt data computed with the SNORM geochemical program indicate three principal components that explain 95% of the total variance. A three-dimensional plot of PC 1, 2 and 3 scores shows three distinct clusters that correspond to distinct hydrochemical facies denoted as facies 1, 2 and 3. Facies 1 samples are distinguished by water samples without the CaCO 3 simple salt and elevated concentrations of NaCl, CaSO 4 , MgSO 4 and Na 2 SO 4 simple salts relative to water samples in facies 2 and 3. Water samples in facies 2 are distinguished from facies 1 by the absence of the MgSO 4 simple salt and the presence of the CaCO 3 simple salt. Water samples in facies 3 are similar to samples in facies 2, with the absence of both MgSO 4 and CaSO 4 simple salts. Water samples in facies 1 have the largest selenium concentration (10 μg l −1 ), compared to a median concentration of 2·0 μg l −1 and less than 1·0 μg l −1 for samples in facies 2 and 3. A classification model using the soft independent modeling by class analogy (SIMCA) algorithm was constructed with data from the NIWQP study areas. The classification model was successful in identifying water samples with a selenium concentration that is hazardous to some species of water-fowl from a test data set comprised of 2,060 water samples from throughout Utah and Wyoming. Application of chemometric and geochemical techniques during data synthesis analysis of multivariate environmental databases from other national-scale environmental programs such as the NIWQP could also provide useful insights for addressing ‘real world’ environmental problems.

Journal of Chemometrics